| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-29 | |||
| 13:12:24 | mriedem | or not | |
| 13:12:50 | mriedem | you'll have to get the licensed version from the author :) | |
| 13:13:39 | cdent | yeah, that’s where I starte4d | |
| 13:28:44 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:28:44 | johnthetubaguy | fried_rice: added more details on that ironic rebalance bug from my findings this morning https://bugs.launchpad.net/nova/+bug/1714248: | |
| 13:29:32 | johnthetubaguy | added a patch with some comments in to make it clear where I believe the error occurs | |
| 13:30:07 | openstackgerrit | Elod Illes proposed openstack/nova stable/ocata: WIP: Functional test for regression bug #1713783 https://review.openstack.org/505160 | |
| 13:30:08 | openstack | bug 1713783 in OpenStack Compute (nova) ocata "After failed evacuation the recovered source compute tries to delete the instance" [High,Triaged] https://launchpad.net/bugs/1713783 | |
| 13:32:08 | fried_rice | johnthetubaguy Okay, haven't dug in fully, but... | |
| 13:32:14 | fried_rice | johnthetubaguy It is legal to rename a RP. | |
| 13:32:25 | fried_rice | Would that solve it, or do you actually need to change the RP's UUID too? | |
| 13:32:35 | johnthetubaguy | name is the same, uuid is different | |
| 13:32:36 | cdent | mriedem: in case you missed it before gibi and I got http://burndown.peermore.com/nova-notification/ going | |
| 13:32:44 | fried_rice | boo | |
| 13:32:50 | johnthetubaguy | yeah :( | |
| 13:32:52 | cdent | using his little framework would be easy to burn lots of other things down | |
| 13:33:14 | fried_rice | johnthetubaguy And the original UUID comes from where? | |
| 13:33:25 | johnthetubaguy | fried_rice: ComputeNode.uuid | |
| 13:33:35 | fried_rice | Which makes sense, I suppose. | |
| 13:33:41 | johnthetubaguy | ish, yeah | |
| 13:34:40 | fried_rice | johnthetubaguy Is there an entity in this world that stays the same on this failover deal? | |
| 13:35:17 | johnthetubaguy | fried_rice: ironic node uuid, i.e. ComputeNode.hypervisor_hostname and ResourceProvider.name | |
| 13:35:17 | superdan | finucannot: ack, I'll allow you to retain rights to conf/network | |
| 13:37:03 | fried_rice | johnthetubaguy Ohh, so hang on, it really *doesn't* make much sense for the RP UUID to be the compute host UUID. Or at least, it would make just as much sense for it to be the ironic node UUID. | |
| 13:37:39 | fried_rice | Though tbh, I'm a tad confused as to why those aren't the same thing. | |
| 13:37:45 | johnthetubaguy | fried_rice: except when get the thing back, the uuid, we want it to always be the same thing, in our case a ComputeNode | |
| 13:38:07 | johnthetubaguy | I am not 100% sure at this point | |
| 13:38:18 | fried_rice | Dangit, I've slept since we started this conversation, need to get the model straight in my head again. | |
| 13:38:55 | fried_rice | Hardware-wise, there's a set of machines. These are called "ironic nodes". | |
| 13:38:57 | johnthetubaguy | oh wait, that could be the answer, for ironic, we could mess with the ComputeNode.uuid | |
| 13:39:24 | fried_rice | Each "ironic node" gets modeled as a separate RP, right? | |
| 13:39:28 | johnthetubaguy | fried_rice: https://developer.openstack.org/api-ref/baremetal/#list-nodes | |
| 13:39:32 | johnthetubaguy | yeah | |
| 13:39:47 | mriedem | cdent: nice | |
| 13:39:55 | fried_rice | And the nova-compute process runs... where? On a totally separate machine? | |
| 13:40:14 | cdent | mriedem: runs on a cron job every hour. took an _entire five minutes_ to set up | |
| 13:40:37 | mriedem | take a break | |
| 13:41:01 | johnthetubaguy | fried_rice: yeah, noramlly on some controller node | |
| 13:41:42 | fried_rice | johnthetubaguy Which is called ComputeNode? | |
| 13:42:21 | johnthetubaguy | not really | |
| 13:42:27 | johnthetubaguy | ComputeNode is the Nova DB table | |
| 13:42:34 | johnthetubaguy | there is one entry for each ironic node | |
| 13:42:53 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275 | |
| 13:43:03 | johnthetubaguy | i.e. one n-cpu Service has one or more compute nodes | |
| 13:43:13 | fried_rice | johnthetubaguy So why isn't the ComputeNode.uuid the same as the ironic node UUID, if they represent the same thing? | |
| 13:43:32 | johnthetubaguy | fried_rice: I think just history, I am just looking at how hard that would be right now | |
| 13:44:00 | superdan | because computenode is a nova structure | |
| 13:44:10 | superdan | which is used for everything else too, | |
| 13:44:33 | superdan | and because you could have generated those uuids from the existing records when things were offline during a save or forced migration | |
| 13:45:32 | johnthetubaguy | superdan: agreed | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Stop joining on system_metadata when listing instances https://review.openstack.org/508335 | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove old compat code from servers ViewBuilder._get_metadata https://review.openstack.org/508326 | |
| 13:45:48 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove system_metadata loading in Instance._load_flavor https://review.openstack.org/508357 | |
| 13:46:06 | johnthetubaguy | superdan: it could be the simplest fix, but all the fixes seem terrible | |
| 13:46:19 | superdan | I wasn't following what the problem was | |
| 13:46:28 | johnthetubaguy | https://bugs.launchpad.net/nova/+bug/1714248 | |
| 13:46:29 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:46:33 | mriedem | the rp uuid isn't the ironic node uuid | |
| 13:46:36 | mriedem | it's the compute node uuid | |
| 13:46:40 | johnthetubaguy | actually, that fix isn't simple in the upgrade case, naturally | |
| 13:46:47 | mriedem | johnthetubaguy: was going to say, | |
| 13:46:50 | mriedem | existing computes... | |
| 13:47:13 | mriedem | you could change the compute node uuid...but not sure what weirdness would happen from that | |
| 13:47:18 | superdan | that's hard | |
| 13:47:43 | superdan | I don't think changing the uuid of an existing provider, especially with allocations makes sense | |
| 13:48:16 | superdan | the ironic driver probably needs to figure this out some way to avoid having to re-do the allocations | |
| 13:48:37 | johnthetubaguy | problem is the allocation is on the ComputeNode.uuid | |
| 13:48:38 | superdan | otherwise we have a race I think | |
| 13:48:45 | superdan | right I know | |
| 13:48:45 | johnthetubaguy | yeah, the race is horrible | |
| 13:50:37 | superdan | I'm not sure what that means | |
| 13:50:56 | superdan | so one thing we could do, | |
| 13:51:03 | leakypipes | johnthetubaguy: nova host aggs never worked with ironic nodes, though. | |
| 13:51:15 | superdan | is make sure that the RP name is the uuid of the ironic node | |
| 13:51:24 | leakypipes | johnthetubaguy: just got off the call, btw, reading back | |
| 13:51:34 | johnthetubaguy | that is true today | |
| 13:51:34 | fried_rice | Can we do a fake migration, using the "migration_uuid" thingy, to avoid the race? | |
| 13:51:34 | superdan | which would be a path for the ironic driver getting the node to find out the uuid of the matching provider | |
| 13:51:49 | superdan | fried_rice: doesn't really help | |
| 13:52:06 | mriedem | right the ironic rp name is the ironic node uuid today | |
| 13:52:17 | johnthetubaguy | oh, you mean when we create the new compute node, we use the existing uuid, so we just rename the compute node? | |
| 13:52:19 | mriedem | since the rp.name == compute_node.hypervisor_hostname == ironic node.uuid | |
| 13:52:34 | superdan | mriedem: is it? | |
| 13:52:40 | mriedem | yar | |
| 13:52:44 | superdan | well in that case, we can just look it up and reparent the compute node object then | |
| 13:52:44 | johnthetubaguy | yeah, thats correct | |
| 13:52:51 | superdan | computenode.host = self.host | |
| 13:52:53 | superdan | done | |
| 13:52:55 | johnthetubaguy | yeah, instead of add a new one | |
| 13:53:10 | johnthetubaguy | that's what I just attempted to say, badly | |
| 13:53:20 | superdan | I was going to suggest that we set it so we could do that once it was, but if it's already done we should be able to do that right away | |
| 13:54:15 | johnthetubaguy | I will take a look at doing that | |
| 13:54:29 | johnthetubaguy | in the ironic driver, or where we create the compute nodes already? | |
| 13:54:46 | superdan | um | |
| 13:54:59 | superdan | I'd have to go look at that.. I guess there's a bit of a layering violation in there some where | |
| 13:55:19 | superdan | maybe we can just make compute manager check for this situation and do the re-homing | |
| 13:56:36 | johnthetubaguy | yeah, I just worry about hostname clashes | |
| 13:56:38 | johnthetubaguy | in none ironic cases | |
| 13:56:45 | superdan | well, | |
| 13:56:54 | superdan | the hostnames have to be unique today or nothing works anyway | |