| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-29 | |||
| 13:09:40 | johnthetubaguy | yeah, +1 | |
| 13:09:54 | johnthetubaguy | so this bug is probably more important than I first thought: | |
| 13:10:00 | johnthetubaguy | https://bugs.launchpad.net/nova/+bug/1714248 | |
| 13:10:02 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:10:09 | mriedem | we could put another field in the response from the get_inventory method that includes an optional rp_uuid field or something | |
| 13:10:25 | mriedem | to tell the RT what to use | |
| 13:10:35 | mriedem | although, | |
| 13:10:43 | mriedem | that might break how scheduling works | |
| 13:10:45 | johnthetubaguy | but then we will fail looking up the correct rabbit queue I presume? | |
| 13:10:52 | mriedem | because the scheduler will lookup a compute node from placement's response by uuid | |
| 13:11:04 | mriedem | this has nothing to do with mq | |
| 13:11:11 | mriedem | are you thinking the cell mapping stuff? | |
| 13:11:33 | johnthetubaguy | well the mapping from the placement object, to the compute node you build the instance on | |
| 13:11:39 | cdent | leakypipes: are the slides from your scheduler wars presentations on teh internets somewhere (not the video) | |
| 13:11:52 | mriedem | cdent: the slides would also be on the summit video site | |
| 13:12:18 | cdent | hmmm | |
| 13:12:24 | mriedem | or not | |
| 13:12:24 | mriedem | https://www.openstack.org/videos/boston-2017/scheduler-wars-a-new-hope | |
| 13:12:50 | mriedem | you'll have to get the licensed version from the author :) | |
| 13:13:39 | cdent | yeah, that’s where I starte4d | |
| 13:28:44 | johnthetubaguy | fried_rice: added more details on that ironic rebalance bug from my findings this morning https://bugs.launchpad.net/nova/+bug/1714248: | |
| 13:28:44 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:29:32 | johnthetubaguy | added a patch with some comments in to make it clear where I believe the error occurs | |
| 13:30:07 | openstackgerrit | Elod Illes proposed openstack/nova stable/ocata: WIP: Functional test for regression bug #1713783 https://review.openstack.org/505160 | |
| 13:30:08 | openstack | bug 1713783 in OpenStack Compute (nova) ocata "After failed evacuation the recovered source compute tries to delete the instance" [High,Triaged] https://launchpad.net/bugs/1713783 | |
| 13:32:08 | fried_rice | johnthetubaguy Okay, haven't dug in fully, but... | |
| 13:32:14 | fried_rice | johnthetubaguy It is legal to rename a RP. | |
| 13:32:25 | fried_rice | Would that solve it, or do you actually need to change the RP's UUID too? | |
| 13:32:35 | johnthetubaguy | name is the same, uuid is different | |
| 13:32:36 | cdent | mriedem: in case you missed it before gibi and I got http://burndown.peermore.com/nova-notification/ going | |
| 13:32:44 | fried_rice | boo | |
| 13:32:50 | johnthetubaguy | yeah :( | |
| 13:32:52 | cdent | using his little framework would be easy to burn lots of other things down | |
| 13:33:14 | fried_rice | johnthetubaguy And the original UUID comes from where? | |
| 13:33:25 | johnthetubaguy | fried_rice: ComputeNode.uuid | |
| 13:33:35 | fried_rice | Which makes sense, I suppose. | |
| 13:33:41 | johnthetubaguy | ish, yeah | |
| 13:34:40 | fried_rice | johnthetubaguy Is there an entity in this world that stays the same on this failover deal? | |
| 13:35:17 | superdan | finucannot: ack, I'll allow you to retain rights to conf/network | |
| 13:35:17 | johnthetubaguy | fried_rice: ironic node uuid, i.e. ComputeNode.hypervisor_hostname and ResourceProvider.name | |
| 13:37:03 | fried_rice | johnthetubaguy Ohh, so hang on, it really *doesn't* make much sense for the RP UUID to be the compute host UUID. Or at least, it would make just as much sense for it to be the ironic node UUID. | |
| 13:37:39 | fried_rice | Though tbh, I'm a tad confused as to why those aren't the same thing. | |
| 13:37:45 | johnthetubaguy | fried_rice: except when get the thing back, the uuid, we want it to always be the same thing, in our case a ComputeNode | |
| 13:38:07 | johnthetubaguy | I am not 100% sure at this point | |
| 13:38:18 | fried_rice | Dangit, I've slept since we started this conversation, need to get the model straight in my head again. | |
| 13:38:55 | fried_rice | Hardware-wise, there's a set of machines. These are called "ironic nodes". | |
| 13:38:57 | johnthetubaguy | oh wait, that could be the answer, for ironic, we could mess with the ComputeNode.uuid | |
| 13:39:24 | fried_rice | Each "ironic node" gets modeled as a separate RP, right? | |
| 13:39:28 | johnthetubaguy | fried_rice: https://developer.openstack.org/api-ref/baremetal/#list-nodes | |
| 13:39:32 | johnthetubaguy | yeah | |
| 13:39:47 | mriedem | cdent: nice | |
| 13:39:55 | fried_rice | And the nova-compute process runs... where? On a totally separate machine? | |
| 13:40:14 | cdent | mriedem: runs on a cron job every hour. took an _entire five minutes_ to set up | |
| 13:40:37 | mriedem | take a break | |
| 13:41:01 | johnthetubaguy | fried_rice: yeah, noramlly on some controller node | |
| 13:41:42 | fried_rice | johnthetubaguy Which is called ComputeNode? | |
| 13:42:21 | johnthetubaguy | not really | |
| 13:42:27 | johnthetubaguy | ComputeNode is the Nova DB table | |
| 13:42:34 | johnthetubaguy | there is one entry for each ironic node | |
| 13:42:53 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275 | |
| 13:43:03 | johnthetubaguy | i.e. one n-cpu Service has one or more compute nodes | |
| 13:43:13 | fried_rice | johnthetubaguy So why isn't the ComputeNode.uuid the same as the ironic node UUID, if they represent the same thing? | |
| 13:43:32 | johnthetubaguy | fried_rice: I think just history, I am just looking at how hard that would be right now | |
| 13:44:00 | superdan | because computenode is a nova structure | |
| 13:44:10 | superdan | which is used for everything else too, | |
| 13:44:33 | superdan | and because you could have generated those uuids from the existing records when things were offline during a save or forced migration | |
| 13:45:32 | johnthetubaguy | superdan: agreed | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove old compat code from servers ViewBuilder._get_metadata https://review.openstack.org/508326 | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Stop joining on system_metadata when listing instances https://review.openstack.org/508335 | |
| 13:45:48 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove system_metadata loading in Instance._load_flavor https://review.openstack.org/508357 | |
| 13:46:06 | johnthetubaguy | superdan: it could be the simplest fix, but all the fixes seem terrible | |
| 13:46:19 | superdan | I wasn't following what the problem was | |
| 13:46:28 | johnthetubaguy | https://bugs.launchpad.net/nova/+bug/1714248 | |
| 13:46:29 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:46:33 | mriedem | the rp uuid isn't the ironic node uuid | |
| 13:46:36 | mriedem | it's the compute node uuid | |
| 13:46:40 | johnthetubaguy | actually, that fix isn't simple in the upgrade case, naturally | |
| 13:46:47 | mriedem | johnthetubaguy: was going to say, | |
| 13:46:50 | mriedem | existing computes... | |
| 13:47:13 | mriedem | you could change the compute node uuid...but not sure what weirdness would happen from that | |
| 13:47:18 | superdan | that's hard | |
| 13:47:43 | superdan | I don't think changing the uuid of an existing provider, especially with allocations makes sense | |
| 13:48:16 | superdan | the ironic driver probably needs to figure this out some way to avoid having to re-do the allocations | |
| 13:48:37 | johnthetubaguy | problem is the allocation is on the ComputeNode.uuid | |
| 13:48:38 | superdan | otherwise we have a race I think | |
| 13:48:45 | johnthetubaguy | yeah, the race is horrible | |
| 13:48:45 | superdan | right I know | |
| 13:50:37 | superdan | I'm not sure what that means | |
| 13:50:56 | superdan | so one thing we could do, | |
| 13:51:03 | leakypipes | johnthetubaguy: nova host aggs never worked with ironic nodes, though. | |
| 13:51:15 | superdan | is make sure that the RP name is the uuid of the ironic node | |
| 13:51:24 | leakypipes | johnthetubaguy: just got off the call, btw, reading back | |
| 13:51:34 | superdan | which would be a path for the ironic driver getting the node to find out the uuid of the matching provider | |
| 13:51:34 | fried_rice | Can we do a fake migration, using the "migration_uuid" thingy, to avoid the race? | |
| 13:51:34 | johnthetubaguy | that is true today | |
| 13:51:49 | superdan | fried_rice: doesn't really help | |
| 13:52:06 | mriedem | right the ironic rp name is the ironic node uuid today | |
| 13:52:17 | johnthetubaguy | oh, you mean when we create the new compute node, we use the existing uuid, so we just rename the compute node? | |
| 13:52:19 | mriedem | since the rp.name == compute_node.hypervisor_hostname == ironic node.uuid | |
| 13:52:34 | superdan | mriedem: is it? | |