| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-29 | |||
| 13:44:33 | superdan | and because you could have generated those uuids from the existing records when things were offline during a save or forced migration | |
| 13:45:32 | johnthetubaguy | superdan: agreed | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove old compat code from servers ViewBuilder._get_metadata https://review.openstack.org/508326 | |
| 13:45:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Stop joining on system_metadata when listing instances https://review.openstack.org/508335 | |
| 13:45:48 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove system_metadata loading in Instance._load_flavor https://review.openstack.org/508357 | |
| 13:46:06 | johnthetubaguy | superdan: it could be the simplest fix, but all the fixes seem terrible | |
| 13:46:19 | superdan | I wasn't following what the problem was | |
| 13:46:28 | johnthetubaguy | https://bugs.launchpad.net/nova/+bug/1714248 | |
| 13:46:29 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 13:46:33 | mriedem | the rp uuid isn't the ironic node uuid | |
| 13:46:36 | mriedem | it's the compute node uuid | |
| 13:46:40 | johnthetubaguy | actually, that fix isn't simple in the upgrade case, naturally | |
| 13:46:47 | mriedem | johnthetubaguy: was going to say, | |
| 13:46:50 | mriedem | existing computes... | |
| 13:47:13 | mriedem | you could change the compute node uuid...but not sure what weirdness would happen from that | |
| 13:47:18 | superdan | that's hard | |
| 13:47:43 | superdan | I don't think changing the uuid of an existing provider, especially with allocations makes sense | |
| 13:48:16 | superdan | the ironic driver probably needs to figure this out some way to avoid having to re-do the allocations | |
| 13:48:37 | johnthetubaguy | problem is the allocation is on the ComputeNode.uuid | |
| 13:48:38 | superdan | otherwise we have a race I think | |
| 13:48:45 | johnthetubaguy | yeah, the race is horrible | |
| 13:48:45 | superdan | right I know | |
| 13:50:37 | superdan | I'm not sure what that means | |
| 13:50:56 | superdan | so one thing we could do, | |
| 13:51:03 | leakypipes | johnthetubaguy: nova host aggs never worked with ironic nodes, though. | |
| 13:51:15 | superdan | is make sure that the RP name is the uuid of the ironic node | |
| 13:51:24 | leakypipes | johnthetubaguy: just got off the call, btw, reading back | |
| 13:51:34 | superdan | which would be a path for the ironic driver getting the node to find out the uuid of the matching provider | |
| 13:51:34 | fried_rice | Can we do a fake migration, using the "migration_uuid" thingy, to avoid the race? | |
| 13:51:34 | johnthetubaguy | that is true today | |
| 13:51:49 | superdan | fried_rice: doesn't really help | |
| 13:52:06 | mriedem | right the ironic rp name is the ironic node uuid today | |
| 13:52:17 | johnthetubaguy | oh, you mean when we create the new compute node, we use the existing uuid, so we just rename the compute node? | |
| 13:52:19 | mriedem | since the rp.name == compute_node.hypervisor_hostname == ironic node.uuid | |
| 13:52:34 | superdan | mriedem: is it? | |
| 13:52:40 | mriedem | yar | |
| 13:52:44 | johnthetubaguy | yeah, thats correct | |
| 13:52:44 | superdan | well in that case, we can just look it up and reparent the compute node object then | |
| 13:52:51 | superdan | computenode.host = self.host | |
| 13:52:53 | superdan | done | |
| 13:52:55 | johnthetubaguy | yeah, instead of add a new one | |
| 13:53:10 | johnthetubaguy | that's what I just attempted to say, badly | |
| 13:53:20 | superdan | I was going to suggest that we set it so we could do that once it was, but if it's already done we should be able to do that right away | |
| 13:54:15 | johnthetubaguy | I will take a look at doing that | |
| 13:54:29 | johnthetubaguy | in the ironic driver, or where we create the compute nodes already? | |
| 13:54:46 | superdan | um | |
| 13:54:59 | superdan | I'd have to go look at that.. I guess there's a bit of a layering violation in there some where | |
| 13:55:19 | superdan | maybe we can just make compute manager check for this situation and do the re-homing | |
| 13:56:36 | johnthetubaguy | yeah, I just worry about hostname clashes | |
| 13:56:38 | johnthetubaguy | in none ironic cases | |
| 13:56:45 | superdan | well, | |
| 13:56:54 | superdan | the hostnames have to be unique today or nothing works anyway | |
| 13:57:03 | superdan | but you could also is_uuid_like() on it for a total hack :) | |
| 13:57:22 | johnthetubaguy | good point, placement has that constraint now, hence the bug in the first place | |
| 13:57:23 | superdan | or set some capability attribute on the driver: CAN_REBALANCE_NODES | |
| 13:57:27 | superdan | LIKES_TO_SHARE=True | |
| 13:57:32 | johnthetubaguy | heh | |
| 13:57:35 | superdan | PLAYS_WELL_WITH_OTHERS=True | |
| 13:57:48 | johnthetubaguy | that doesn't sound right... | |
| 13:58:03 | superdan | we have to have unique hostnames for rpc to work anyway | |
| 13:58:30 | superdan | maybe not hypervisor_hostname I guess, but they should be the same for non-ironic ones anyway right? | |
| 13:58:50 | johnthetubaguy | isn't that one on the service, but yeah thats the same for regular virt | |
| 13:58:59 | johnthetubaguy | OK, I am convinced, will give that a whirl | |
| 13:59:36 | superdan | yes it's service.host that has to be unique but should be the same for virt systems | |
| 14:00:39 | johnthetubaguy | cool, just making sure, keep getting it all twisted | |
| 14:03:31 | cdent | register headline: OpenStack switches back to twisted | |
| 14:06:24 | superdan | mriedem: I hadn't seen your second flavor sysmeta patch.. I didn't test with it applied as well | |
| 14:06:55 | johnthetubaguy | cdent: heh | |
| 14:07:02 | mriedem | superdan: that shouldn't make a difference during instance list | |
| 14:07:15 | mriedem | since it's just used during (1) save and (2) lazy-loading of flavor i think | |
| 14:07:28 | superdan | mriedem: well, unless we started lazy-loading it, but yeah, I was checking for that in the logs anyway | |
| 14:11:44 | figleaf | cdent: heh, I remember the discussions that led to the switch from twisted to eventlet | |
| 14:13:42 | openstackgerrit | Merged openstack/nova-specs master: Update a URL https://review.openstack.org/489028 | |
| 14:25:16 | bauzas | superdan: https://review.openstack.org/#/c/498948/10/nova/compute/manager.py@3640 I'm not super expert of all the migration states but I guess a migration revert is not having an 'in-progress state' ? | |
| 14:25:45 | bauzas | I just want to make sure that we don't do the math if we're at the middle of a revert | |
| 14:26:14 | bauzas | wait | |
| 14:26:33 | bauzas | if we get allocations, we know we have nothing to do | |
| 14:26:46 | bauzas | so we shouldn't care of the migration state then ? | |
| 14:29:46 | superdan | bauzas: right that's just a short-circuit | |
| 14:48:36 | openstackgerrit | priyaduggirala proposed openstack/nova master: Rename arguments in call() of nova/image/glance.py https://review.openstack.org/508533 | |
| 14:50:10 | bauzas | superdan: corrolary question, if we are reverting and if we have allocations related to the migration, we of course don't remove them. But then, we will remove them once it's calling finish_revert_resize() right? | |
| 14:50:29 | leakypipes | cdent: I presume you found the scheduler presentation links? the links to the presentations are http://bit.ly/scheduler-wars-a-new-hope and http://bit.ly/scheduler-wars-revenge-of-the-split | |
| 14:50:39 | bauzas | superdan: nevermind, just saw https://review.openstack.org/#/c/498949/11/nova/compute/manager.py | |
| 14:50:45 | bauzas | okay, so +W | |
| 14:50:52 | cdent | leakypipes: i did not, thank you | |
| 14:51:06 | superdan | bauzas: check out the functional tests.. unless you find a gap, they're pretty obsessive about the allocations at all the steps | |
| 14:51:18 | bauzas | cool with me | |
| 14:51:43 | leakypipes | cdent: BTW, thanks again for continuing your resource provider update emails. they really are very helpful. | |
| 14:52:29 | leakypipes | cdent: also, nice use of the word "passel" | |
| 14:52:34 | cdent | leakypipes: you’re welcome. I decided this week I’d go short and focused, in part because that seemed like good timing, but also because sometimes I’m so far behind myself that I haven’t got time to do the big version. we are moving a _ton_ of code these days | |
| 14:58:15 | melwitt | superdan: I dunno if you saw this but I think I found a bug with the get_instance_object_sorted when there are faults. I explained it here https://review.openstack.org/#/c/505417/7/nova/tests/functional/compute/test_instance_list.py@453 | |
| 14:59:50 | superdan | melwitt: so we're not going to merge the patch that pre-joins faults | |
| 14:59:57 | superdan | so that test can go away | |
| 15:00:10 | superdan | is there actually a problem based on how the API works now, or just that that test failed? | |
| 15:01:12 | melwitt | superdan: it's that _from_db_object doesn't accept pre-joined attrs, i.e. it won't set them. so even after you pre-join it lazy-loads it. and with an untargeted _context, it won't be able to and will get a None fault | |
| 15:01:16 | superdan | and that test was intentionally not 3 cells (see mriedem's comment above) but only because it wasn't really needed | |
| 15:01:32 | superdan | melwitt: okay but the api isn't asking for fault to be joined | |
| 15:02:02 | superdan | melwitt: the api loads the faults it wants in a batch after having collected the full list | |
| 15:02:07 | melwitt | superdan: you mean nova-api? | |
| 15:02:15 | superdan | compute/api and above but yeah | |