Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-26
19:42:40 dansmith the test _is_ asserting that the dest host's allocation was cleaned up by drop_move_claim
19:42:43 mriedem we don't ever want to remove allocations for a *rebuild*
19:42:55 mriedem the tests are specifically for evacuate
19:43:01 mriedem where the scheduler creates allocations on the dest host
19:43:13 mriedem we fail the evacuate on the dest host, so we need to remove those allocatoins created by the scheduler
19:44:04 dansmith there's a specific reason why I made this change,
19:44:11 mriedem i'm sorry for being grouchy about this,
19:44:11 dansmith and I talked it through with jaypipes which is why I made this
19:44:23 mriedem but i've spent the better part of the last 6 weeks fixing these allocation bugs,
19:44:26 dansmith so I'll have to go re-load all my context on this before I can really think about it
19:44:29 mriedem so being told i don't understand the test pisses me off
19:44:36 jaypipes mriedem: understood. and you're saying that you want the drop_move_claim() to remove those resources when ComputeResourcesUnvailable is raised but you want update_available_resource() to delete the allocations when a virt driver exception is raised?
19:44:47 dansmith jaypipes: no
19:45:48 dansmith mriedem: without offending your test sensibilities, you see this right? https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2812
19:45:59 dansmith that's what *should* have been deleting the target's allocation
19:46:05 dansmith only in the unavailable case
19:46:16 dansmith but in reality, we were always doing it for the other cases as well
19:47:14 mriedem how?
19:47:30 dansmith how? because we always ran drop_move_claim
19:48:16 dansmith reset the test to where it was and run it with the oddball exception and we'll assert that the dest host claim is zero, but it no longer is after this change
19:48:40 dansmith this: https://pastebin.com/fMUmmgMC
19:49:12 mriedem i don't know if we're talking about the same thing,
19:49:23 mriedem i wrote https://review.openstack.org/#/c/499877/ to show that we didn't need to manually remove allocations on the dest when the driver failed
19:49:27 mriedem because of drop_move_claim
19:49:42 mriedem the other test was to show that we needed to add https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2812
19:49:45 mriedem when the claim itself fails
19:49:51 dansmith right, but that doesn't make sense right?
19:49:51 mriedem raising ComputeResourcesUnavailable
19:50:08 dansmith if we fail for some driver reason, we're now "on" that dest host and should be able to run a same-host rebuild on it
19:50:11 dansmith which won't re-claim for us
19:50:26 mriedem no, we're not on that host
19:50:32 mriedem if driver.spawn fails, we're not on that host
19:50:43 mriedem the instance is only on the dest host if we get here https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2836
19:50:48 mriedem which doesn't happen if the driver fails
19:50:58 dansmith why is this here? https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2812
19:51:15 mriedem see https://review.openstack.org/#/c/499878/
19:51:40 mriedem plus the comment above it
19:54:32 mriedem so i don't know what else is going on in this patch, i didn't get that far, i saw the commit message and change to the test and wanted to bring that up since it's approved
19:54:57 dansmith mriedem: okay yeah I really thought that the rt claim would set host and node
19:55:07 dansmith mriedem: you better -2 that or something so it doesn't merge
19:55:18 mriedem i don't think that works
19:55:31 dansmith mriedem: -2 will make it not merge I think, it just won't kick it
19:55:39 openstackgerrit Matt Riedemann proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
19:55:45 openstackgerrit Dan Smith proposed openstack/nova master: BUMPMove allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
19:55:47 mriedem i know a new commit will do it
19:55:49 mriedem heh
19:55:57 mikal I have a cold, hold me
19:56:05 mriedem gdi mikal
19:56:11 mriedem you've stepped into the wrong room at the wrong time
19:56:17 mriedem way out of line donny
19:56:36 jaypipes you're outta your element, mikal
19:56:53 mikal jaypipes: that's always been true though
19:56:57 jaypipes :P
19:57:03 mriedem mikal, btw, given your plentiful rackspaceness, do you know if rax ever used this thing http://lists.openstack.org/pipermail/openstack-operators/2017-September/014267.html ?
19:57:05 mikal mriedem: my name isn't donny?
19:57:10 mriedem gdi mikal
19:57:23 mriedem https://www.youtube.com/watch?v=AS8X2Qp_6aA
19:57:48 mriedem i'm walter in this scenario
19:57:50 mriedem filled with rage
19:57:58 mikal mriedem: so, Rackspace is definitely running code newer than kilo, so if they haven't noticed they don't need it?
19:58:22 mikal mriedem: private cloud didn't use it, public might but johnthetubaguy would know more about that
19:58:36 mikal mriedem: ahhh, ok, I shall study before next time
19:58:50 mriedem mikal: anyone that was relying on it and has newer than kilo, and didn't notice, then yeah i guess we don't need it
19:59:04 mriedem that's the assertion in the ML thread and commit to remove it anyway
20:03:35 melwitt does anything ever set 'group_members' in filter_properties? I'm not finding anything https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L205
20:04:42 sdague mikal: I retool your ploop patch with the fixes from the virtuozo folks
20:04:59 melwitt this line is making reschedule fail with "'NoneType' object is not iterable" after a late-affinity-check failure. I don't see how reschedule after late check ever could work
20:08:23 mriedem melwitt: it could be a case of group_members being set on the request spec initially, and then it's transformed into the primitive filter_properties stuff which doesn't include the group_members?
20:08:32 mriedem i've seen some wonky stuff with how the request spec transforms to/from the legacy filter props
20:09:19 mikal sdague: ta, looking at it now
20:09:37 mriedem melwitt: see _to_legacy_group_info ?
20:09:49 mriedem it sets group_updated=True but doesn't include group_members
20:09:51 melwitt mriedem: I don't see that RequestSpec has any group_members in it. I grepped for "group_members" in nova and found nothing that ever sets it
20:10:19 melwitt yeah, I see that. that's the only thing that looks like it could be related
20:10:21 mikal sdague: looks like there is a rebase error there though? The console pty stuff is now in that patch.
20:10:43 mtreinish dansmith: http://logs.openstack.org/57/507657/1/check/gate-nova-python27-ubuntu-xenial/55aebe1/console.html#_2017-09-26_19_42_08_172719
20:10:50 melwitt mriedem: late affinity check reschedule just looks totally broken unless I'm blind. which may very well be the case
20:10:50 mikal sdague: I shall rectify
20:11:00 mriedem melwitt: this added it https://review.openstack.org/#/c/148277/
20:11:43 melwitt thanks
20:12:08 mriedem https://review.openstack.org/#/c/148277/64/nova/scheduler/utils.py@349
20:12:14 melwitt looks like that line in scheduler/utils that sets it is gone now
20:12:36 openstackgerrit Ed Leafe proposed openstack/nova master: Add Selection objects https://review.openstack.org/499239
20:12:36 openstackgerrit Ed Leafe proposed openstack/nova master: Add alternate hosts https://review.openstack.org/486215
20:12:37 openstackgerrit Ed Leafe proposed openstack/nova master: Return Selection objects from the scheduler driver https://review.openstack.org/495854
20:13:06 melwitt trying to find what removed it
20:13:08 mriedem melwitt: https://review.openstack.org/#/c/469037/
20:13:22 mriedem pike ^
20:13:46 mriedem https://review.openstack.org/#/c/469037/6/nova/scheduler/utils.py
20:13:48 melwitt thanks
20:14:50 melwitt okay, so the conductor logic is still relying on stuff being in filter_properties via RequestSpec.from_primitives
20:15:39 mriedem looks like it, left some comments in https://review.openstack.org/#/c/469037/6/nova/objects/request_spec.py
20:15:46 mriedem likely need a functional regression test to show the failure
20:15:54 mriedem then fix on top and backport both to pike
20:16:02 mriedem melwitt: do you have a bug for this?
20:17:42 melwitt mriedem: no, I can open one. wanted to sanity check with yall first
20:17:56 mriedem someone was saying they hit this exact same thing this morning to bauzas
20:18:45 melwitt hah, what a coinky dink
20:19:09 mriedem please keep that sailor talk for at home
20:20:17 melwitt aye aye sir!

Earlier   Later