Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-20
15:54:11 gibi could one of the cores look at test only patch? https://review.openstack.org/#/c/496202/ only needs a second +2
15:54:17 cdent fix on MiniDNS
15:55:58 openstackgerrit Evgeny Antyshev proposed openstack/nova master: Vzstorage: synchronize volume connect/disconnect https://review.openstack.org/505708
15:57:34 stephenfin gibi: Looking
15:59:10 gibi stephenfin: thanks
16:05:50 openstackgerrit OpenStack Proposal Bot proposed openstack/os-traits master: Updated from global requirements https://review.openstack.org/503646
16:05:54 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif master: Updated from global requirements https://review.openstack.org/502708
16:08:21 mriedem gibi: stephenfin: commented on https://review.openstack.org/#/c/496202/
16:08:29 mriedem that seems to be munging together a few different things
16:08:37 openstackgerrit Stephen Finucane proposed openstack/nova master: Add functional migrate force_complete test https://review.openstack.org/496202
16:09:26 stephenfin mriedem: Aye, but they all seemed valid
16:09:41 stephenfin There could be merit in splitting them out though, so I've rebased to take it out of the gate queue
16:09:56 mriedem if we're going to whitelist legacy notifications because they were missed, the change that adds them in should have a test to make sure they are sent
16:10:16 mriedem which i guess is the referenced patch...but those are linked together in any way
16:10:28 mriedem here is an example https://review.openstack.org/#/c/504978/
16:13:05 dansmith mriedem: right on right on right on
16:13:36 gibi mriedem: OK, let's make a separate bug for the missing force_complete test coverage and move the whitelisting into that
16:15:50 mriedem sounds good to me
16:16:01 mriedem dansmith: you're getting older and the patches are staying the same age?
16:16:02 gibi sorry for the mess
16:16:10 dansmith mriedem: lol
16:19:11 bauzas dansmith: sdague: sorry, was running a meeting, but like I said in my comment in https://review.openstack.org/#/c/419502/2 "Modifying an AZ name should not be possible if instances are still in the AZ"
16:19:23 mriedem sdague: jamespage: looks like the devstack patch with the qemu version hack is passing the neutron multinode job that runs live migration
16:19:41 bauzas dansmith: sdague: the only point I have is that if we agree on that, should we have a microversion for that?
16:20:00 dansmith bauzas: we can't support the old thing, so no, IMHO
16:20:27 bauzas dansmith: sdague: I mean, say I wanna update my aggregate and modify the AZ value, should we need a microversion for returning a 40x now ?
16:20:39 bauzas instead of a HTTP200
16:20:54 bauzas in theory, it should, but I guess it's an already broken behaviour
16:21:46 bauzas also, it would require to get the full list of all per-cell instances when updating the aggregate, not a huge deal but still something to do
16:23:05 gibi mriedem, stephenfin: reported the bug https://bugs.launchpad.net/nova/+bug/1718485
16:23:06 openstack Launchpad bug 1718485 in OpenStack Compute (nova) "instance.live.migration.force.complete is not a versioned notification and not whitelisted" [Undecided,New]
16:24:15 mriedem thanks
16:25:21 gibi will push the fix soon
16:32:27 stephenfin gibi: Cool. Ping me when you do, sure
16:34:00 openstackgerrit Matt Riedemann proposed openstack/nova master: WIP: check qemu version when calling qemu-img info https://review.openstack.org/505673
16:41:30 mriedem dansmith: i assume the -2 on https://review.openstack.org/#/c/504983/ is just a placeholder for the whole thing to be reviewed and approved?
16:42:03 dansmith mriedem: it was because the top level wasn't wired in, but yeah I figure no reason to land things that aren't reachable until we have some reasonable reviews on the rest of it
16:42:47 mriedem efried: fyi for your ksa endpoint discovery stuff https://review.openstack.org/#/c/485121/
16:42:56 mriedem dansmith: ok, will hit those after lunh
16:42:58 mriedem *lunch
16:44:01 dansmith mriedem: cool thanks
16:44:03 dansmith mriedem: note that I found a bug in a functional test for pagination with that series
16:44:22 dansmith which wasn't an issue before because we were inefficient, but it's nice that it was a test bug and not a functional one
16:45:45 mriedem efried: edmondsw: if the operator has to configure something different because of https://review.openstack.org/#/c/505546/ that's a big no-no for a backport
16:46:08 edmondsw mriedem they don't have to
16:46:15 edmondsw mriedem they can... they don't have to
16:49:36 openstackgerrit James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748
16:50:38 cdent edleafe: a) o/ b) if bp/return-selection-objects the right topic for the real thing?
16:51:23 cdent s/if/is/
16:52:10 openstackgerrit Merged openstack/nova master: [placement] Unregister the UsageList object https://review.openstack.org/502156
16:53:13 jamespage mriedem: https://review.openstack.org/505748 but I think I prefer sdague's approach - calls to qemu_img_info are in alot of places...
16:54:21 openstackgerrit Merged openstack/nova master: [placement] Unregister the Usage object https://review.openstack.org/502157
16:55:02 mriedem jamespage: yeah just reviewed it
16:55:04 mriedem left some comments
16:55:21 mriedem there are other places it's going to fail because you're not passing that flag, like fetch_to_raw
16:55:30 openstackgerrit Merged openstack/nova master: Use symbolic names for capabilities, expand sys_admin context. https://review.openstack.org/504193
16:55:31 mriedem it really becomes a lot of whack a mole
16:56:00 openstackgerrit Merged openstack/nova master: [placement] Unregister the AllocationList object https://review.openstack.org/502158
16:56:50 openstackgerrit Merged openstack/nova master: [placement] Unregister the Allocation object https://review.openstack.org/502159
16:57:24 dansmith mriedem: jamespage: Is this the fix for the live migration job?
16:58:00 jamespage mriedem: yeah - felt like pulling at a ball of string
16:58:10 jamespage dansmith: a start on at least
16:58:40 dansmith okay, I feel like we need a quicker resolution in the meantime.. are we reverting the repo for devstack in the interim or something?
16:58:47 dansmith apologies if I missed it
16:59:51 mriedem dansmith: the revert was merged last night
17:00:04 dansmith oh? I thought I saw fails from this morning
17:00:34 mriedem 3:37am i guess https://review.openstack.org/#/c/505446/
17:01:11 dansmith okay maybe these ran before that
17:09:14 openstackgerrit James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748
17:11:55 openstackgerrit Merged openstack/nova master: [placement] Unregister the InventoryList object https://review.openstack.org/502160
17:12:10 edleafe cdent: a_) \o b) don't understand the question
17:12:29 openstackgerrit Merged openstack/nova stable/pike: Add @targets_cell for live_migrate_instance method in conductor https://review.openstack.org/505285
17:12:32 cdent edleafe: I’m trying to confirm that that’s the correct for review
17:12:58 openstackgerrit Merged openstack/nova master: [placement] Unregister the Inventory object https://review.openstack.org/502161
17:12:59 edleafe cdent: yes, that's the one
17:13:04 cdent thanks
17:13:33 openstackgerrit Merged openstack/nova master: [placement] Unregister the ResourceProviderList object https://review.openstack.org/502162
17:14:04 openstackgerrit Merged openstack/nova master: [placement] Unregister the ResourceProvider object https://review.openstack.org/502163
17:16:50 openstackgerrit Merged openstack/nova master: [placement] Removing versioning from resource_provider objects https://review.openstack.org/502164
17:19:32 tasker mriedem: the only entry in the nova-conductor is a reply from the scheduler saying "no valid host was found". the scheduler shows "host [u'compute-2'] fails" but doesn't explain _why_ it failed. and the computes have no record or log of the request because it never gets past the scheduler.
17:21:23 mriedem tasker: the scheduler will dump the filter it failed on
17:21:30 mriedem i'd have to see if that's logged at info or debug
17:21:40 melwitt tasker: was the target host removed by a scheduler filter or are you saying the request was sent to the target compute host and then failed there?
17:22:30 melwitt if the former, the debug logs should show which filter removed the host from consideration. if the latter, the nova-compute logs should contain some message about why the request failed
17:22:43 mriedem tasker: you should see something like this at INFO level
17:22:44 mriedem LOG.info(_LI("Filter %s returned 0 hosts"), cls_name)
17:22:52 mriedem so for that request, figure out which filter kicked it out
17:23:52 mriedem you should also see something like, "Filtering removed all hosts for the request with"
17:24:00 mriedem if you have INFO level logging enabled
17:24:07 tasker Filter results: ['RetryFilter: (start: 1, end: 0)']
17:24:30 tasker ok .. after your explanation, that line makes sense.
17:25:14 melwitt it sounds like the request landed on a compute host and then failed and came back to the scheduler to retry a different host, and then it failed with NoValidHost
17:25:41 mriedem i don't think live migration actually does a retry from the compute
17:25:49 mriedem conductor will retry if the pre-migration checks on the chosen target host fail
17:26:05 mriedem we only retry from compute -> conductor for (1) initial create and (2) cold migrate/resize
17:26:56 mriedem this is where conductor is asking the scheduler for a host during live migrate https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L274
17:27:01 tasker other migrations succeed.
17:27:04 tasker just not this one.
17:28:12 melwitt there should be some message about the request in the compute log. if filtering didn't remove the target host, then it must have gone to the host and then failed
17:28:17 tasker on a successful migrtion, RetryFilter returns 1 host.

Earlier   Later