Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-20
16:20:27 bauzas dansmith: sdague: I mean, say I wanna update my aggregate and modify the AZ value, should we need a microversion for returning a 40x now ?
16:20:39 bauzas instead of a HTTP200
16:20:54 bauzas in theory, it should, but I guess it's an already broken behaviour
16:21:46 bauzas also, it would require to get the full list of all per-cell instances when updating the aggregate, not a huge deal but still something to do
16:23:05 gibi mriedem, stephenfin: reported the bug https://bugs.launchpad.net/nova/+bug/1718485
16:23:06 openstack Launchpad bug 1718485 in OpenStack Compute (nova) "instance.live.migration.force.complete is not a versioned notification and not whitelisted" [Undecided,New]
16:24:15 mriedem thanks
16:25:21 gibi will push the fix soon
16:32:27 stephenfin gibi: Cool. Ping me when you do, sure
16:34:00 openstackgerrit Matt Riedemann proposed openstack/nova master: WIP: check qemu version when calling qemu-img info https://review.openstack.org/505673
16:41:30 mriedem dansmith: i assume the -2 on https://review.openstack.org/#/c/504983/ is just a placeholder for the whole thing to be reviewed and approved?
16:42:03 dansmith mriedem: it was because the top level wasn't wired in, but yeah I figure no reason to land things that aren't reachable until we have some reasonable reviews on the rest of it
16:42:47 mriedem efried: fyi for your ksa endpoint discovery stuff https://review.openstack.org/#/c/485121/
16:42:56 mriedem dansmith: ok, will hit those after lunh
16:42:58 mriedem *lunch
16:44:01 dansmith mriedem: cool thanks
16:44:03 dansmith mriedem: note that I found a bug in a functional test for pagination with that series
16:44:22 dansmith which wasn't an issue before because we were inefficient, but it's nice that it was a test bug and not a functional one
16:45:45 mriedem efried: edmondsw: if the operator has to configure something different because of https://review.openstack.org/#/c/505546/ that's a big no-no for a backport
16:46:08 edmondsw mriedem they don't have to
16:46:15 edmondsw mriedem they can... they don't have to
16:49:36 openstackgerrit James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748
16:50:38 cdent edleafe: a) o/ b) if bp/return-selection-objects the right topic for the real thing?
16:51:23 cdent s/if/is/
16:52:10 openstackgerrit Merged openstack/nova master: [placement] Unregister the UsageList object https://review.openstack.org/502156
16:53:13 jamespage mriedem: https://review.openstack.org/505748 but I think I prefer sdague's approach - calls to qemu_img_info are in alot of places...
16:54:21 openstackgerrit Merged openstack/nova master: [placement] Unregister the Usage object https://review.openstack.org/502157
16:55:02 mriedem jamespage: yeah just reviewed it
16:55:04 mriedem left some comments
16:55:21 mriedem there are other places it's going to fail because you're not passing that flag, like fetch_to_raw
16:55:30 openstackgerrit Merged openstack/nova master: Use symbolic names for capabilities, expand sys_admin context. https://review.openstack.org/504193
16:55:31 mriedem it really becomes a lot of whack a mole
16:56:00 openstackgerrit Merged openstack/nova master: [placement] Unregister the AllocationList object https://review.openstack.org/502158
16:56:50 openstackgerrit Merged openstack/nova master: [placement] Unregister the Allocation object https://review.openstack.org/502159
16:57:24 dansmith mriedem: jamespage: Is this the fix for the live migration job?
16:58:00 jamespage mriedem: yeah - felt like pulling at a ball of string
16:58:10 jamespage dansmith: a start on at least
16:58:40 dansmith okay, I feel like we need a quicker resolution in the meantime.. are we reverting the repo for devstack in the interim or something?
16:58:47 dansmith apologies if I missed it
16:59:51 mriedem dansmith: the revert was merged last night
17:00:04 dansmith oh? I thought I saw fails from this morning
17:00:34 mriedem 3:37am i guess https://review.openstack.org/#/c/505446/
17:01:11 dansmith okay maybe these ran before that
17:09:14 openstackgerrit James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748
17:11:55 openstackgerrit Merged openstack/nova master: [placement] Unregister the InventoryList object https://review.openstack.org/502160
17:12:10 edleafe cdent: a_) \o b) don't understand the question
17:12:29 openstackgerrit Merged openstack/nova stable/pike: Add @targets_cell for live_migrate_instance method in conductor https://review.openstack.org/505285
17:12:32 cdent edleafe: I’m trying to confirm that that’s the correct for review
17:12:58 openstackgerrit Merged openstack/nova master: [placement] Unregister the Inventory object https://review.openstack.org/502161
17:12:59 edleafe cdent: yes, that's the one
17:13:04 cdent thanks
17:13:33 openstackgerrit Merged openstack/nova master: [placement] Unregister the ResourceProviderList object https://review.openstack.org/502162
17:14:04 openstackgerrit Merged openstack/nova master: [placement] Unregister the ResourceProvider object https://review.openstack.org/502163
17:16:50 openstackgerrit Merged openstack/nova master: [placement] Removing versioning from resource_provider objects https://review.openstack.org/502164
17:19:32 tasker mriedem: the only entry in the nova-conductor is a reply from the scheduler saying "no valid host was found". the scheduler shows "host [u'compute-2'] fails" but doesn't explain _why_ it failed. and the computes have no record or log of the request because it never gets past the scheduler.
17:21:23 mriedem tasker: the scheduler will dump the filter it failed on
17:21:30 mriedem i'd have to see if that's logged at info or debug
17:21:40 melwitt tasker: was the target host removed by a scheduler filter or are you saying the request was sent to the target compute host and then failed there?
17:22:30 melwitt if the former, the debug logs should show which filter removed the host from consideration. if the latter, the nova-compute logs should contain some message about why the request failed
17:22:43 mriedem tasker: you should see something like this at INFO level
17:22:44 mriedem LOG.info(_LI("Filter %s returned 0 hosts"), cls_name)
17:22:52 mriedem so for that request, figure out which filter kicked it out
17:23:52 mriedem you should also see something like, "Filtering removed all hosts for the request with"
17:24:00 mriedem if you have INFO level logging enabled
17:24:07 tasker Filter results: ['RetryFilter: (start: 1, end: 0)']
17:24:30 tasker ok .. after your explanation, that line makes sense.
17:25:14 melwitt it sounds like the request landed on a compute host and then failed and came back to the scheduler to retry a different host, and then it failed with NoValidHost
17:25:41 mriedem i don't think live migration actually does a retry from the compute
17:25:49 mriedem conductor will retry if the pre-migration checks on the chosen target host fail
17:26:05 mriedem we only retry from compute -> conductor for (1) initial create and (2) cold migrate/resize
17:26:56 mriedem this is where conductor is asking the scheduler for a host during live migrate https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L274
17:27:01 tasker other migrations succeed.
17:27:04 tasker just not this one.
17:28:12 melwitt there should be some message about the request in the compute log. if filtering didn't remove the target host, then it must have gone to the host and then failed
17:28:17 tasker on a successful migrtion, RetryFilter returns 1 host.
17:29:15 mriedem tasker: how many hosts do you have? does the instance have anything special about it, like pci requests or numa affinity?
17:29:36 tasker on the instance that fails to live-migrate there is no log on the target host. filtering removes _all_ targets regardless of which host I send it to.
17:29:52 melwitt RetryFilter removes previously tried hosts from consideration. so if it removes anything, that means it already tried to run it on the host it selected last time
17:29:54 tasker 3; no. it and all other instances were built to the same requirements.
17:30:33 tasker melwitt: does that have a memory? or is each migration request independent of the previous?
17:30:43 melwitt each request is independent
17:31:30 tasker the only thing different is that this instance was built while the cluster was at M. now my cluster is at N.
17:32:06 melwitt I would grep for the request-id of the failing live migration request in nova-conductor and nova-compute logs and see if there's anything about it
17:32:31 tasker I am live watching all the logs.
17:32:46 tasker conductor says "no valid host" because that's what scheduler says.
17:32:53 mriedem i'm wondering if the retry filter shouldn't be ignored here...
17:33:16 tasker if i send another instance to the same target, retryfilter passes the host.
17:33:42 melwitt RetryFilter removing a host implies that it tried to run the pre-migration on a host already
17:34:01 melwitt so something is weird here
17:34:21 tasker nothing logged at the compute at INFO or lower.
17:34:25 mriedem i'm wondering if the original instance request spec has retries set on it, and that's goofing up the retry filter during live migration
17:34:55 tasker something I could see in the database?
17:35:30 mriedem yeah, you should be able to pull the serialized json request spec out of the db
17:36:04 mriedem select spec from nova_api.request_specs where instance_uuid=xyz;
17:36:34 mriedem melwitt: this is where we get the request spec during live migration https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L227
17:36:40 mriedem we don't set a retry attribute on the request spec
17:37:16 mriedem we pass it along from the api if we are able to look it up https://github.com/openstack/nova/blob/master/nova/compute/api.py#L3860
17:37:36 mriedem which would be the original request spec from when we built the instance, which would have retry stuff set on it
17:38:03 mriedem if that wasn't set, we'd ignore it here https://github.com/openstack/nova/blob/master/nova/scheduler/filters/retry_filter.py#L31
17:38:42 mriedem which, this is all kinds of weird because the conductor task is handling it's own retry logic https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L344

Earlier   Later