| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-20 | |||
| 16:05:50 | openstackgerrit | OpenStack Proposal Bot proposed openstack/os-traits master: Updated from global requirements https://review.openstack.org/503646 | |
| 16:05:54 | openstackgerrit | OpenStack Proposal Bot proposed openstack/os-vif master: Updated from global requirements https://review.openstack.org/502708 | |
| 16:08:21 | mriedem | gibi: stephenfin: commented on https://review.openstack.org/#/c/496202/ | |
| 16:08:29 | mriedem | that seems to be munging together a few different things | |
| 16:08:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Add functional migrate force_complete test https://review.openstack.org/496202 | |
| 16:09:26 | stephenfin | mriedem: Aye, but they all seemed valid | |
| 16:09:41 | stephenfin | There could be merit in splitting them out though, so I've rebased to take it out of the gate queue | |
| 16:09:56 | mriedem | if we're going to whitelist legacy notifications because they were missed, the change that adds them in should have a test to make sure they are sent | |
| 16:10:16 | mriedem | which i guess is the referenced patch...but those are linked together in any way | |
| 16:10:28 | mriedem | here is an example https://review.openstack.org/#/c/504978/ | |
| 16:13:05 | dansmith | mriedem: right on right on right on | |
| 16:13:36 | gibi | mriedem: OK, let's make a separate bug for the missing force_complete test coverage and move the whitelisting into that | |
| 16:15:50 | mriedem | sounds good to me | |
| 16:16:01 | mriedem | dansmith: you're getting older and the patches are staying the same age? | |
| 16:16:02 | gibi | sorry for the mess | |
| 16:16:10 | dansmith | mriedem: lol | |
| 16:19:11 | bauzas | dansmith: sdague: sorry, was running a meeting, but like I said in my comment in https://review.openstack.org/#/c/419502/2 "Modifying an AZ name should not be possible if instances are still in the AZ" | |
| 16:19:23 | mriedem | sdague: jamespage: looks like the devstack patch with the qemu version hack is passing the neutron multinode job that runs live migration | |
| 16:19:41 | bauzas | dansmith: sdague: the only point I have is that if we agree on that, should we have a microversion for that? | |
| 16:20:00 | dansmith | bauzas: we can't support the old thing, so no, IMHO | |
| 16:20:27 | bauzas | dansmith: sdague: I mean, say I wanna update my aggregate and modify the AZ value, should we need a microversion for returning a 40x now ? | |
| 16:20:39 | bauzas | instead of a HTTP200 | |
| 16:20:54 | bauzas | in theory, it should, but I guess it's an already broken behaviour | |
| 16:21:46 | bauzas | also, it would require to get the full list of all per-cell instances when updating the aggregate, not a huge deal but still something to do | |
| 16:23:05 | gibi | mriedem, stephenfin: reported the bug https://bugs.launchpad.net/nova/+bug/1718485 | |
| 16:23:06 | openstack | Launchpad bug 1718485 in OpenStack Compute (nova) "instance.live.migration.force.complete is not a versioned notification and not whitelisted" [Undecided,New] | |
| 16:24:15 | mriedem | thanks | |
| 16:25:21 | gibi | will push the fix soon | |
| 16:32:27 | stephenfin | gibi: Cool. Ping me when you do, sure | |
| 16:34:00 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: check qemu version when calling qemu-img info https://review.openstack.org/505673 | |
| 16:41:30 | mriedem | dansmith: i assume the -2 on https://review.openstack.org/#/c/504983/ is just a placeholder for the whole thing to be reviewed and approved? | |
| 16:42:03 | dansmith | mriedem: it was because the top level wasn't wired in, but yeah I figure no reason to land things that aren't reachable until we have some reasonable reviews on the rest of it | |
| 16:42:47 | mriedem | efried: fyi for your ksa endpoint discovery stuff https://review.openstack.org/#/c/485121/ | |
| 16:42:56 | mriedem | dansmith: ok, will hit those after lunh | |
| 16:42:58 | mriedem | *lunch | |
| 16:44:01 | dansmith | mriedem: cool thanks | |
| 16:44:03 | dansmith | mriedem: note that I found a bug in a functional test for pagination with that series | |
| 16:44:22 | dansmith | which wasn't an issue before because we were inefficient, but it's nice that it was a test bug and not a functional one | |
| 16:45:45 | mriedem | efried: edmondsw: if the operator has to configure something different because of https://review.openstack.org/#/c/505546/ that's a big no-no for a backport | |
| 16:46:08 | edmondsw | mriedem they don't have to | |
| 16:46:15 | edmondsw | mriedem they can... they don't have to | |
| 16:49:36 | openstackgerrit | James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748 | |
| 16:50:38 | cdent | edleafe: a) o/ b) if bp/return-selection-objects the right topic for the real thing? | |
| 16:51:23 | cdent | s/if/is/ | |
| 16:52:10 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the UsageList object https://review.openstack.org/502156 | |
| 16:53:13 | jamespage | mriedem: https://review.openstack.org/505748 but I think I prefer sdague's approach - calls to qemu_img_info are in alot of places... | |
| 16:54:21 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Usage object https://review.openstack.org/502157 | |
| 16:55:02 | mriedem | jamespage: yeah just reviewed it | |
| 16:55:04 | mriedem | left some comments | |
| 16:55:21 | mriedem | there are other places it's going to fail because you're not passing that flag, like fetch_to_raw | |
| 16:55:30 | openstackgerrit | Merged openstack/nova master: Use symbolic names for capabilities, expand sys_admin context. https://review.openstack.org/504193 | |
| 16:55:31 | mriedem | it really becomes a lot of whack a mole | |
| 16:56:00 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the AllocationList object https://review.openstack.org/502158 | |
| 16:56:50 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Allocation object https://review.openstack.org/502159 | |
| 16:57:24 | dansmith | mriedem: jamespage: Is this the fix for the live migration job? | |
| 16:58:00 | jamespage | mriedem: yeah - felt like pulling at a ball of string | |
| 16:58:10 | jamespage | dansmith: a start on at least | |
| 16:58:40 | dansmith | okay, I feel like we need a quicker resolution in the meantime.. are we reverting the repo for devstack in the interim or something? | |
| 16:58:47 | dansmith | apologies if I missed it | |
| 16:59:51 | mriedem | dansmith: the revert was merged last night | |
| 17:00:04 | dansmith | oh? I thought I saw fails from this morning | |
| 17:00:34 | mriedem | 3:37am i guess https://review.openstack.org/#/c/505446/ | |
| 17:01:11 | dansmith | okay maybe these ran before that | |
| 17:09:14 | openstackgerrit | James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748 | |
| 17:11:55 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the InventoryList object https://review.openstack.org/502160 | |
| 17:12:10 | edleafe | cdent: a_) \o b) don't understand the question | |
| 17:12:29 | openstackgerrit | Merged openstack/nova stable/pike: Add @targets_cell for live_migrate_instance method in conductor https://review.openstack.org/505285 | |
| 17:12:32 | cdent | edleafe: I’m trying to confirm that that’s the correct for review | |
| 17:12:58 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Inventory object https://review.openstack.org/502161 | |
| 17:12:59 | edleafe | cdent: yes, that's the one | |
| 17:13:04 | cdent | thanks | |
| 17:13:33 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the ResourceProviderList object https://review.openstack.org/502162 | |
| 17:14:04 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the ResourceProvider object https://review.openstack.org/502163 | |
| 17:16:50 | openstackgerrit | Merged openstack/nova master: [placement] Removing versioning from resource_provider objects https://review.openstack.org/502164 | |
| 17:19:32 | tasker | mriedem: the only entry in the nova-conductor is a reply from the scheduler saying "no valid host was found". the scheduler shows "host [u'compute-2'] fails" but doesn't explain _why_ it failed. and the computes have no record or log of the request because it never gets past the scheduler. | |
| 17:21:23 | mriedem | tasker: the scheduler will dump the filter it failed on | |
| 17:21:30 | mriedem | i'd have to see if that's logged at info or debug | |
| 17:21:40 | melwitt | tasker: was the target host removed by a scheduler filter or are you saying the request was sent to the target compute host and then failed there? | |
| 17:22:30 | melwitt | if the former, the debug logs should show which filter removed the host from consideration. if the latter, the nova-compute logs should contain some message about why the request failed | |
| 17:22:43 | mriedem | tasker: you should see something like this at INFO level | |
| 17:22:44 | mriedem | LOG.info(_LI("Filter %s returned 0 hosts"), cls_name) | |
| 17:22:52 | mriedem | so for that request, figure out which filter kicked it out | |
| 17:23:52 | mriedem | you should also see something like, "Filtering removed all hosts for the request with" | |
| 17:24:00 | mriedem | if you have INFO level logging enabled | |
| 17:24:07 | tasker | Filter results: ['RetryFilter: (start: 1, end: 0)'] | |
| 17:24:30 | tasker | ok .. after your explanation, that line makes sense. | |
| 17:25:14 | melwitt | it sounds like the request landed on a compute host and then failed and came back to the scheduler to retry a different host, and then it failed with NoValidHost | |
| 17:25:41 | mriedem | i don't think live migration actually does a retry from the compute | |
| 17:25:49 | mriedem | conductor will retry if the pre-migration checks on the chosen target host fail | |
| 17:26:05 | mriedem | we only retry from compute -> conductor for (1) initial create and (2) cold migrate/resize | |
| 17:26:56 | mriedem | this is where conductor is asking the scheduler for a host during live migrate https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L274 | |
| 17:27:01 | tasker | other migrations succeed. | |
| 17:27:04 | tasker | just not this one. | |
| 17:28:12 | melwitt | there should be some message about the request in the compute log. if filtering didn't remove the target host, then it must have gone to the host and then failed | |
| 17:28:17 | tasker | on a successful migrtion, RetryFilter returns 1 host. | |
| 17:29:15 | mriedem | tasker: how many hosts do you have? does the instance have anything special about it, like pci requests or numa affinity? | |
| 17:29:36 | tasker | on the instance that fails to live-migrate there is no log on the target host. filtering removes _all_ targets regardless of which host I send it to. | |
| 17:29:52 | melwitt | RetryFilter removes previously tried hosts from consideration. so if it removes anything, that means it already tried to run it on the host it selected last time | |
| 17:29:54 | tasker | 3; no. it and all other instances were built to the same requirements. | |
| 17:30:33 | tasker | melwitt: does that have a memory? or is each migration request independent of the previous? | |