| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-20 | |||
| 16:42:47 | mriedem | efried: fyi for your ksa endpoint discovery stuff https://review.openstack.org/#/c/485121/ | |
| 16:42:56 | mriedem | dansmith: ok, will hit those after lunh | |
| 16:42:58 | mriedem | *lunch | |
| 16:44:01 | dansmith | mriedem: cool thanks | |
| 16:44:03 | dansmith | mriedem: note that I found a bug in a functional test for pagination with that series | |
| 16:44:22 | dansmith | which wasn't an issue before because we were inefficient, but it's nice that it was a test bug and not a functional one | |
| 16:45:45 | mriedem | efried: edmondsw: if the operator has to configure something different because of https://review.openstack.org/#/c/505546/ that's a big no-no for a backport | |
| 16:46:08 | edmondsw | mriedem they don't have to | |
| 16:46:15 | edmondsw | mriedem they can... they don't have to | |
| 16:49:36 | openstackgerrit | James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748 | |
| 16:50:38 | cdent | edleafe: a) o/ b) if bp/return-selection-objects the right topic for the real thing? | |
| 16:51:23 | cdent | s/if/is/ | |
| 16:52:10 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the UsageList object https://review.openstack.org/502156 | |
| 16:53:13 | jamespage | mriedem: https://review.openstack.org/505748 but I think I prefer sdague's approach - calls to qemu_img_info are in alot of places... | |
| 16:54:21 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Usage object https://review.openstack.org/502157 | |
| 16:55:02 | mriedem | jamespage: yeah just reviewed it | |
| 16:55:04 | mriedem | left some comments | |
| 16:55:21 | mriedem | there are other places it's going to fail because you're not passing that flag, like fetch_to_raw | |
| 16:55:30 | openstackgerrit | Merged openstack/nova master: Use symbolic names for capabilities, expand sys_admin context. https://review.openstack.org/504193 | |
| 16:55:31 | mriedem | it really becomes a lot of whack a mole | |
| 16:56:00 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the AllocationList object https://review.openstack.org/502158 | |
| 16:56:50 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Allocation object https://review.openstack.org/502159 | |
| 16:57:24 | dansmith | mriedem: jamespage: Is this the fix for the live migration job? | |
| 16:58:00 | jamespage | mriedem: yeah - felt like pulling at a ball of string | |
| 16:58:10 | jamespage | dansmith: a start on at least | |
| 16:58:40 | dansmith | okay, I feel like we need a quicker resolution in the meantime.. are we reverting the repo for devstack in the interim or something? | |
| 16:58:47 | dansmith | apologies if I missed it | |
| 16:59:51 | mriedem | dansmith: the revert was merged last night | |
| 17:00:04 | dansmith | oh? I thought I saw fails from this morning | |
| 17:00:34 | mriedem | 3:37am i guess https://review.openstack.org/#/c/505446/ | |
| 17:01:11 | dansmith | okay maybe these ran before that | |
| 17:09:14 | openstackgerrit | James Page proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505748 | |
| 17:11:55 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the InventoryList object https://review.openstack.org/502160 | |
| 17:12:10 | edleafe | cdent: a_) \o b) don't understand the question | |
| 17:12:29 | openstackgerrit | Merged openstack/nova stable/pike: Add @targets_cell for live_migrate_instance method in conductor https://review.openstack.org/505285 | |
| 17:12:32 | cdent | edleafe: I’m trying to confirm that that’s the correct for review | |
| 17:12:58 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the Inventory object https://review.openstack.org/502161 | |
| 17:12:59 | edleafe | cdent: yes, that's the one | |
| 17:13:04 | cdent | thanks | |
| 17:13:33 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the ResourceProviderList object https://review.openstack.org/502162 | |
| 17:14:04 | openstackgerrit | Merged openstack/nova master: [placement] Unregister the ResourceProvider object https://review.openstack.org/502163 | |
| 17:16:50 | openstackgerrit | Merged openstack/nova master: [placement] Removing versioning from resource_provider objects https://review.openstack.org/502164 | |
| 17:19:32 | tasker | mriedem: the only entry in the nova-conductor is a reply from the scheduler saying "no valid host was found". the scheduler shows "host [u'compute-2'] fails" but doesn't explain _why_ it failed. and the computes have no record or log of the request because it never gets past the scheduler. | |
| 17:21:23 | mriedem | tasker: the scheduler will dump the filter it failed on | |
| 17:21:30 | mriedem | i'd have to see if that's logged at info or debug | |
| 17:21:40 | melwitt | tasker: was the target host removed by a scheduler filter or are you saying the request was sent to the target compute host and then failed there? | |
| 17:22:30 | melwitt | if the former, the debug logs should show which filter removed the host from consideration. if the latter, the nova-compute logs should contain some message about why the request failed | |
| 17:22:43 | mriedem | tasker: you should see something like this at INFO level | |
| 17:22:44 | mriedem | LOG.info(_LI("Filter %s returned 0 hosts"), cls_name) | |
| 17:22:52 | mriedem | so for that request, figure out which filter kicked it out | |
| 17:23:52 | mriedem | you should also see something like, "Filtering removed all hosts for the request with" | |
| 17:24:00 | mriedem | if you have INFO level logging enabled | |
| 17:24:07 | tasker | Filter results: ['RetryFilter: (start: 1, end: 0)'] | |
| 17:24:30 | tasker | ok .. after your explanation, that line makes sense. | |
| 17:25:14 | melwitt | it sounds like the request landed on a compute host and then failed and came back to the scheduler to retry a different host, and then it failed with NoValidHost | |
| 17:25:41 | mriedem | i don't think live migration actually does a retry from the compute | |
| 17:25:49 | mriedem | conductor will retry if the pre-migration checks on the chosen target host fail | |
| 17:26:05 | mriedem | we only retry from compute -> conductor for (1) initial create and (2) cold migrate/resize | |
| 17:26:56 | mriedem | this is where conductor is asking the scheduler for a host during live migrate https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L274 | |
| 17:27:01 | tasker | other migrations succeed. | |
| 17:27:04 | tasker | just not this one. | |
| 17:28:12 | melwitt | there should be some message about the request in the compute log. if filtering didn't remove the target host, then it must have gone to the host and then failed | |
| 17:28:17 | tasker | on a successful migrtion, RetryFilter returns 1 host. | |
| 17:29:15 | mriedem | tasker: how many hosts do you have? does the instance have anything special about it, like pci requests or numa affinity? | |
| 17:29:36 | tasker | on the instance that fails to live-migrate there is no log on the target host. filtering removes _all_ targets regardless of which host I send it to. | |
| 17:29:52 | melwitt | RetryFilter removes previously tried hosts from consideration. so if it removes anything, that means it already tried to run it on the host it selected last time | |
| 17:29:54 | tasker | 3; no. it and all other instances were built to the same requirements. | |
| 17:30:33 | tasker | melwitt: does that have a memory? or is each migration request independent of the previous? | |
| 17:30:43 | melwitt | each request is independent | |
| 17:31:30 | tasker | the only thing different is that this instance was built while the cluster was at M. now my cluster is at N. | |
| 17:32:06 | melwitt | I would grep for the request-id of the failing live migration request in nova-conductor and nova-compute logs and see if there's anything about it | |
| 17:32:31 | tasker | I am live watching all the logs. | |
| 17:32:46 | tasker | conductor says "no valid host" because that's what scheduler says. | |
| 17:32:53 | mriedem | i'm wondering if the retry filter shouldn't be ignored here... | |
| 17:33:16 | tasker | if i send another instance to the same target, retryfilter passes the host. | |
| 17:33:42 | melwitt | RetryFilter removing a host implies that it tried to run the pre-migration on a host already | |
| 17:34:01 | melwitt | so something is weird here | |
| 17:34:21 | tasker | nothing logged at the compute at INFO or lower. | |
| 17:34:25 | mriedem | i'm wondering if the original instance request spec has retries set on it, and that's goofing up the retry filter during live migration | |
| 17:34:55 | tasker | something I could see in the database? | |
| 17:35:30 | mriedem | yeah, you should be able to pull the serialized json request spec out of the db | |
| 17:36:04 | mriedem | select spec from nova_api.request_specs where instance_uuid=xyz; | |
| 17:36:34 | mriedem | melwitt: this is where we get the request spec during live migration https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L227 | |
| 17:36:40 | mriedem | we don't set a retry attribute on the request spec | |
| 17:37:16 | mriedem | we pass it along from the api if we are able to look it up https://github.com/openstack/nova/blob/master/nova/compute/api.py#L3860 | |
| 17:37:36 | mriedem | which would be the original request spec from when we built the instance, which would have retry stuff set on it | |
| 17:38:03 | mriedem | if that wasn't set, we'd ignore it here https://github.com/openstack/nova/blob/master/nova/scheduler/filters/retry_filter.py#L31 | |
| 17:38:42 | mriedem | which, this is all kinds of weird because the conductor task is handling it's own retry logic https://github.com/openstack/nova/blob/master/nova/conductor/tasks/live_migrate.py#L344 | |
| 17:38:53 | mriedem | which tells me, when doing a live migration, we don't want to even ask the RetryFilter | |
| 17:39:01 | mriedem | bauzas: are you around? | |
| 17:39:44 | mriedem | tasker: if you can get that request spec json blob out of the db, that should tell us if the retry attribute is set and what hosts and how many times it's been retried during the initial build | |
| 17:41:43 | tasker | mriedem: yes, "retry" is set and is populated with both hypervisors that I'm trying to send it to. | |
| 17:42:16 | tasker | would you like to see it? | |
| 17:42:43 | mriedem | sure, throw it in a paste | |
| 17:46:32 | mriedem | i think this would actually fix the problem http://paste.openstack.org/show/621556/ | |
| 17:46:40 | melwitt | mriedem: so the first time you build an instance, if it fails host X, it will save that to the RequestSpec so host X will never be tried again ever? I didn't know it persisted something like that | |
| 17:46:47 | mriedem | oops one issue there | |
| 17:47:00 | tasker | http://paste.openstack.org/show/621557/ | |
| 17:47:09 | mriedem | melwitt: yeah | |
| 17:48:54 | mriedem | melwitt: i think this https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L583 | |