| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-11-09 | |||
| 17:20:46 | mriedem | sean-k-mooney: i suggested a config option as an option because this sounds very hit or miss | |
| 17:21:31 | sean-k-mooney | well the backing store for libvirt instance is only going to be on one mountpoint | |
| 17:21:33 | mriedem | this is definitely not something i've got a lot of experience in though | |
| 17:22:18 | sean-k-mooney | presumable it will all have the same alignment/sector size so we could jsut have a single valus and defualt it to 512 and they could set it to 4k or 8k if they have something else | |
| 17:23:02 | sean-k-mooney | the other option is just super over align to like a 64K bondary | |
| 17:23:28 | sean-k-mooney | that said im sure someone will have a 128K lun now that i have said that | |
| 17:24:25 | cfriesen | so guaranteed setting it to 4K will work for both 4K and 512b disks, so I think 4K should be the default | |
| 17:24:58 | sean-k-mooney | cfriesen: its also the most common sector size on most new disks so ya that should work | |
| 17:25:10 | dansmith | for years now | |
| 17:25:25 | cfriesen | I'd be okay with a config option if someone has weird hardware | |
| 17:26:20 | sean-k-mooney | cfriesen: so your going to submit a patch :)_ | |
| 17:26:47 | cfriesen | there's already a patch in progress | |
| 17:26:53 | cfriesen | by someone else | |
| 17:27:09 | mriedem | wee https://bugs.launchpad.net/nova/+bug/1798688 | |
| 17:27:10 | openstack | Launchpad bug 1798688 in OpenStack Compute (nova) "AllocationUpdateFailed_Remote: Failed to update allocations for consumer. Error: another process changed the consumer after the report client read the consumer state during the claim" [Undecided,Triaged] | |
| 17:27:21 | mriedem | looks like our scheduler allocation claim races have shot up since nov 4 | |
| 17:28:31 | cfriesen | mriedem: did you want to get the person to update the 4K patch to add a config option? or change the hardcoded number to 4K (which is still an improvement) and add the config option later if someone complains? | |
| 17:29:11 | dansmith | mriedem: what does that mean exactly? just the scheduler having to retry the allocation part? | |
| 17:29:29 | mriedem | dansmith: we already retry on PUT allocations in the scheduler | |
| 17:29:39 | mriedem | maybe the error message changed and we're not retrying properly now? | |
| 17:29:41 | mriedem | i haven't dug in yet | |
| 17:29:53 | mriedem | cfriesen: the patch already just changes 512 to 4k right? | |
| 17:30:06 | mriedem | cfriesen: i believe i just said we might want a 'fixes' release note for it | |
| 17:30:07 | mriedem | as a heads up | |
| 17:30:09 | melwitt | cfriesen: this is the method we have to checking for directio, if that's the same thing you mentioned earlier https://github.com/openstack/nova/blob/master/nova/privsep/utils.py#L34 | |
| 17:30:10 | dansmith | mriedem: I know we do, I'm wondering if you mean it's just having to retry more lately or if it's failing | |
| 17:30:18 | mriedem | haven't dug into the logs yet | |
| 17:30:23 | mriedem | hopefully we log if we are retrying | |
| 17:30:45 | sean-k-mooney | mriedem: ya i was under teh impression we had planned at least to retry in this case which is why we have the generation on the resouce providers in the first place | |
| 17:31:01 | dansmith | sean-k-mooney: we do retry | |
| 17:31:06 | mriedem | the scheduler doesn't do anything with generations for this as far as i know | |
| 17:31:38 | mriedem | just duplicated another bug in triage to this if someone is looking for work https://bugs.launchpad.net/nova/+bug/1783338 | |
| 17:31:38 | openstack | Launchpad bug 1783338 in OpenStack Compute (nova) "Unexpected exception in API method: ValueError: year is out of range" [Medium,Confirmed] - Assigned to Ghanshyam Mann (ghanshyammann) | |
| 17:31:48 | mriedem | something in the simple tenant usage code | |
| 17:32:01 | cfriesen | mriedem: in irc you were talking about a config option I thought. but yeah, I'd be cool with just a release note for now. | |
| 17:32:48 | mriedem | gmann: looks like https://bugs.launchpad.net/nova/+bug/1783338 was due to bad data in the db? you're probably traveling, but if you don't plan on handling this we should unassign you https://bugs.launchpad.net/nova/+bug/1783338 | |
| 17:32:48 | openstack | Launchpad bug 1783338 in OpenStack Compute (nova) "Unexpected exception in API method: ValueError: year is out of range" [Medium,Confirmed] - Assigned to Ghanshyam Mann (ghanshyammann) | |
| 17:32:49 | cfriesen | melwitt: yes, that's the one. there's a patch in review to change the 512 to 4096 in there. which is good, but maybe not sufficient for exotic hardware | |
| 17:33:02 | mriedem | i'm mriedem | |
| 17:33:26 | melwitt | I said something to him earlier | |
| 17:34:06 | melwitt | cfriesen: ok, cool. *looks for the patch* | |
| 17:34:07 | mriedem | oh missed that | |
| 17:34:32 | cfriesen | melwitt: https://review.openstack.org/#/c/616580 | |
| 17:34:45 | melwitt | probably because our nicks blend together. maybe I need to be jgwentworth all the time | |
| 17:35:44 | mriedem | dansmith: looks like, from the logs, that we're not retrying | |
| 17:36:30 | dansmith | mriedem: maybe something changed recently then? | |
| 17:39:26 | melwitt | cfriesen: ok, so looks like trying to decide which value to use for the check | |
| 17:40:36 | mriedem | dansmith: my guess would be https://review.openstack.org/#/c/583667/ | |
| 17:40:54 | mriedem | because the scheduler logs are saying we're doing a double up allocation | |
| 17:40:54 | mriedem | Nov 06 19:48:36.969356 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: DEBUG nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] Doubling-up allocation_request for move operation. {{(pid=13677) _move_operation_alloc_request /opt/stack/nova/nova/scheduler/client/report.py:203}} | |
| 17:41:09 | mriedem | but in this test, we're just unshelving a shelved offloaded server | |
| 17:41:15 | mriedem | so that shouldn't really double up any allocatoins | |
| 17:42:08 | cfriesen | melwitt: basically, yes. 4096 would work for the vast majority of systems | |
| 17:42:49 | mriedem | Nov 06 19:48:36.969659 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: DEBUG nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] New allocation_request containing both source and destination hosts in move operation: {'allocations': {u'3ceb7eab-549c-40ba-a70c-320822c310ab': {'resources': {u'VCPU': 2, u'MEMORY | |
| 17:42:49 | mriedem | : 128}}}} {{(pid=13677) _move_operation_alloc_request /opt/stack/nova/nova/scheduler/client/report.py:234}} | |
| 17:42:50 | dansmith | mriedem: it also touches the code near where we raise retry... | |
| 17:42:52 | mriedem | ^ is definitely wrong | |
| 17:43:00 | mriedem | there is only one provider in that log | |
| 17:43:32 | dansmith | mriedem: it seems to specifically exclude the consumer generation conflict from the case where we retry | |
| 17:46:23 | dansmith | mriedem: do you see "another process changed the consumer" in the log? | |
| 17:46:43 | mriedem | yes | |
| 17:46:53 | dansmith | then it's hitting that consumer case and not retrying | |
| 17:46:55 | mriedem | that's why we don't retry | |
| 17:46:58 | dansmith | yeah | |
| 17:47:09 | mriedem | i also don't know why it thinks we're starting with existing allocations for a shelved offloaded server | |
| 17:47:44 | dansmith | and that's causing it to try to double? | |
| 17:48:10 | mriedem | well, it goes into _move_operation_alloc_request but doesn't actually double anything | |
| 17:48:14 | mriedem | Nov 06 19:48:36.969659 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: DEBUG nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] New allocation_request containing both source and destination hosts in move operation: {'allocations': {u'3ceb7eab-549c-40ba-a70c-320822c310ab': {'resources': {u'VCPU': 2, u'MEMORY | |
| 17:48:14 | mriedem | : 128}}}} {{(pid=13677) _move_operation_alloc_request /opt/stack/nova/nova/scheduler/client/report.py:234}} | |
| 17:48:27 | mriedem | there is only one provider in that body | |
| 17:51:52 | mriedem | this is the error from placement | |
| 17:51:52 | mriedem | Nov 06 19:48:37.013780 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: WARNING nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] Failed to save allocation for 6665f00a-dcf1-4286-b075-d7dcd7c37487. Got HTTP 409: {"errors": [{"status": 409, "request_id": "req-c9ba6cbd-3b6e-4e5d-b550-9588be8a49d2", "code": "p | |
| 17:51:52 | mriedem | ment.concurrent_update", "detail": "There was a conflict when trying to complete your request.\n\n consumer generation conflict - expected null but got 1 ", "title": "Conflict"}]} | |
| 17:58:38 | mriedem | idk wtf is going on, but i see 3 different PUT allocations in the placement logs for that consumer | |
| 17:58:52 | mriedem | first is probably for the initial scheduler, and then we offload and remove allocations | |
| 17:59:03 | mriedem | 2nd is for the unshelve | |
| 18:07:11 | mriedem | aha | |
| 18:07:16 | mriedem | the allocatoin delete on unshelve changed with this patch https://review.openstack.org/#/c/591597/ | |
| 18:07:23 | mriedem | so we no longer actually delete allocations, we PUT {} | |
| 18:07:49 | sean-k-mooney | quick question. in what cases does nova update the network info cache? | |
| 18:08:13 | sean-k-mooney | i know it does it in respconce to neutron notification. there is also a periodic heal task right | |
| 18:08:17 | sean-k-mooney | is that it? | |
| 18:08:20 | mriedem | in all cases | |
| 18:08:21 | mriedem | attach vifs | |
| 18:08:22 | mriedem | etc | |
| 18:08:58 | mriedem | dansmith: yeah so there are 3 PUTs for allocations, 1st for initial schedule, 2nd for shelve offload (PUT /allocations with {}) and then the 3rd is scheduling during unshelve | |
| 18:09:12 | mriedem | since the allocations aren't deleted on shelve offload, the consumer must persist in placement with it's existing consumer generation | |
| 18:09:17 | mriedem | which the sheduler in the 3rd PUT doesn't account for | |
| 18:09:20 | dansmith | okay | |
| 18:09:38 | dansmith | the scheduler assumes that the consumer is gone? | |
| 18:09:52 | dansmith | I'm surprised it would care, | |
| 18:10:15 | dansmith | because it doesn't know if we're doing an initial boot or a move right? | |
| 18:10:43 | mriedem | it would know if we're doing a move if the consumer already has allocations elsewhere | |
| 18:10:56 | dansmith | right, but otherwise it doesn't, | |
| 18:11:02 | dansmith | and in this case there are no remaining allocations right? | |
| 18:11:09 | mriedem | well, | |
| 18:11:15 | dansmith | or are you saying it assumes that if you have no allocations the consumer can't exist? | |
| 18:11:22 | mriedem | that's not what the scheduler thinks, because it goes down that _move_operation_alloc_request path | |
| 18:11:41 | mriedem | maybe the tempest test isn't really waiting for the instance be fully shelved offloaded before it unshelves | |
| 18:11:59 | mriedem | ah indeed, | |