| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-11-09 | |||
| 17:46:55 | mriedem | that's why we don't retry | |
| 17:46:58 | dansmith | yeah | |
| 17:47:09 | mriedem | i also don't know why it thinks we're starting with existing allocations for a shelved offloaded server | |
| 17:47:44 | dansmith | and that's causing it to try to double? | |
| 17:48:10 | mriedem | well, it goes into _move_operation_alloc_request but doesn't actually double anything | |
| 17:48:14 | mriedem | Nov 06 19:48:36.969659 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: DEBUG nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] New allocation_request containing both source and destination hosts in move operation: {'allocations': {u'3ceb7eab-549c-40ba-a70c-320822c310ab': {'resources': {u'VCPU': 2, u'MEMORY | |
| 17:48:14 | mriedem | : 128}}}} {{(pid=13677) _move_operation_alloc_request /opt/stack/nova/nova/scheduler/client/report.py:234}} | |
| 17:48:27 | mriedem | there is only one provider in that body | |
| 17:51:52 | mriedem | this is the error from placement | |
| 17:51:52 | mriedem | Nov 06 19:48:37.013780 ubuntu-xenial-inap-mtl01-0000379614 nova-scheduler[12154]: WARNING nova.scheduler.client.report [None req-f266a0ff-2840-413d-9877-4500e61512f5 tempest-ServersNegativeTestJSON-477704048 tempest-ServersNegativeTestJSON-477704048] Failed to save allocation for 6665f00a-dcf1-4286-b075-d7dcd7c37487. Got HTTP 409: {"errors": [{"status": 409, "request_id": "req-c9ba6cbd-3b6e-4e5d-b550-9588be8a49d2", "code": "p | |
| 17:51:52 | mriedem | ment.concurrent_update", "detail": "There was a conflict when trying to complete your request.\n\n consumer generation conflict - expected null but got 1 ", "title": "Conflict"}]} | |
| 17:58:38 | mriedem | idk wtf is going on, but i see 3 different PUT allocations in the placement logs for that consumer | |
| 17:58:52 | mriedem | first is probably for the initial scheduler, and then we offload and remove allocations | |
| 17:59:03 | mriedem | 2nd is for the unshelve | |
| 18:07:11 | mriedem | aha | |
| 18:07:16 | mriedem | the allocatoin delete on unshelve changed with this patch https://review.openstack.org/#/c/591597/ | |
| 18:07:23 | mriedem | so we no longer actually delete allocations, we PUT {} | |
| 18:07:49 | sean-k-mooney | quick question. in what cases does nova update the network info cache? | |
| 18:08:13 | sean-k-mooney | i know it does it in respconce to neutron notification. there is also a periodic heal task right | |
| 18:08:17 | sean-k-mooney | is that it? | |
| 18:08:20 | mriedem | in all cases | |
| 18:08:21 | mriedem | attach vifs | |
| 18:08:22 | mriedem | etc | |
| 18:08:58 | mriedem | dansmith: yeah so there are 3 PUTs for allocations, 1st for initial schedule, 2nd for shelve offload (PUT /allocations with {}) and then the 3rd is scheduling during unshelve | |
| 18:09:12 | mriedem | since the allocations aren't deleted on shelve offload, the consumer must persist in placement with it's existing consumer generation | |
| 18:09:17 | mriedem | which the sheduler in the 3rd PUT doesn't account for | |
| 18:09:20 | dansmith | okay | |
| 18:09:38 | dansmith | the scheduler assumes that the consumer is gone? | |
| 18:09:52 | dansmith | I'm surprised it would care, | |
| 18:10:15 | dansmith | because it doesn't know if we're doing an initial boot or a move right? | |
| 18:10:43 | mriedem | it would know if we're doing a move if the consumer already has allocations elsewhere | |
| 18:10:56 | dansmith | right, but otherwise it doesn't, | |
| 18:11:02 | dansmith | and in this case there are no remaining allocations right? | |
| 18:11:09 | mriedem | well, | |
| 18:11:15 | dansmith | or are you saying it assumes that if you have no allocations the consumer can't exist? | |
| 18:11:22 | mriedem | that's not what the scheduler thinks, because it goes down that _move_operation_alloc_request path | |
| 18:11:41 | mriedem | maybe the tempest test isn't really waiting for the instance be fully shelved offloaded before it unshelves | |
| 18:11:59 | mriedem | ah indeed, | |
| 18:12:13 | mriedem | we set the instance vm_state to SHELVED_OFFLOADED *before* we remove allocations | |
| 18:12:14 | dansmith | I guess I'm still surprised it cares | |
| 18:12:16 | mriedem | which is definitely a race | |
| 18:16:29 | dansmith | why? | |
| 18:16:59 | dansmith | just because it signals to tempest that it's done? | |
| 18:17:04 | dansmith | if we didn't, | |
| 18:17:25 | dansmith | and we crashed right between deleting the allocations and makring it as offloaded we'd have lost some information I would think | |
| 18:17:46 | mriedem | well, this also doesn't seem right | |
| 18:17:47 | mriedem | https://review.openstack.org/#/c/591597/8/nova/scheduler/client/report.py@2091 | |
| 18:17:52 | mriedem | "# removing all resources from the allocation will auto delete the | |
| 18:17:52 | mriedem | # consumer in placement" | |
| 18:18:01 | dansmith | I guess maybe your point is that it's a race between starting the unshelve and there still being allocations... | |
| 18:18:02 | mriedem | maybe that is happening, i'm not sure | |
| 18:18:06 | mriedem | correct | |
| 18:18:19 | mriedem | b/c i'm seeing this being true during unshelve https://github.com/openstack/nova/blob/e27905f482ba26d2bbf3ae5d948dee37523042d5/nova/scheduler/client/report.py#L1824 | |
| 18:18:21 | dansmith | not sure that's better either way | |
| 18:18:22 | mriedem | which shouldn't be the case | |
| 18:19:52 | dansmith | it's certainly possible that the move to PUT{} from DELETE didn't bring over some "and delete the consumer" part | |
| 18:20:14 | dansmith | but again, I'm not sure why it should matter to the scheduler that it exists | |
| 18:20:15 | dansmith | although... | |
| 18:20:26 | dansmith | we have no api for looking at the consumer to get the generation if it already exists, IIRC | |
| 18:20:43 | dansmith | so maybe without seeing an allocation, and not being able to see the consumer directly, we have no alternative? | |
| 18:20:45 | mriedem | right, it's supposed to come back on the GET /allocations/{consumer} call | |
| 18:20:59 | dansmith | I expect jaypipes to pop in here any second and say "ah hah!" | |
| 18:21:19 | mriedem | jaypipes is busy chefing it up | |
| 18:21:26 | dansmith | his fave | |
| 18:21:33 | mriedem | and getting ready to sleep for a week while the rest of us are in berlin | |
| 18:21:42 | dansmith | lucky bastard | |
| 18:21:57 | dansmith | I'm going to be landing pretty soon, FYI | |
| 18:22:04 | mriedem | i'm overdue for lunch as well | |
| 18:22:50 | mriedem | so, can't claim_resources in the scheduler still just retry if it hits that consumer generatoin conflict? | |
| 18:22:57 | dansmith | well, | |
| 18:23:01 | dansmith | not if it doesn't know what the generation is | |
| 18:23:10 | dansmith | that's what I was saying.. it might not be able to find out what it is, | |
| 18:23:19 | dansmith | with no consumer api and no existing allocation to look at | |
| 18:23:52 | mriedem | i would expect GET /allocations/{consumer_uuid} to return the consumer generation | |
| 18:23:57 | mriedem | even if allocations are {} | |
| 18:24:01 | dansmith | with no alloc records? | |
| 18:24:02 | dansmith | I dunno | |
| 18:24:04 | mriedem | but i guess i'd have to dig into the placement code | |
| 18:24:21 | dansmith | I would expect that code returns 404 if none come back, | |
| 18:24:28 | mriedem | should probably also log in placement when the consumer is deleted b/c allocations went to 0 | |
| 18:24:38 | dansmith | because it would only get the consume through the join, or afterwards I would expect | |
| 18:24:44 | mriedem | no it doesn't 404, you get {"allocations": {}} | |
| 18:24:52 | mriedem | if there are no allocations for the consumer | |
| 18:24:52 | dansmith | oh? | |
| 18:24:57 | mriedem | yeah it's confusing | |
| 18:25:06 | dansmith | that seems supremely weird to me, but okay | |
| 18:25:37 | mriedem | what does taylor think about all this? | |
| 18:26:03 | dansmith | she's busy with her own work | |
| 18:26:19 | jaypipes | mriedem: fuck chef. fuck ansible. fuck docker. it's all a bunch of complete assbaggery. | |
| 18:26:31 | dansmith | my little mobile wifi router lets us share the same crappy airline wifi, so after that came online, I might as well not be sitting next to her | |
| 18:26:39 | mriedem | jaypipes: but salt?! | |
| 18:27:02 | dansmith | jaypipes: you're gonna love it | |
| 18:27:03 | sean-k-mooney | mriedem: i think jaypipes has enough salt in his life right now | |
| 18:27:24 | mriedem | jaypipes: notes are in https://bugs.launchpad.net/nova/+bug/1798688 | |
| 18:27:25 | openstack | Launchpad bug 1798688 in OpenStack Compute (nova) "AllocationUpdateFailed_Remote: Failed to update allocations for consumer. Error: another process changed the consumer after the report client read the consumer state during the claim" [Undecided,Triaged] | |
| 18:27:37 | dansmith | jaypipes: question.. should placement have a consumers endpoint? | |
| 18:27:48 | mriedem | looks like we're racing between shelve offload (allocation removal) and unshelve (put new allocations) and hitting a consumer generation conflict | |
| 18:28:21 | sean-k-mooney | jaypipes: at least you dont have to use tripplo where we use yaml to drive heat to drive puppet to drive ansible to deploy docker containers ... | |
| 18:28:50 | sean-k-mooney | or maybe the ansible drives puppet its hard to keep track of | |
| 18:29:21 | dansmith | sean-k-mooney: it drives ... me insane | |