| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 19:11:19 | mriedem | so i think unit tests are covering _move_operation_alloc_request | |
| 19:11:27 | mriedem | New allocation request containing both source and destination hosts in move operation: {'allocations': [{'resource_provider': {'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 2, u'MEMORY_MB': 192}}]} | |
| 19:11:30 | mriedem | that's from that same method, | |
| 19:11:34 | mriedem | after merging the allocations | |
| 19:11:37 | mriedem | by reference | |
| 19:12:43 | mriedem | ah a red herring | |
| 19:12:53 | mriedem | Successfully claimed resources for instance 6fa7b953-c1fe-4520-8564-aeba8e90aece using allocation request {u'allocations': [{u'resource_provider': {u'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, u'resources': {u'VCPU': 1, u'MEMORY_MB': 128}}]} {{(pid=17280) _claim_resources /opt/stack/new/nova/nova/scheduler/filter_scheduler.py:289}} | |
| 19:12:56 | jaypipes | not for same-host-resize, though, right? | |
| 19:12:56 | mriedem | ^ is stale | |
| 19:13:11 | mriedem | no we doubled correctly | |
| 19:13:13 | mriedem | https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L201 | |
| 19:13:20 | mriedem | New allocation request containing both source and destination hosts in move operation: {'allocations': [{'resource_provider': {'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 2, u'MEMORY_MB': 192}}]} | |
| 19:13:22 | mriedem | that's doubled | |
| 19:13:45 | mriedem | what we're logging here: https://github.com/openstack/nova/blob/master/nova/scheduler/filter_scheduler.py#L293 | |
| 19:13:47 | mriedem | is stale | |
| 19:13:50 | mriedem | that's why i got confused | |
| 19:14:17 | mriedem | however, when we get to the RT, and subtract, we lost the doubled allocation somewhere still | |
| 19:14:33 | mriedem | jaypipes: so i don't know what the bug is yet | |
| 19:14:37 | mriedem | we just have really misleading logging | |
| 19:15:01 | jaypipes | mriedem: I don't agree that the above line is "stale". | |
| 19:15:09 | mriedem | of course it is | |
| 19:15:15 | mriedem | alloc_req comes from placement | |
| 19:15:18 | mriedem | we double it | |
| 19:15:27 | mriedem | then we log the original alloc_req from placement | |
| 19:15:29 | mriedem | which is not double | |
| 19:15:50 | mriedem | follow the 4 log messages starting here http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-sch.txt.gz#_Aug_09_16_27_37_563749 | |
| 19:17:02 | jaypipes | mriedem: yeah, right. | |
| 19:17:30 | mriedem | so, we're still losing the doubled allocation somewhere | |
| 19:17:33 | mriedem | my guess is the RT is overwriting it? | |
| 19:17:46 | jaypipes | mriedem: I will change placement client claim_resources() to return the possibly-updated-for-move-operation alloc_req. | |
| 19:18:16 | mriedem | jaypipes: we don't have to change all of that, just drop alloc_req from this log message https://github.com/openstack/nova/blob/master/nova/scheduler/filter_scheduler.py#L293 | |
| 19:18:25 | mriedem | we already logged the original and new thing | |
| 19:18:36 | jaypipes | k | |
| 19:18:38 | mriedem | well, we logged the new thing | |
| 19:23:51 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Require Placement 1.0 in nova-status upgrade check https://review.openstack.org/492234 | |
| 19:25:38 | mriedem | dansmith: ^ thar she blar | |
| 19:26:11 | sdague | mriedem: commit message? | |
| 19:26:23 | sdague | 1.10 right? | |
| 19:26:27 | dansmith | mriedem: I got distracted by a rant opportunity.. did you figure out the doubling thing? | |
| 19:26:38 | mriedem | crap | |
| 19:26:42 | jaypipes | dansmith: still working on it. | |
| 19:27:06 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Require Placement 1.10 in nova-status upgrade check https://review.openstack.org/492234 | |
| 19:27:12 | mriedem | dansmith: are you ranting in the ops midcycle upgrades etherpad? | |
| 19:27:14 | mriedem | perhaps | |
| 19:27:22 | mriedem | oh -dev, | |
| 19:27:22 | mriedem | i see | |
| 19:27:26 | dansmith | yes | |
| 19:27:44 | dansmith | plus screaming obscenities at the monitor that only Taylor and I can hear | |
| 19:28:25 | dansmith | I can tell it's bad when she just starts slipping me candy to calm me down | |
| 19:29:05 | jaypipes | lol | |
| 19:30:05 | dansmith | jaypipes: "working on it" meaning "working on figuring it out" or "working on putting the fix into code" ? | |
| 19:30:17 | mriedem | we don't have it figured out | |
| 19:30:20 | mriedem | where the doubled allocation went | |
| 19:30:24 | jaypipes | dansmith: what mriedem said. | |
| 19:30:38 | dansmith | and it's not what I said? | |
| 19:30:45 | dansmith | that healing is still blind in this patch | |
| 19:30:50 | mriedem | it might bethat | |
| 19:30:54 | mriedem | that's what i suspect anyway | |
| 19:31:05 | mriedem | one node is ovewriting the doubled allocation | |
| 19:31:12 | mriedem | i should be able to tell that from the logs | |
| 19:31:13 | dansmith | well we're 0.25 days away from getting a run of the top one anyway | |
| 19:32:20 | openstackgerrit | Sean Dague proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517 | |
| 19:33:09 | sdague | dansmith: what's the non gate exposure of this? | |
| 19:33:20 | dansmith | sdague: what? | |
| 19:33:31 | sdague | like which jobs are showing the issue | |
| 19:33:41 | dansmith | none of them are failing | |
| 19:33:47 | dansmith | if that is what you mean | |
| 19:34:04 | sdague | dansmith: one of them is showing a funny allocation though? | |
| 19:34:10 | dansmith | I imagine it's the single-node tempest one that's hitting the issue though, although theoretically the multinode ones should too | |
| 19:34:35 | dansmith | sdague: logging errors as they fail to do their accounting, but nothing is fatal from tempest's point of view it sounds like | |
| 19:34:38 | sdague | just thinking if it's more effective to do local run / debug | |
| 19:34:55 | sdague | given the gate turn around time isn't going to get better any time soon | |
| 19:35:09 | dansmith | probably, but I have so little time left, it'd take me all that to just stack once and start looking I think | |
| 19:35:10 | mriedem | gd we love to lazy-load pci_requests and pci_devices | |
| 19:35:26 | dansmith | I assume jaypipes has been running this locally | |
| 19:36:42 | jaypipes | dansmith: not in the last few weeks, no. been relying on functional tests and logging. | |
| 19:39:04 | jaypipes | mriedem: ok if I remove that log message in scheduler _claim_resources() (with the stale alloc_req) in the bottom patch of this series? the patch call "refactor heal..." | |
| 19:39:31 | jaypipes | nm, I'll just throw it in another patch | |
| 19:39:34 | jaypipes | patches are cheap. | |
| 19:39:46 | mriedem | http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_46_100320 | |
| 19:40:37 | mriedem | yeah so on resize to same host, it's the _instance_to_allocations_dict that overwrites the doubled allocation | |
| 19:40:46 | mriedem | dansmith: which is what you were saying | |
| 19:40:54 | mriedem | Sending allocation for instance {'VCPU': 1, 'MEMORY_MB': 64} {{(pid=19567) _allocate_for_instance /opt/stack/new/nova/nova/scheduler/client/report.py:924}} | |
| 19:40:59 | dansmith | yeah | |
| 19:41:01 | jaypipes | mriedem: right, and that's the thing that doens't get "corrected" until the later patch | |
| 19:41:14 | mriedem | that's at Aug 09 16:27:46.100320 | |
| 19:41:29 | mriedem | the doubled up allocation in the scheduler was at Aug 09 16:27:37.633725 | |
| 19:41:58 | mriedem | ok so maybe not a turd furgeson after all | |
| 19:42:20 | mriedem | although, an ocata compute will trample that and then the last patch will double the allocations again? | |
| 19:42:40 | dansmith | mriedem: while we're healing, | |
| 19:42:47 | dansmith | we'll fix it eventually on the destination node | |
| 19:43:26 | dansmith | once we're all pike, we don't trample and thus we're good | |
| 19:44:09 | mriedem | but we'll continue to log ERRORs? | |
| 19:44:41 | dansmith | I dunno I didn't look at the actual site where we log that, but probably so | |
| 19:45:09 | mriedem | ok, and we probably don't have a job that would test it either, | |
| 19:45:23 | mriedem | would require a grenade multinode job that runs a migration | |
| 19:45:34 | mriedem | we do have a multinode grenade job that runs live migration... | |
| 19:45:45 | mriedem | back and forth too so that actually might show it | |
| 19:45:48 | dansmith | mriedem: you can borrow my pistol when I'm done with it | |
| 19:46:35 | mriedem | i'll be classy and do Seppuku | |