Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
19:08:12 mriedem which i guess could be follow ups
19:08:22 mriedem i'm just kind of getting tired of making myself TODOs about chasing follow ups
19:08:25 mriedem but realize we're short on time
19:08:58 jaypipes mriedem: I'm happy to address your comments on the second patch. just need to know if dansmith is actively working on these.
19:09:08 dansmith nope
19:09:44 jaypipes dansmith: gotcha. I'll get on this then.
19:10:37 cdent k, in that case I’ll look at the existing main stack in the morning and see where we’re at
19:10:40 cdent g’night
19:10:46 mriedem o/
19:11:19 mriedem so i think unit tests are covering _move_operation_alloc_request
19:11:27 mriedem New allocation request containing both source and destination hosts in move operation: {'allocations': [{'resource_provider': {'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 2, u'MEMORY_MB': 192}}]}
19:11:30 mriedem that's from that same method,
19:11:34 mriedem after merging the allocations
19:11:37 mriedem by reference
19:12:43 mriedem ah a red herring
19:12:53 mriedem Successfully claimed resources for instance 6fa7b953-c1fe-4520-8564-aeba8e90aece using allocation request {u'allocations': [{u'resource_provider': {u'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, u'resources': {u'VCPU': 1, u'MEMORY_MB': 128}}]} {{(pid=17280) _claim_resources /opt/stack/new/nova/nova/scheduler/filter_scheduler.py:289}}
19:12:56 jaypipes not for same-host-resize, though, right?
19:12:56 mriedem ^ is stale
19:13:11 mriedem no we doubled correctly
19:13:13 mriedem https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L201
19:13:20 mriedem New allocation request containing both source and destination hosts in move operation: {'allocations': [{'resource_provider': {'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 2, u'MEMORY_MB': 192}}]}
19:13:22 mriedem that's doubled
19:13:45 mriedem what we're logging here: https://github.com/openstack/nova/blob/master/nova/scheduler/filter_scheduler.py#L293
19:13:47 mriedem is stale
19:13:50 mriedem that's why i got confused
19:14:17 mriedem however, when we get to the RT, and subtract, we lost the doubled allocation somewhere still
19:14:33 mriedem jaypipes: so i don't know what the bug is yet
19:14:37 mriedem we just have really misleading logging
19:15:01 jaypipes mriedem: I don't agree that the above line is "stale".
19:15:09 mriedem of course it is
19:15:15 mriedem alloc_req comes from placement
19:15:18 mriedem we double it
19:15:27 mriedem then we log the original alloc_req from placement
19:15:29 mriedem which is not double
19:15:50 mriedem follow the 4 log messages starting here http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-sch.txt.gz#_Aug_09_16_27_37_563749
19:17:02 jaypipes mriedem: yeah, right.
19:17:30 mriedem so, we're still losing the doubled allocation somewhere
19:17:33 mriedem my guess is the RT is overwriting it?
19:17:46 jaypipes mriedem: I will change placement client claim_resources() to return the possibly-updated-for-move-operation alloc_req.
19:18:16 mriedem jaypipes: we don't have to change all of that, just drop alloc_req from this log message https://github.com/openstack/nova/blob/master/nova/scheduler/filter_scheduler.py#L293
19:18:25 mriedem we already logged the original and new thing
19:18:36 jaypipes k
19:18:38 mriedem well, we logged the new thing
19:23:51 openstackgerrit Matt Riedemann proposed openstack/nova master: Require Placement 1.0 in nova-status upgrade check https://review.openstack.org/492234
19:25:38 mriedem dansmith: ^ thar she blar
19:26:11 sdague mriedem: commit message?
19:26:23 sdague 1.10 right?
19:26:27 dansmith mriedem: I got distracted by a rant opportunity.. did you figure out the doubling thing?
19:26:38 mriedem crap
19:26:42 jaypipes dansmith: still working on it.
19:27:06 openstackgerrit Matt Riedemann proposed openstack/nova master: Require Placement 1.10 in nova-status upgrade check https://review.openstack.org/492234
19:27:12 mriedem dansmith: are you ranting in the ops midcycle upgrades etherpad?
19:27:14 mriedem perhaps
19:27:22 mriedem oh -dev,
19:27:22 mriedem i see
19:27:26 dansmith yes
19:27:44 dansmith plus screaming obscenities at the monitor that only Taylor and I can hear
19:28:25 dansmith I can tell it's bad when she just starts slipping me candy to calm me down
19:29:05 jaypipes lol
19:30:05 dansmith jaypipes: "working on it" meaning "working on figuring it out" or "working on putting the fix into code" ?
19:30:17 mriedem we don't have it figured out
19:30:20 mriedem where the doubled allocation went
19:30:24 jaypipes dansmith: what mriedem said.
19:30:38 dansmith and it's not what I said?
19:30:45 dansmith that healing is still blind in this patch
19:30:50 mriedem it might bethat
19:30:54 mriedem that's what i suspect anyway
19:31:05 mriedem one node is ovewriting the doubled allocation
19:31:12 mriedem i should be able to tell that from the logs
19:31:13 dansmith well we're 0.25 days away from getting a run of the top one anyway
19:32:20 openstackgerrit Sean Dague proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517
19:33:09 sdague dansmith: what's the non gate exposure of this?
19:33:20 dansmith sdague: what?
19:33:31 sdague like which jobs are showing the issue
19:33:41 dansmith none of them are failing
19:33:47 dansmith if that is what you mean
19:34:04 sdague dansmith: one of them is showing a funny allocation though?
19:34:10 dansmith I imagine it's the single-node tempest one that's hitting the issue though, although theoretically the multinode ones should too
19:34:35 dansmith sdague: logging errors as they fail to do their accounting, but nothing is fatal from tempest's point of view it sounds like
19:34:38 sdague just thinking if it's more effective to do local run / debug
19:34:55 sdague given the gate turn around time isn't going to get better any time soon
19:35:09 dansmith probably, but I have so little time left, it'd take me all that to just stack once and start looking I think
19:35:10 mriedem gd we love to lazy-load pci_requests and pci_devices
19:35:26 dansmith I assume jaypipes has been running this locally
19:36:42 jaypipes dansmith: not in the last few weeks, no. been relying on functional tests and logging.
19:39:04 jaypipes mriedem: ok if I remove that log message in scheduler _claim_resources() (with the stale alloc_req) in the bottom patch of this series? the patch call "refactor heal..."
19:39:31 jaypipes nm, I'll just throw it in another patch
19:39:34 jaypipes patches are cheap.
19:39:46 mriedem http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_46_100320
19:40:37 mriedem yeah so on resize to same host, it's the _instance_to_allocations_dict that overwrites the doubled allocation
19:40:46 mriedem dansmith: which is what you were saying
19:40:54 mriedem Sending allocation for instance {'VCPU': 1, 'MEMORY_MB': 64} {{(pid=19567) _allocate_for_instance /opt/stack/new/nova/nova/scheduler/client/report.py:924}}
19:40:59 dansmith yeah
19:41:01 jaypipes mriedem: right, and that's the thing that doens't get "corrected" until the later patch
19:41:14 mriedem that's at Aug 09 16:27:46.100320
19:41:29 mriedem the doubled up allocation in the scheduler was at Aug 09 16:27:37.633725
19:41:58 mriedem ok so maybe not a turd furgeson after all
19:42:20 mriedem although, an ocata compute will trample that and then the last patch will double the allocations again?
19:42:40 dansmith mriedem: while we're healing,
19:42:47 dansmith we'll fix it eventually on the destination node

Earlier   Later