Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
19:26:38 mriedem crap
19:26:42 jaypipes dansmith: still working on it.
19:27:06 openstackgerrit Matt Riedemann proposed openstack/nova master: Require Placement 1.10 in nova-status upgrade check https://review.openstack.org/492234
19:27:12 mriedem dansmith: are you ranting in the ops midcycle upgrades etherpad?
19:27:14 mriedem perhaps
19:27:22 mriedem oh -dev,
19:27:22 mriedem i see
19:27:26 dansmith yes
19:27:44 dansmith plus screaming obscenities at the monitor that only Taylor and I can hear
19:28:25 dansmith I can tell it's bad when she just starts slipping me candy to calm me down
19:29:05 jaypipes lol
19:30:05 dansmith jaypipes: "working on it" meaning "working on figuring it out" or "working on putting the fix into code" ?
19:30:17 mriedem we don't have it figured out
19:30:20 mriedem where the doubled allocation went
19:30:24 jaypipes dansmith: what mriedem said.
19:30:38 dansmith and it's not what I said?
19:30:45 dansmith that healing is still blind in this patch
19:30:50 mriedem it might bethat
19:30:54 mriedem that's what i suspect anyway
19:31:05 mriedem one node is ovewriting the doubled allocation
19:31:12 mriedem i should be able to tell that from the logs
19:31:13 dansmith well we're 0.25 days away from getting a run of the top one anyway
19:32:20 openstackgerrit Sean Dague proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517
19:33:09 sdague dansmith: what's the non gate exposure of this?
19:33:20 dansmith sdague: what?
19:33:31 sdague like which jobs are showing the issue
19:33:41 dansmith none of them are failing
19:33:47 dansmith if that is what you mean
19:34:04 sdague dansmith: one of them is showing a funny allocation though?
19:34:10 dansmith I imagine it's the single-node tempest one that's hitting the issue though, although theoretically the multinode ones should too
19:34:35 dansmith sdague: logging errors as they fail to do their accounting, but nothing is fatal from tempest's point of view it sounds like
19:34:38 sdague just thinking if it's more effective to do local run / debug
19:34:55 sdague given the gate turn around time isn't going to get better any time soon
19:35:09 dansmith probably, but I have so little time left, it'd take me all that to just stack once and start looking I think
19:35:10 mriedem gd we love to lazy-load pci_requests and pci_devices
19:35:26 dansmith I assume jaypipes has been running this locally
19:36:42 jaypipes dansmith: not in the last few weeks, no. been relying on functional tests and logging.
19:39:04 jaypipes mriedem: ok if I remove that log message in scheduler _claim_resources() (with the stale alloc_req) in the bottom patch of this series? the patch call "refactor heal..."
19:39:31 jaypipes nm, I'll just throw it in another patch
19:39:34 jaypipes patches are cheap.
19:39:46 mriedem http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_46_100320
19:40:37 mriedem yeah so on resize to same host, it's the _instance_to_allocations_dict that overwrites the doubled allocation
19:40:46 mriedem dansmith: which is what you were saying
19:40:54 mriedem Sending allocation for instance {'VCPU': 1, 'MEMORY_MB': 64} {{(pid=19567) _allocate_for_instance /opt/stack/new/nova/nova/scheduler/client/report.py:924}}
19:40:59 dansmith yeah
19:41:01 jaypipes mriedem: right, and that's the thing that doens't get "corrected" until the later patch
19:41:14 mriedem that's at Aug 09 16:27:46.100320
19:41:29 mriedem the doubled up allocation in the scheduler was at Aug 09 16:27:37.633725
19:41:58 mriedem ok so maybe not a turd furgeson after all
19:42:20 mriedem although, an ocata compute will trample that and then the last patch will double the allocations again?
19:42:40 dansmith mriedem: while we're healing,
19:42:47 dansmith we'll fix it eventually on the destination node
19:43:26 dansmith once we're all pike, we don't trample and thus we're good
19:44:09 mriedem but we'll continue to log ERRORs?
19:44:41 dansmith I dunno I didn't look at the actual site where we log that, but probably so
19:45:09 mriedem ok, and we probably don't have a job that would test it either,
19:45:23 mriedem would require a grenade multinode job that runs a migration
19:45:34 mriedem we do have a multinode grenade job that runs live migration...
19:45:45 mriedem back and forth too so that actually might show it
19:45:48 dansmith mriedem: you can borrow my pistol when I'm done with it
19:46:35 mriedem i'll be classy and do Seppuku
19:46:51 dansmith heh, less mess, how midwestern of you
19:47:07 mriedem disemboweling is still a mess
19:47:26 mriedem i'll do it in the soaker tub
19:48:15 mriedem this error log case wouldn't be the end of the world, we could release with that and backport a fix later,
19:48:26 mriedem adding some logic to determine we're about to push a 0 allocation and just not do that
19:49:03 dansmith well, normally that would be a legit error case,
19:49:07 dansmith and worth the log
19:49:26 dansmith so I dunno.. maybe if not has_ocata and ==0, then log error or something
19:51:08 mriedem my point is, it doesn't cause things to fail, so not a regression, so we could cleanup the error log later and backport if necessary, yeah?
19:51:26 dansmith yes
19:51:51 openstackgerrit Merged openstack/nova-specs master: doc: Remove crud from Sphinx conf.py https://review.openstack.org/456129
19:52:10 mriedem jaypipes: so you're putting a patch at the bottom of the series to remove that bogus log in the filter scheduler, and then fixing the pep8 in https://review.openstack.org/#/c/491850/ on top of that?
19:52:40 mriedem and then i think we can limp https://review.openstack.org/#/c/488510/ in and fix my other problems in a follow up
19:53:11 mriedem and then once we get a good run on https://review.openstack.org/#/c/491012/ we call it a ball for rc1
19:53:19 mriedem oh except this of course https://review.openstack.org/#/c/491491/
19:53:42 dansmith jay needs to look at that one I think
19:53:53 dansmith it makes sense to me, but I figure he should look
19:53:54 mriedem gibi's?
19:53:57 mriedem yes agree
19:53:59 dansmith yeah
19:53:59 mriedem he wrote it
19:54:03 jaypipes I'm still addressing mriedem's review comments on the confirm/reverrt resize patch.
19:54:06 mriedem i mean, the code that's changing
19:54:07 jaypipes should be done shortly.
19:54:11 dansmith mriedem: right
19:54:29 mriedem although it does say, "this should really raise NoValidHost but tests...."
19:58:51 mriedem what i don't understand in his test https://review.openstack.org/#/c/490814/ is how the resize doesn't fail on the compute during the claim
19:58:55 mriedem the claim should fail
20:00:02 dansmith allocation ratio?
20:00:23 dansmith he's not setting it so I would expect it's defaulting to 16 or whatever
20:00:24 dansmith for vcpu
20:00:34 mriedem default=0.0,
20:00:40 dansmith which means 16
20:00:48 mriedem if set to 0.0, the value
20:00:48 mriedem set on the scheduler node(s) or compute node(s) will be used
20:00:48 mriedem and defaulted to 16.0.
20:00:48 mriedem yeah
20:01:15 mriedem haha
20:01:47 mriedem i'll pull his test down, set that to 1.0 and make sure it fails the resize, and then if that's the case i'll be cool with the test

Earlier   Later