Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
19:30:24 jaypipes dansmith: what mriedem said.
19:30:38 dansmith and it's not what I said?
19:30:45 dansmith that healing is still blind in this patch
19:30:50 mriedem it might bethat
19:30:54 mriedem that's what i suspect anyway
19:31:05 mriedem one node is ovewriting the doubled allocation
19:31:12 mriedem i should be able to tell that from the logs
19:31:13 dansmith well we're 0.25 days away from getting a run of the top one anyway
19:32:20 openstackgerrit Sean Dague proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517
19:33:09 sdague dansmith: what's the non gate exposure of this?
19:33:20 dansmith sdague: what?
19:33:31 sdague like which jobs are showing the issue
19:33:41 dansmith none of them are failing
19:33:47 dansmith if that is what you mean
19:34:04 sdague dansmith: one of them is showing a funny allocation though?
19:34:10 dansmith I imagine it's the single-node tempest one that's hitting the issue though, although theoretically the multinode ones should too
19:34:35 dansmith sdague: logging errors as they fail to do their accounting, but nothing is fatal from tempest's point of view it sounds like
19:34:38 sdague just thinking if it's more effective to do local run / debug
19:34:55 sdague given the gate turn around time isn't going to get better any time soon
19:35:09 dansmith probably, but I have so little time left, it'd take me all that to just stack once and start looking I think
19:35:10 mriedem gd we love to lazy-load pci_requests and pci_devices
19:35:26 dansmith I assume jaypipes has been running this locally
19:36:42 jaypipes dansmith: not in the last few weeks, no. been relying on functional tests and logging.
19:39:04 jaypipes mriedem: ok if I remove that log message in scheduler _claim_resources() (with the stale alloc_req) in the bottom patch of this series? the patch call "refactor heal..."
19:39:31 jaypipes nm, I'll just throw it in another patch
19:39:34 jaypipes patches are cheap.
19:39:46 mriedem http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_46_100320
19:40:37 mriedem yeah so on resize to same host, it's the _instance_to_allocations_dict that overwrites the doubled allocation
19:40:46 mriedem dansmith: which is what you were saying
19:40:54 mriedem Sending allocation for instance {'VCPU': 1, 'MEMORY_MB': 64} {{(pid=19567) _allocate_for_instance /opt/stack/new/nova/nova/scheduler/client/report.py:924}}
19:40:59 dansmith yeah
19:41:01 jaypipes mriedem: right, and that's the thing that doens't get "corrected" until the later patch
19:41:14 mriedem that's at Aug 09 16:27:46.100320
19:41:29 mriedem the doubled up allocation in the scheduler was at Aug 09 16:27:37.633725
19:41:58 mriedem ok so maybe not a turd furgeson after all
19:42:20 mriedem although, an ocata compute will trample that and then the last patch will double the allocations again?
19:42:40 dansmith mriedem: while we're healing,
19:42:47 dansmith we'll fix it eventually on the destination node
19:43:26 dansmith once we're all pike, we don't trample and thus we're good
19:44:09 mriedem but we'll continue to log ERRORs?
19:44:41 dansmith I dunno I didn't look at the actual site where we log that, but probably so
19:45:09 mriedem ok, and we probably don't have a job that would test it either,
19:45:23 mriedem would require a grenade multinode job that runs a migration
19:45:34 mriedem we do have a multinode grenade job that runs live migration...
19:45:45 mriedem back and forth too so that actually might show it
19:45:48 dansmith mriedem: you can borrow my pistol when I'm done with it
19:46:35 mriedem i'll be classy and do Seppuku
19:46:51 dansmith heh, less mess, how midwestern of you
19:47:07 mriedem disemboweling is still a mess
19:47:26 mriedem i'll do it in the soaker tub
19:48:15 mriedem this error log case wouldn't be the end of the world, we could release with that and backport a fix later,
19:48:26 mriedem adding some logic to determine we're about to push a 0 allocation and just not do that
19:49:03 dansmith well, normally that would be a legit error case,
19:49:07 dansmith and worth the log
19:49:26 dansmith so I dunno.. maybe if not has_ocata and ==0, then log error or something
19:51:08 mriedem my point is, it doesn't cause things to fail, so not a regression, so we could cleanup the error log later and backport if necessary, yeah?
19:51:26 dansmith yes
19:51:51 openstackgerrit Merged openstack/nova-specs master: doc: Remove crud from Sphinx conf.py https://review.openstack.org/456129
19:52:10 mriedem jaypipes: so you're putting a patch at the bottom of the series to remove that bogus log in the filter scheduler, and then fixing the pep8 in https://review.openstack.org/#/c/491850/ on top of that?
19:52:40 mriedem and then i think we can limp https://review.openstack.org/#/c/488510/ in and fix my other problems in a follow up
19:53:11 mriedem and then once we get a good run on https://review.openstack.org/#/c/491012/ we call it a ball for rc1
19:53:19 mriedem oh except this of course https://review.openstack.org/#/c/491491/
19:53:42 dansmith jay needs to look at that one I think
19:53:53 dansmith it makes sense to me, but I figure he should look
19:53:54 mriedem gibi's?
19:53:57 mriedem yes agree
19:53:59 dansmith yeah
19:53:59 mriedem he wrote it
19:54:03 jaypipes I'm still addressing mriedem's review comments on the confirm/reverrt resize patch.
19:54:06 mriedem i mean, the code that's changing
19:54:07 jaypipes should be done shortly.
19:54:11 dansmith mriedem: right
19:54:29 mriedem although it does say, "this should really raise NoValidHost but tests...."
19:58:51 mriedem what i don't understand in his test https://review.openstack.org/#/c/490814/ is how the resize doesn't fail on the compute during the claim
19:58:55 mriedem the claim should fail
20:00:02 dansmith allocation ratio?
20:00:23 dansmith he's not setting it so I would expect it's defaulting to 16 or whatever
20:00:24 dansmith for vcpu
20:00:34 mriedem default=0.0,
20:00:40 dansmith which means 16
20:00:48 mriedem if set to 0.0, the value
20:00:48 mriedem set on the scheduler node(s) or compute node(s) will be used
20:00:48 mriedem and defaulted to 16.0.
20:00:48 mriedem yeah
20:01:15 mriedem haha
20:01:47 mriedem i'll pull his test down, set that to 1.0 and make sure it fails the resize, and then if that's the case i'll be cool with the test
20:02:41 dansmith so, um
20:02:49 dansmith we really fall back to legacy scheduling in that case?
20:02:54 dansmith I don't think I knew that was the plan
20:03:33 openstackgerrit Jay Pipes proposed openstack/nova master: placement: refactor healing of allocations in RT https://review.openstack.org/491850
20:03:33 openstackgerrit Jay Pipes proposed openstack/nova master: Remove provider allocs in confirm/revert resize https://review.openstack.org/488510
20:03:34 openstackgerrit Jay Pipes proposed openstack/nova master: Resource tracker compatibility with Ocata and Pike https://review.openstack.org/491012
20:03:35 openstackgerrit Jay Pipes proposed openstack/nova master: remove log message with potential stale info https://review.openstack.org/492242
20:03:46 jaypipes mriedem, dansmith: ^.
20:03:51 mriedem well we wouldn't with gibi's patch
20:04:04 dansmith jaypipes: I'm not reviewing any more things today unless mriedem +2s them first
20:04:08 dansmith my stats and ego just can't take it
20:04:59 mriedem if it makes you feel better, maya told laura that she didn't want to go to camp this morning b/c she thought i'd forget her again
20:06:32 dansmith hah that's pretty awesome
20:06:57 jaypipes lol Dad of the Yea.

Earlier   Later