Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
18:05:49 dansmith caching scheduler is filter scheduler, with cached host list data
18:06:18 mriedem dansmith: so this resize/rt series has a pep8 failure in the bottom change,
18:06:31 mriedem dansmith: but i guess hold off on fixing that so i can go through them now
18:06:33 mriedem the last 2 i mean
18:06:41 dansmith okay
18:06:46 mriedem dansmith: it's T-~3 hours to 0 dan time yeah?
18:06:53 dansmith check queue is 7h long
18:07:26 dansmith mriedem: I think you mean T-3 hours until free-of-dan time
18:07:27 dansmith but yeah
18:07:55 dansmith I assume everyone brought liquor and cake to work today for the post-2pm PDT party
18:08:14 mriedem brought? i think it's just generally available
18:08:47 dansmith I was trying to act like it was the 90s
18:13:40 mnaser cfriesen thanks for all your feedback, i think the best solution for now is going to be use 0,16 to have predictable performance for the host and assign hugepages with a slight bias towards node1 because it has 2 more cores so things can get properly packed
18:23:37 sdague dansmith: yeh, we're down to 800 nodes in ci
18:23:54 sdague which is unfortunately not sufficient
18:24:16 dansmith well, it was fun while it lasted
18:25:50 sdague if anyone wants to do doc reviews to hopefully complete the import - https://review.openstack.org/#/q/status:open+project:openstack/nova+branch:master+topic:bp/doc-migration
18:30:06 cfriesen mnaser: sorry you hit a bug. :) we should be able to get it sorted out once sfinucan comes back from holidays
18:31:07 cfriesen mnaser: ping me with the bug number once you file it please
18:41:12 mriedem dansmith: jaypipes: cdent: edleafe: problems https://review.openstack.org/#/c/488510/
18:42:23 dansmith mriedem: I +2d before we had jenkins runs of the latest version
18:42:27 dansmith and it was passing (tempest) before
18:42:35 mriedem it's not that it's not passing tempest
18:42:41 mriedem it's just spewing ERRORs
18:42:46 jaypipes hey guys, just got back.
18:43:33 jaypipes hmm, yes, the dreaded ./nova/compute/resource_tracker.py:1112:17: N352 LOG.warn is deprecated, please use LOG.warning!
18:44:12 mriedem heh i even pointed that out
18:44:17 mriedem it's in the bottom change btw
18:44:23 mriedem so the entire series needs to be rebaesd
18:44:24 mriedem *rebased
18:44:32 mriedem but i have issues with the 2nd change in the series
18:45:02 jaypipes yeah, I see that
18:45:14 dansmith mriedem: we have the placement upgrade before nova thing required in the prelude of the renos
18:45:21 dansmith which should be enough to merge this, IMHO
18:45:44 jaypipes "prelude of the renos" sounds like some classical music composition.
18:46:25 mriedem dansmith: where?
18:46:26 dansmith I probably shouldn't be reviewing this anyway since I was elbow deep in it myself
18:46:28 mriedem i know it's in the release notes
18:46:35 mriedem but it's not in the prelude
18:46:53 mriedem https://docs.openstack.org/releasenotes/nova/unreleased.html#id1
18:46:58 mriedem A new 1.10 API microversion is added to the Placement REST API. This microversion adds support for the GET /allocation_candidates resource endpoint. This endpoint returns information about possible allocation requests that callers can make which meet a set of resource constraints supplied as query string parameters. Also returned is some inventory and capacity information for the resource providers involved in the allocation
18:46:58 mriedem idates.
18:47:19 mriedem wrong one
18:47:19 mriedem https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes
18:47:25 mriedem The scheduler now requests allocation candidates from the Placement service during scheduling. The allocation candidates information was introduced in the Placement API 1.10 microversion, so you should upgrade the placement service before the Nova scheduler service so that the scheduler can take advantage of the allocation candidate information.
18:47:42 mriedem the code totally falls back though if 1.10 isn't available
18:48:08 mriedem if we want to make 1.10 required, that's fine, but we need to also update nova-status, which i could do lickety split
18:48:10 dansmith yeah I thought it was in the prelude that I read this morning
18:48:28 cdent I think we should require 1.10
18:48:34 dansmith there's something that specifically says "be sure to upgrade placement first"
18:48:41 dansmith I really thought that was the prelude
18:48:48 mriedem that's https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes
18:49:04 mriedem the language there is more clear, the behavior in the scheduler code is not
18:49:09 dansmith ah it's thjs: https://review.openstack.org/#/c/491900/2/doc/source/user/placement.rst
18:49:17 dansmith pike placement upgrade notes
18:49:32 cdent if we haven’t got 1.10, we haven’t got alloc_candidates and we’ve not forced things forward, let’s do that
18:49:43 mriedem i'll update nova-status quick
18:50:23 cdent If nobody gets to it tonight, I can and will look into the stuff where tempest is spewing ERROR tomorrow morning
18:50:38 mriedem it's because we have a small flavor, 1 VCPU
18:50:39 dansmith mriedem: did you say you think the errors during single-host migration come from the bottom patch?
18:50:53 mriedem when we resize on the same host, that subtracts the flavor from itself, resulting in 0 VCPU
18:50:55 mriedem which is an invalid allocation
18:51:03 mriedem dansmith: from the middle patch
18:51:07 mriedem the pep8 is on the bottom patch
18:51:08 dansmith okay I was going to say..
18:51:21 dansmith ah
18:51:44 mriedem so we'd need some additional floor type logic
18:51:46 cdent isn’t the allocation supposed to be doubling to 2 already? where’s that getting lost?
18:51:48 mriedem such that an allocation can't dip below 0
18:52:08 mriedem cdent: good question
18:52:24 mriedem Instance 6fa7b953-c1fe-4520-8564-aeba8e90aece has resources on 1 compute nodes
18:52:28 mriedem Original resources from same-host allocation: {u'VCPU': 1, u'MEMORY_MB': 64}
18:52:33 mriedem Subtracting old resources from same-host allocation: {u'VCPU': 0, u'MEMORY_MB': 0, 'DISK_GB': 0}
18:52:36 mriedem then kablammo
18:52:39 dansmith mriedem: we're doing the -1 merge which should be subtracting us from the doubled allocation
18:53:20 mriedem yeah the allocation doesn't appear to be doubled
18:53:21 mriedem is the problem
18:53:44 dansmith so
18:53:45 dansmith in that patch,
18:53:52 dansmith we're still healing on both sides with reckless abandon
18:53:56 dansmith does it show up in thelast patch?
18:53:59 dansmith that's where we stop doing that
18:54:09 dansmith although that patch has a typo in it, so probably isn't running properly anyway
18:54:23 dansmith er I mean, had until I updated it a few minutes ago
18:56:50 mriedem umm
18:56:51 mriedem http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_56_298572
18:56:58 mriedem Sending updated allocation [{'resource_provider': {'uuid': '97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 0, u'MEMORY_MB': 0, 'DISK_GB': 0}}] for instance 6fa7b953-c1fe-4520-8564-aeba8e90aece after removing resources for 97ce076a-2644-465b-95dc-fc0674152976. {{(pid=19567) remove_provider_from_instance_allocation /opt/stack/new/nova/nova/scheduler/client/report.py:1091}}
18:56:58 cdent ah, so doubling gets cleaned up too soon (until later in the series)?
18:57:12 mriedem oh nvm
18:57:17 mriedem that's known, i was thinking that was the bottom change
18:59:17 mriedem the current_allocs get logged but they don't have 2 VCPU
18:59:20 jaypipes dansmith: are you actively making changes to these patches? if not, I can address mriedem's comments.
18:59:26 mriedem so the doubling isn't happening it appears
19:00:43 mriedem this is the resize in the scheduler http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-sch.txt.gz#_Aug_09_16_27_37_363471
19:00:53 mriedem req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4
19:01:12 openstackgerrit Merged openstack/nova master: doc: provide more details on scheduling with placement https://review.openstack.org/491900
19:01:49 mriedem req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4
19:01:51 mriedem oops
19:01:56 mriedem Aug 09 16:27:37.633395 ubuntu-xenial-internap-mtl01-10347573 nova-scheduler[17280]: DEBUG nova.scheduler.client.report [None req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 tempest-MigrationsAdminTest-402050297 tempest-MigrationsAdminTest-402050297] Doubling-up allocation request for move operation. {{(pid=17280) _move_operation_alloc_request /opt/stack/new/nova/nova/scheduler/client/report.py:162}}
19:02:01 mriedem it says it's double stuffing it

Earlier   Later