Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
17:58:58 dansmith cfriesen: it should end up being much better
17:59:24 cfriesen mnaser: if you've got 0/1 reserved then I think arguably there should be a set with just 16 and another set with just 17
18:00:03 cfriesen it should be noted though that the performance of 16/17 is going to be variable, depending on what load is on 0/1
18:00:40 cfriesen dansmith: better? isn't CachingScheduler in memory for the most part?
18:00:42 mnaser of course, that's a given
18:01:38 dansmith cfriesen: we've been discussing this in this channel for like I year, so I hope I don't have to summarize all of it, but if you count the ability to multi-thread the scheduler with placement, and zero reschedules, yes, it should be much better
18:01:54 dansmith s/I/a
18:02:02 mriedem roman numeral I
18:02:05 mriedem you meant
18:02:11 dansmith heh, yeah
18:02:11 sdague dansmith: I did not push the deprecation of the other schedulers
18:02:17 sdague dansmith: if that's a patch that is needed, I could do it this afternoon
18:02:18 dansmith sdague: I just did
18:02:23 mriedem https://review.openstack.org/#/c/492210/1
18:02:25 sdague dansmith: ok, coolio
18:02:35 cfriesen dansmith: ah, you mean overall scheduler load. got it. (I was thinking time to perform a single schedule operation.)
18:02:42 mnaser https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L5632-L5633 -- bingo.. now to see why the decision of filtering out singles is there
18:02:51 mriedem cfriesen: it's also that
18:03:00 mriedem time to schedule should be faster when you don't have reschedules
18:03:02 mriedem b/c of bad decisions
18:03:41 mriedem cfriesen: your team seems to do some performance testing, it would be cool if your guy(s) could benchmark and compare the two now in pike
18:04:00 mriedem like single caching scheduler vs multiple filter schedulers
18:04:00 dansmith if you really care about the milliseconds requried to place something, you should be able to do better with a bunch of filterschedulers than one in-memory caching scheduler I would think
18:04:08 cdent cfriesen: it may be the case, but I’m not sure that anyone has proven it yet, that single schedule operations will speed up because placement will help limit the initial set of available candidates
18:04:51 dansmith cdent: cachingscheduler only refreshes the list when it runs out of spots I think, so it varies, but it also is using very outdated information which means fast but poor decisions
18:05:09 cfriesen I'd expect the FilterScheduler to speed up due to Placement. Wasn't sure that it'd be faster than CachingScheduler. Anyways, we've been using FilterScheduler since we care about accurate resource tracking.
18:05:16 mriedem caching scheduler is also on a periodic
18:05:45 cfriesen (We modified stuff like resize and migration to update the resource tracking immediately rather than wait for the resource audit.)
18:05:49 dansmith caching scheduler is filter scheduler, with cached host list data
18:06:18 mriedem dansmith: so this resize/rt series has a pep8 failure in the bottom change,
18:06:31 mriedem dansmith: but i guess hold off on fixing that so i can go through them now
18:06:33 mriedem the last 2 i mean
18:06:41 dansmith okay
18:06:46 mriedem dansmith: it's T-~3 hours to 0 dan time yeah?
18:06:53 dansmith check queue is 7h long
18:07:26 dansmith mriedem: I think you mean T-3 hours until free-of-dan time
18:07:27 dansmith but yeah
18:07:55 dansmith I assume everyone brought liquor and cake to work today for the post-2pm PDT party
18:08:14 mriedem brought? i think it's just generally available
18:08:47 dansmith I was trying to act like it was the 90s
18:13:40 mnaser cfriesen thanks for all your feedback, i think the best solution for now is going to be use 0,16 to have predictable performance for the host and assign hugepages with a slight bias towards node1 because it has 2 more cores so things can get properly packed
18:23:37 sdague dansmith: yeh, we're down to 800 nodes in ci
18:23:54 sdague which is unfortunately not sufficient
18:24:16 dansmith well, it was fun while it lasted
18:25:50 sdague if anyone wants to do doc reviews to hopefully complete the import - https://review.openstack.org/#/q/status:open+project:openstack/nova+branch:master+topic:bp/doc-migration
18:30:06 cfriesen mnaser: sorry you hit a bug. :) we should be able to get it sorted out once sfinucan comes back from holidays
18:31:07 cfriesen mnaser: ping me with the bug number once you file it please
18:41:12 mriedem dansmith: jaypipes: cdent: edleafe: problems https://review.openstack.org/#/c/488510/
18:42:23 dansmith mriedem: I +2d before we had jenkins runs of the latest version
18:42:27 dansmith and it was passing (tempest) before
18:42:35 mriedem it's not that it's not passing tempest
18:42:41 mriedem it's just spewing ERRORs
18:42:46 jaypipes hey guys, just got back.
18:43:33 jaypipes hmm, yes, the dreaded ./nova/compute/resource_tracker.py:1112:17: N352 LOG.warn is deprecated, please use LOG.warning!
18:44:12 mriedem heh i even pointed that out
18:44:17 mriedem it's in the bottom change btw
18:44:23 mriedem so the entire series needs to be rebaesd
18:44:24 mriedem *rebased
18:44:32 mriedem but i have issues with the 2nd change in the series
18:45:02 jaypipes yeah, I see that
18:45:14 dansmith mriedem: we have the placement upgrade before nova thing required in the prelude of the renos
18:45:21 dansmith which should be enough to merge this, IMHO
18:45:44 jaypipes "prelude of the renos" sounds like some classical music composition.
18:46:25 mriedem dansmith: where?
18:46:26 dansmith I probably shouldn't be reviewing this anyway since I was elbow deep in it myself
18:46:28 mriedem i know it's in the release notes
18:46:35 mriedem but it's not in the prelude
18:46:53 mriedem https://docs.openstack.org/releasenotes/nova/unreleased.html#id1
18:46:58 mriedem idates.
18:46:58 mriedem A new 1.10 API microversion is added to the Placement REST API. This microversion adds support for the GET /allocation_candidates resource endpoint. This endpoint returns information about possible allocation requests that callers can make which meet a set of resource constraints supplied as query string parameters. Also returned is some inventory and capacity information for the resource providers involved in the allocation
18:47:19 mriedem https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes
18:47:19 mriedem wrong one
18:47:25 mriedem The scheduler now requests allocation candidates from the Placement service during scheduling. The allocation candidates information was introduced in the Placement API 1.10 microversion, so you should upgrade the placement service before the Nova scheduler service so that the scheduler can take advantage of the allocation candidate information.
18:47:42 mriedem the code totally falls back though if 1.10 isn't available
18:48:08 mriedem if we want to make 1.10 required, that's fine, but we need to also update nova-status, which i could do lickety split
18:48:10 dansmith yeah I thought it was in the prelude that I read this morning
18:48:28 cdent I think we should require 1.10
18:48:34 dansmith there's something that specifically says "be sure to upgrade placement first"
18:48:41 dansmith I really thought that was the prelude
18:48:48 mriedem that's https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes
18:49:04 mriedem the language there is more clear, the behavior in the scheduler code is not
18:49:09 dansmith ah it's thjs: https://review.openstack.org/#/c/491900/2/doc/source/user/placement.rst
18:49:17 dansmith pike placement upgrade notes
18:49:32 cdent if we haven’t got 1.10, we haven’t got alloc_candidates and we’ve not forced things forward, let’s do that
18:49:43 mriedem i'll update nova-status quick
18:50:23 cdent If nobody gets to it tonight, I can and will look into the stuff where tempest is spewing ERROR tomorrow morning
18:50:38 mriedem it's because we have a small flavor, 1 VCPU
18:50:39 dansmith mriedem: did you say you think the errors during single-host migration come from the bottom patch?
18:50:53 mriedem when we resize on the same host, that subtracts the flavor from itself, resulting in 0 VCPU
18:50:55 mriedem which is an invalid allocation
18:51:03 mriedem dansmith: from the middle patch
18:51:07 mriedem the pep8 is on the bottom patch
18:51:08 dansmith okay I was going to say..
18:51:21 dansmith ah
18:51:44 mriedem so we'd need some additional floor type logic
18:51:46 cdent isn’t the allocation supposed to be doubling to 2 already? where’s that getting lost?
18:51:48 mriedem such that an allocation can't dip below 0
18:52:08 mriedem cdent: good question
18:52:24 mriedem Instance 6fa7b953-c1fe-4520-8564-aeba8e90aece has resources on 1 compute nodes
18:52:28 mriedem Original resources from same-host allocation: {u'VCPU': 1, u'MEMORY_MB': 64}

Earlier   Later