Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-07
14:06:08 bauzas if we make clear the fact that it gets a list of RPs provided by Placement
14:06:20 bauzas it == the scheduler driver interface
14:06:28 mriedem the chance scheduler shouldn't even be asking placement for anything
14:06:37 mriedem it sets a flag for that (or doesn't set it rather)
14:06:40 dansmith I don't know what outsource means in this context
14:06:46 bauzas dansmith: out-of-tree
14:06:51 dansmith mriedem: well, the problem is it won't claim
14:07:02 dansmith mriedem: so if we have it in tree it needs to at least do that
14:07:14 mriedem dansmith: claim where? the scheduler?
14:07:20 mriedem neither does the caching scheduler
14:07:25 dansmith mriedem: right, same issue
14:07:37 mriedem why is that a problem? that was intentional
14:07:53 dansmith mriedem: because if we remove the claiming on the compute, we have no way of knowing if something will fit before sending it
14:08:06 dansmith which isn't so much a problem for chance since chance doesn't care about that at all,
14:08:06 mriedem we aren't removing claims in the compute in pike
14:08:33 dansmith mriedem: I think we're talking about the long-term strategy here
14:08:35 bauzas I really 'd love to have a whiteboard
14:08:47 mriedem dansmith: ok - not sure why it's a problem or fire drill for this week
14:08:49 mriedem maybe it's not
14:08:55 mriedem sounds like a distraction
14:08:55 bauzas I just see drivers as a black box from a placement perspective
14:08:58 dansmith mriedem: I don't think it is
14:09:17 bauzas #1 scheduler manager passes a list of candidates to the driver
14:09:31 bauzas #2 driver does its black magic to find the perfect candidate
14:09:46 bauzas #3 scheduler manager would claim the allocation
14:09:58 bauzas that is the loved interface I'd like
14:10:32 bauzas so, one day, we could just do better things than just looping over the list of instances which loops over the list of hosts
14:12:06 dansmith filter scheduler with no filters or weights configured is O(n) right?
14:12:28 dansmith and is effectively chance, but without the pathological sending of instances to full computes
14:12:40 openstackgerrit Merged openstack/nova-specs master: Amend spec for "Allow custom resource classes in flavor extra specs" https://review.openstack.org/481748
14:13:39 bauzas dansmith: in theory, we also loop over each filter
14:13:59 bauzas dansmith: but since we have far less filters than hosts, I'm just keeping it O(n2)
14:14:13 dansmith bauzas: well (a) that's O(1) and (b) that's irrelevant if there are no filters :)
14:14:15 bauzas dansmith: the point is, what would be the interest of filter scheduler if we don't filter anything ?
14:14:33 dansmith bauzas: there is zero interest of the chance scheduler, so I'm not sure what your point is :)
14:14:40 bauzas dansmith: LOL
14:15:26 dansmith filter scheduler with no filters gives you at least selection by "would fit at all" in constant time (per instance) with proper claiming and random selection of a host within the possible set
14:15:58 bauzas dansmith: I see your point
14:16:13 bauzas if you want to deprecate chance, that would the solution, I agree
14:16:15 dansmith I'd be willing to bet a number of people assume chance does at least "would fit" and then once they read more and realize it doesn't, move on to something sane as they work through their POC
14:16:18 bauzas would be
14:20:02 gibi 20% of func tests are failing after change chance_scheduler to filter_scheduler
14:20:22 gibi let's see if it is something that easy to fix
14:21:02 cdent gibi++
14:26:59 mriedem mikal: http://eavesdrop.openstack.org/meetings/nova/2017/nova.2017-07-27-14.00.log.html#l-67
14:27:09 mriedem mikal: if you'd read the meeting log
14:37:54 edleafe mriedem: you'd be the longest-serving PTL that nobody voted for
14:38:18 cdent let’s get mdbooth to run for PTL
14:38:28 cdent for variety
14:39:06 cdent mriedem: I went ahead and backported that at-least-one allocation fix: v
14:39:07 cdent https://review.openstack.org/#/c/491487/
14:40:58 bauzas cdent: +2d
14:41:05 cdent thanks
14:48:49 jaypipes mriedem: k, commented on https://review.openstack.org/#/c/491098/1
14:54:39 mriedem jaypipes: thanks, replied
14:55:49 openstack Launchpad bug 1708920 in OpenStack Compute (nova) "Cold migration fails" [Undecided,New]
14:55:49 lennyb dansmith: pls take a look #link https://bugs.launchpad.net/nova/+bug/1708920
14:56:28 jaypipes mriedem: no, Matt, we *are* removing the call to put_allocations() in Pike.
14:56:57 jaypipes mriedem: that's what this does. https://review.openstack.org/#/c/491012/4/nova/compute/resource_tracker.py
14:57:30 jaypipes mriedem: what it doesn't do is stop the heal allocations thing from happening when ocata computes are in the mix.
14:57:48 jaypipes mriedem: because we'll need to continually correct the mistake that the ocata compute will make.
14:58:04 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove dead function, variable https://review.openstack.org/491512
14:59:00 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove some single use function from utils https://review.openstack.org/491513
14:59:00 jaypipes mriedem: so, w.r.t. the problem of shared storage, I think instead of adding the type of code you added in your patch, maybe we should just wait for Queens when we can assume no compute hosts are attempting to heal allocations in the compute node.
14:59:12 jaypipes mriedem: I'd like dansmith's thoughts on that, though.
14:59:40 mriedem jaypipes: i don't think we can claim support for shared storage at this point regardless
14:59:45 mriedem we're basically implementing a feature now
14:59:56 mriedem which is why i un-completed the shared storage bp
15:00:07 dansmith lennyb: that's not a lost connection to the cell database, that's a lack of configuration of some node for the api database. are you running conductors on your second (non-allinone) node?
15:00:25 jaypipes mriedem: another thing I've been thinking of is instead of further mucking with the resource tracker and/or making the scheduler report client that we do a nova-status or nova-manage audit function that cleans up allocations based on the cell DB states.
15:01:28 jaypipes mriedem: and once done, we have all computes on service version 22 (or whatever) and after that point, we remove all allocation stuff from the compute node entirely (other than the aforementioned delete allocations pieces)
15:02:37 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove "vif" script https://review.openstack.org/491443
15:02:37 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove files from 'tools' https://review.openstack.org/491444
15:03:05 lennyb dansmith, thanks, I will recheck it.
15:03:08 jaypipes mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :(
15:03:09 mriedem jaypipes: the service version won't change until the actual code is upgraded though
15:03:18 mriedem jaypipes: i agree with that
15:03:38 dansmith is it alternate-ego mondays? why is jaypipes being "Frank" ?
15:03:45 jaypipes heh
15:05:04 jaypipes mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information.
15:06:11 jaypipes mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client.
15:06:14 dansmith jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing?
15:06:19 jaypipes mriedem: choose your poison I think?
15:06:27 jaypipes dansmith: yeah.
15:06:29 sdague mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work
15:06:42 dansmith jaypipes: that would be really unfortunate, IMHO
15:06:49 dansmith jaypipes: some people take a looong time to do that upgrade
15:06:50 jaypipes dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again.
15:08:11 mriedem sdague: i can look in a bit
15:09:13 jaypipes dansmith: do you think we should even attempt to fix the shared storage reporting in Pike?
15:09:27 jaypipes dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1
15:09:39 dansmith jaypipes: well, tbh, I thought that was out the window a while ago
15:09:51 dansmith the computes have no idea which shared storage uuid a given thing is using anyway right?
15:10:06 jaypipes dansmith: right. and ocatas are just going to pummel the allocations either way.
15:11:22 mriedem jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes
15:11:49 jaypipes dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike.
15:11:55 mriedem kind of weird to say, "this feature doesn't yet work" in a known issues section though
15:12:07 mriedem i don't know of omission is better than clearly saying it's not supported
15:12:23 mriedem i'd rather be clear personally
15:12:28 mriedem of what is supported and what's not

Earlier   Later