| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 14:01:26 | dansmith | heh | |
| 14:01:38 | jaypipes | bauzas: err, you and I have different ideas of efficiency I guess. | |
| 14:01:46 | bauzas | yeah, I know, it's orthogonal to what we do | |
| 14:02:03 | jaypipes | bauzas, dansmith: sched meeting in meeting-alt | |
| 14:02:03 | bauzas | I'm just describing *why* people use it | |
| 14:03:07 | dansmith | bauzas: can you find any reference to support your claim that some people use chance successfully and recommend it in anything other than toy environments? | |
| 14:03:31 | bauzas | dansmith: unfortunately, only hallways talks | |
| 14:03:39 | dansmith | the only such mention on the first page of google is a RAX presentation from 2013 that says "filter scheduler is the only legit one" | |
| 14:04:22 | bauzas | anyway, I'm not particularly attached to chance_scheduler | |
| 14:04:29 | bauzas | it's even broken if you use ironic | |
| 14:04:33 | stephenfin | cdent: https://review.openstack.org/#/c/490952/ is pretty much done, yeah. All that needs to be checked if if I've missed anything from the import process | |
| 14:04:59 | dansmith | bauzas: oh okay, seems like you're arguing to keep maintaining it | |
| 14:05:01 | bauzas | I'm just trying to explain that I want to make sure we keep a clear separation in between what placement gives to drivers, and how driver work | |
| 14:05:28 | bauzas | dansmith: sorry, I'm unclear, I'm just giving food for thoughts about why we have it in-tree | |
| 14:05:47 | bauzas | we could outsource chance_scheduler | |
| 14:06:08 | bauzas | if we make clear the fact that it gets a list of RPs provided by Placement | |
| 14:06:20 | bauzas | it == the scheduler driver interface | |
| 14:06:28 | mriedem | the chance scheduler shouldn't even be asking placement for anything | |
| 14:06:37 | mriedem | it sets a flag for that (or doesn't set it rather) | |
| 14:06:40 | dansmith | I don't know what outsource means in this context | |
| 14:06:46 | bauzas | dansmith: out-of-tree | |
| 14:06:51 | dansmith | mriedem: well, the problem is it won't claim | |
| 14:07:02 | dansmith | mriedem: so if we have it in tree it needs to at least do that | |
| 14:07:14 | mriedem | dansmith: claim where? the scheduler? | |
| 14:07:20 | mriedem | neither does the caching scheduler | |
| 14:07:25 | dansmith | mriedem: right, same issue | |
| 14:07:37 | mriedem | why is that a problem? that was intentional | |
| 14:07:53 | dansmith | mriedem: because if we remove the claiming on the compute, we have no way of knowing if something will fit before sending it | |
| 14:08:06 | mriedem | we aren't removing claims in the compute in pike | |
| 14:08:06 | dansmith | which isn't so much a problem for chance since chance doesn't care about that at all, | |
| 14:08:33 | dansmith | mriedem: I think we're talking about the long-term strategy here | |
| 14:08:35 | bauzas | I really 'd love to have a whiteboard | |
| 14:08:47 | mriedem | dansmith: ok - not sure why it's a problem or fire drill for this week | |
| 14:08:49 | mriedem | maybe it's not | |
| 14:08:55 | bauzas | I just see drivers as a black box from a placement perspective | |
| 14:08:55 | mriedem | sounds like a distraction | |
| 14:08:58 | dansmith | mriedem: I don't think it is | |
| 14:09:17 | bauzas | #1 scheduler manager passes a list of candidates to the driver | |
| 14:09:31 | bauzas | #2 driver does its black magic to find the perfect candidate | |
| 14:09:46 | bauzas | #3 scheduler manager would claim the allocation | |
| 14:09:58 | bauzas | that is the loved interface I'd like | |
| 14:10:32 | bauzas | so, one day, we could just do better things than just looping over the list of instances which loops over the list of hosts | |
| 14:12:06 | dansmith | filter scheduler with no filters or weights configured is O(n) right? | |
| 14:12:28 | dansmith | and is effectively chance, but without the pathological sending of instances to full computes | |
| 14:12:40 | openstackgerrit | Merged openstack/nova-specs master: Amend spec for "Allow custom resource classes in flavor extra specs" https://review.openstack.org/481748 | |
| 14:13:39 | bauzas | dansmith: in theory, we also loop over each filter | |
| 14:13:59 | bauzas | dansmith: but since we have far less filters than hosts, I'm just keeping it O(n2) | |
| 14:14:13 | dansmith | bauzas: well (a) that's O(1) and (b) that's irrelevant if there are no filters :) | |
| 14:14:15 | bauzas | dansmith: the point is, what would be the interest of filter scheduler if we don't filter anything ? | |
| 14:14:33 | dansmith | bauzas: there is zero interest of the chance scheduler, so I'm not sure what your point is :) | |
| 14:14:40 | bauzas | dansmith: LOL | |
| 14:15:26 | dansmith | filter scheduler with no filters gives you at least selection by "would fit at all" in constant time (per instance) with proper claiming and random selection of a host within the possible set | |
| 14:15:58 | bauzas | dansmith: I see your point | |
| 14:16:13 | bauzas | if you want to deprecate chance, that would the solution, I agree | |
| 14:16:15 | dansmith | I'd be willing to bet a number of people assume chance does at least "would fit" and then once they read more and realize it doesn't, move on to something sane as they work through their POC | |
| 14:16:18 | bauzas | would be | |
| 14:20:02 | gibi | 20% of func tests are failing after change chance_scheduler to filter_scheduler | |
| 14:20:22 | gibi | let's see if it is something that easy to fix | |
| 14:21:02 | cdent | gibi++ | |
| 14:26:59 | mriedem | mikal: http://eavesdrop.openstack.org/meetings/nova/2017/nova.2017-07-27-14.00.log.html#l-67 | |
| 14:27:09 | mriedem | mikal: if you'd read the meeting log | |
| 14:37:54 | edleafe | mriedem: you'd be the longest-serving PTL that nobody voted for | |
| 14:38:18 | cdent | let’s get mdbooth to run for PTL | |
| 14:38:28 | cdent | for variety | |
| 14:39:06 | cdent | mriedem: I went ahead and backported that at-least-one allocation fix: v | |
| 14:39:07 | cdent | https://review.openstack.org/#/c/491487/ | |
| 14:40:58 | bauzas | cdent: +2d | |
| 14:41:05 | cdent | thanks | |
| 14:48:49 | jaypipes | mriedem: k, commented on https://review.openstack.org/#/c/491098/1 | |
| 14:54:39 | mriedem | jaypipes: thanks, replied | |
| 14:55:49 | lennyb | dansmith: pls take a look #link https://bugs.launchpad.net/nova/+bug/1708920 | |
| 14:55:49 | openstack | Launchpad bug 1708920 in OpenStack Compute (nova) "Cold migration fails" [Undecided,New] | |
| 14:56:28 | jaypipes | mriedem: no, Matt, we *are* removing the call to put_allocations() in Pike. | |
| 14:56:57 | jaypipes | mriedem: that's what this does. https://review.openstack.org/#/c/491012/4/nova/compute/resource_tracker.py | |
| 14:57:30 | jaypipes | mriedem: what it doesn't do is stop the heal allocations thing from happening when ocata computes are in the mix. | |
| 14:57:48 | jaypipes | mriedem: because we'll need to continually correct the mistake that the ocata compute will make. | |
| 14:58:04 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove dead function, variable https://review.openstack.org/491512 | |
| 14:59:00 | jaypipes | mriedem: so, w.r.t. the problem of shared storage, I think instead of adding the type of code you added in your patch, maybe we should just wait for Queens when we can assume no compute hosts are attempting to heal allocations in the compute node. | |
| 14:59:00 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove some single use function from utils https://review.openstack.org/491513 | |
| 14:59:12 | jaypipes | mriedem: I'd like dansmith's thoughts on that, though. | |
| 14:59:40 | mriedem | jaypipes: i don't think we can claim support for shared storage at this point regardless | |
| 14:59:45 | mriedem | we're basically implementing a feature now | |
| 14:59:56 | mriedem | which is why i un-completed the shared storage bp | |
| 15:00:07 | dansmith | lennyb: that's not a lost connection to the cell database, that's a lack of configuration of some node for the api database. are you running conductors on your second (non-allinone) node? | |
| 15:00:25 | jaypipes | mriedem: another thing I've been thinking of is instead of further mucking with the resource tracker and/or making the scheduler report client that we do a nova-status or nova-manage audit function that cleans up allocations based on the cell DB states. | |
| 15:01:28 | jaypipes | mriedem: and once done, we have all computes on service version 22 (or whatever) and after that point, we remove all allocation stuff from the compute node entirely (other than the aforementioned delete allocations pieces) | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove files from 'tools' https://review.openstack.org/491444 | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove "vif" script https://review.openstack.org/491443 | |
| 15:03:05 | lennyb | dansmith, thanks, I will recheck it. | |
| 15:03:08 | jaypipes | mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :( | |
| 15:03:09 | mriedem | jaypipes: the service version won't change until the actual code is upgraded though | |
| 15:03:18 | mriedem | jaypipes: i agree with that | |
| 15:03:38 | dansmith | is it alternate-ego mondays? why is jaypipes being "Frank" ? | |
| 15:03:45 | jaypipes | heh | |
| 15:05:04 | jaypipes | mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information. | |
| 15:06:11 | jaypipes | mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client. | |
| 15:06:14 | dansmith | jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing? | |
| 15:06:19 | jaypipes | mriedem: choose your poison I think? | |
| 15:06:27 | jaypipes | dansmith: yeah. | |
| 15:06:29 | sdague | mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work | |