| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 14:06:28 | mriedem | the chance scheduler shouldn't even be asking placement for anything | |
| 14:06:37 | mriedem | it sets a flag for that (or doesn't set it rather) | |
| 14:06:40 | dansmith | I don't know what outsource means in this context | |
| 14:06:46 | bauzas | dansmith: out-of-tree | |
| 14:06:51 | dansmith | mriedem: well, the problem is it won't claim | |
| 14:07:02 | dansmith | mriedem: so if we have it in tree it needs to at least do that | |
| 14:07:14 | mriedem | dansmith: claim where? the scheduler? | |
| 14:07:20 | mriedem | neither does the caching scheduler | |
| 14:07:25 | dansmith | mriedem: right, same issue | |
| 14:07:37 | mriedem | why is that a problem? that was intentional | |
| 14:07:53 | dansmith | mriedem: because if we remove the claiming on the compute, we have no way of knowing if something will fit before sending it | |
| 14:08:06 | mriedem | we aren't removing claims in the compute in pike | |
| 14:08:06 | dansmith | which isn't so much a problem for chance since chance doesn't care about that at all, | |
| 14:08:33 | dansmith | mriedem: I think we're talking about the long-term strategy here | |
| 14:08:35 | bauzas | I really 'd love to have a whiteboard | |
| 14:08:47 | mriedem | dansmith: ok - not sure why it's a problem or fire drill for this week | |
| 14:08:49 | mriedem | maybe it's not | |
| 14:08:55 | bauzas | I just see drivers as a black box from a placement perspective | |
| 14:08:55 | mriedem | sounds like a distraction | |
| 14:08:58 | dansmith | mriedem: I don't think it is | |
| 14:09:17 | bauzas | #1 scheduler manager passes a list of candidates to the driver | |
| 14:09:31 | bauzas | #2 driver does its black magic to find the perfect candidate | |
| 14:09:46 | bauzas | #3 scheduler manager would claim the allocation | |
| 14:09:58 | bauzas | that is the loved interface I'd like | |
| 14:10:32 | bauzas | so, one day, we could just do better things than just looping over the list of instances which loops over the list of hosts | |
| 14:12:06 | dansmith | filter scheduler with no filters or weights configured is O(n) right? | |
| 14:12:28 | dansmith | and is effectively chance, but without the pathological sending of instances to full computes | |
| 14:12:40 | openstackgerrit | Merged openstack/nova-specs master: Amend spec for "Allow custom resource classes in flavor extra specs" https://review.openstack.org/481748 | |
| 14:13:39 | bauzas | dansmith: in theory, we also loop over each filter | |
| 14:13:59 | bauzas | dansmith: but since we have far less filters than hosts, I'm just keeping it O(n2) | |
| 14:14:13 | dansmith | bauzas: well (a) that's O(1) and (b) that's irrelevant if there are no filters :) | |
| 14:14:15 | bauzas | dansmith: the point is, what would be the interest of filter scheduler if we don't filter anything ? | |
| 14:14:33 | dansmith | bauzas: there is zero interest of the chance scheduler, so I'm not sure what your point is :) | |
| 14:14:40 | bauzas | dansmith: LOL | |
| 14:15:26 | dansmith | filter scheduler with no filters gives you at least selection by "would fit at all" in constant time (per instance) with proper claiming and random selection of a host within the possible set | |
| 14:15:58 | bauzas | dansmith: I see your point | |
| 14:16:13 | bauzas | if you want to deprecate chance, that would the solution, I agree | |
| 14:16:15 | dansmith | I'd be willing to bet a number of people assume chance does at least "would fit" and then once they read more and realize it doesn't, move on to something sane as they work through their POC | |
| 14:16:18 | bauzas | would be | |
| 14:20:02 | gibi | 20% of func tests are failing after change chance_scheduler to filter_scheduler | |
| 14:20:22 | gibi | let's see if it is something that easy to fix | |
| 14:21:02 | cdent | gibi++ | |
| 14:26:59 | mriedem | mikal: http://eavesdrop.openstack.org/meetings/nova/2017/nova.2017-07-27-14.00.log.html#l-67 | |
| 14:27:09 | mriedem | mikal: if you'd read the meeting log | |
| 14:37:54 | edleafe | mriedem: you'd be the longest-serving PTL that nobody voted for | |
| 14:38:18 | cdent | let’s get mdbooth to run for PTL | |
| 14:38:28 | cdent | for variety | |
| 14:39:06 | cdent | mriedem: I went ahead and backported that at-least-one allocation fix: v | |
| 14:39:07 | cdent | https://review.openstack.org/#/c/491487/ | |
| 14:40:58 | bauzas | cdent: +2d | |
| 14:41:05 | cdent | thanks | |
| 14:48:49 | jaypipes | mriedem: k, commented on https://review.openstack.org/#/c/491098/1 | |
| 14:54:39 | mriedem | jaypipes: thanks, replied | |
| 14:55:49 | lennyb | dansmith: pls take a look #link https://bugs.launchpad.net/nova/+bug/1708920 | |
| 14:55:49 | openstack | Launchpad bug 1708920 in OpenStack Compute (nova) "Cold migration fails" [Undecided,New] | |
| 14:56:28 | jaypipes | mriedem: no, Matt, we *are* removing the call to put_allocations() in Pike. | |
| 14:56:57 | jaypipes | mriedem: that's what this does. https://review.openstack.org/#/c/491012/4/nova/compute/resource_tracker.py | |
| 14:57:30 | jaypipes | mriedem: what it doesn't do is stop the heal allocations thing from happening when ocata computes are in the mix. | |
| 14:57:48 | jaypipes | mriedem: because we'll need to continually correct the mistake that the ocata compute will make. | |
| 14:58:04 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove dead function, variable https://review.openstack.org/491512 | |
| 14:59:00 | jaypipes | mriedem: so, w.r.t. the problem of shared storage, I think instead of adding the type of code you added in your patch, maybe we should just wait for Queens when we can assume no compute hosts are attempting to heal allocations in the compute node. | |
| 14:59:00 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove some single use function from utils https://review.openstack.org/491513 | |
| 14:59:12 | jaypipes | mriedem: I'd like dansmith's thoughts on that, though. | |
| 14:59:40 | mriedem | jaypipes: i don't think we can claim support for shared storage at this point regardless | |
| 14:59:45 | mriedem | we're basically implementing a feature now | |
| 14:59:56 | mriedem | which is why i un-completed the shared storage bp | |
| 15:00:07 | dansmith | lennyb: that's not a lost connection to the cell database, that's a lack of configuration of some node for the api database. are you running conductors on your second (non-allinone) node? | |
| 15:00:25 | jaypipes | mriedem: another thing I've been thinking of is instead of further mucking with the resource tracker and/or making the scheduler report client that we do a nova-status or nova-manage audit function that cleans up allocations based on the cell DB states. | |
| 15:01:28 | jaypipes | mriedem: and once done, we have all computes on service version 22 (or whatever) and after that point, we remove all allocation stuff from the compute node entirely (other than the aforementioned delete allocations pieces) | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove files from 'tools' https://review.openstack.org/491444 | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove "vif" script https://review.openstack.org/491443 | |
| 15:03:05 | lennyb | dansmith, thanks, I will recheck it. | |
| 15:03:08 | jaypipes | mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :( | |
| 15:03:09 | mriedem | jaypipes: the service version won't change until the actual code is upgraded though | |
| 15:03:18 | mriedem | jaypipes: i agree with that | |
| 15:03:38 | dansmith | is it alternate-ego mondays? why is jaypipes being "Frank" ? | |
| 15:03:45 | jaypipes | heh | |
| 15:05:04 | jaypipes | mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information. | |
| 15:06:11 | jaypipes | mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client. | |
| 15:06:14 | dansmith | jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing? | |
| 15:06:19 | jaypipes | mriedem: choose your poison I think? | |
| 15:06:27 | jaypipes | dansmith: yeah. | |
| 15:06:29 | sdague | mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work | |
| 15:06:42 | dansmith | jaypipes: that would be really unfortunate, IMHO | |
| 15:06:49 | dansmith | jaypipes: some people take a looong time to do that upgrade | |
| 15:06:50 | jaypipes | dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again. | |
| 15:08:11 | mriedem | sdague: i can look in a bit | |
| 15:09:13 | jaypipes | dansmith: do you think we should even attempt to fix the shared storage reporting in Pike? | |
| 15:09:27 | jaypipes | dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1 | |
| 15:09:39 | dansmith | jaypipes: well, tbh, I thought that was out the window a while ago | |
| 15:09:51 | dansmith | the computes have no idea which shared storage uuid a given thing is using anyway right? | |
| 15:10:06 | jaypipes | dansmith: right. and ocatas are just going to pummel the allocations either way. | |
| 15:11:22 | mriedem | jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes | |
| 15:11:49 | jaypipes | dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike. | |
| 15:11:55 | mriedem | kind of weird to say, "this feature doesn't yet work" in a known issues section though | |
| 15:12:07 | mriedem | i don't know of omission is better than clearly saying it's not supported | |
| 15:12:23 | mriedem | i'd rather be clear personally | |
| 15:12:28 | mriedem | of what is supported and what's not | |
| 15:12:30 | jaypipes | mriedem: yeah, agreed | |
| 15:13:11 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517 | |