| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 15:46:02 | mriedem | we removed the ability for the compute manager to be pluggable, or do you mean the virt driver? | |
| 15:46:12 | dansmith | that's my point, | |
| 15:46:31 | dansmith | there's basically zero chance you've written your own compute manager that does everything properly | |
| 15:46:55 | dansmith | and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today | |
| 15:47:08 | openstackgerrit | Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469 | |
| 15:47:10 | mriedem | well, you just didn't upgrade from kilo L) | |
| 15:47:11 | mriedem | :) | |
| 15:47:19 | mriedem | which is like 50% of deployments probably | |
| 15:47:24 | bauzas | mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters | |
| 15:47:34 | bauzas | mriedem: I mean, when writing func tests | |
| 15:47:41 | mriedem | huh? | |
| 15:47:41 | jaypipes | bauzas: hmm? | |
| 15:47:48 | dansmith | lol | |
| 15:47:48 | mriedem | bauzas: people use caching scheduler for performance | |
| 15:48:03 | dansmith | he means test writers I think | |
| 15:48:04 | dansmith | I hope. | |
| 15:48:05 | bauzas | mriedem: so if we provide them either a fixture or something easy like dansmith said about just running without filters, that would be sufficient I guess | |
| 15:48:08 | mriedem | i don't think we can remove the caching scheduler until we've shown that the filter scheduler + claims in the scheduler outperforms caching scheduler at scale | |
| 15:48:20 | bauzas | mriedem: jaypipes: huh, s/caching/chance | |
| 15:48:34 | dansmith | mriedem: that means no removing the compute node claiming | |
| 15:48:42 | dansmith | mriedem: which I think is really terrible | |
| 15:49:08 | bauzas | sorry, was just explaining that we can I guess easily remove chance if we just provide a very standard fixture that would just run filterscheduler without all filters | |
| 15:49:12 | dansmith | all the complexity that is going to make jaypipes' head explode is because we're double-claiming with different rules in the RT | |
| 15:49:12 | bauzas | (c) dansmith | |
| 15:49:23 | bauzas | cachingscheduler is a totally different story | |
| 15:49:25 | mriedem | dansmith: we should be able to test that before removal | |
| 15:49:54 | bauzas | we want to remove caching because we do claims only in filter, right? | |
| 15:50:04 | mriedem | anyway, we're not removing those today, if we want to signal i think that's fine | |
| 15:50:06 | dansmith | mriedem: who is going to test that at scale do you think? | |
| 15:50:09 | bauzas | jaypipes: can't we imagine to have caching scheduler to claim as well ? | |
| 15:50:29 | dansmith | OSIC would have been an option, but tbh, I dunno who is likely to do it by the time we need to know in queens | |
| 15:50:34 | bauzas | jaypipes: you said you were blocked when untangling that, but do you think it's doable tho ? | |
| 15:50:45 | mriedem | dansmith: well, i know mirantis + intel has done some scale testing in the past, i don't know if those resources are still available, otherwise i'd like to try and see if huawei can do that, | |
| 15:50:55 | mriedem | because huawei public cloud is using the caching scheduler | |
| 15:50:57 | jaypipes | bauzas: caching scheduler would need to stop using the cached compute node information it has and instead call ComputeNodeList.get_all_by_uuid(), which would make the caching scheduler === the filter scheduler. | |
| 15:51:00 | dansmith | mriedem: okay well, that'd be cool | |
| 15:51:49 | bauzas | jaypipes: caching has different input than filterscheduler, but we could still claim the result right? | |
| 15:52:00 | dansmith | bauzas: not really | |
| 15:52:13 | dansmith | bauzas: if it's not getting allocation candidates, it is missing some of what it needs | |
| 15:52:21 | dansmith | bauzas: we could make it work, but it would be wonky | |
| 15:52:31 | bauzas | I see | |
| 15:52:34 | jaypipes | dansmith: and not worthwhile, imho. | |
| 15:52:36 | dansmith | bauzas: or we'd have to sift through the results of /ac to find the thing we chose based on our cached data | |
| 15:52:39 | dansmith | jaypipes: agreed | |
| 15:52:48 | bauzas | so, yeah, benchmarks | |
| 15:53:07 | dansmith | IIRC, the reason to keep caching scheduler after we moved to placement was because we weren't claiming and thus still making bad (cached) decisions in high traffic | |
| 15:53:23 | bauzas | I need to disappear for family business, but I will watch out thec convo | |
| 15:53:28 | bauzas | dansmith: you're correct | |
| 15:53:29 | dansmith | so claiming now should resolve that I think, and I would expect the gains of being able to run multiple schedulers would be more win than loss | |
| 15:53:30 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: Add a prelude section for Pike https://review.openstack.org/491424 | |
| 15:53:50 | bauzas | dansmith: right, I just mentioned that in the prelude ^ | |
| 15:54:01 | bauzas | dansmith: we can now run multiple scheduler workers | |
| 15:54:10 | bauzas | THAT is the big improvement | |
| 15:54:30 | dansmith | I would expect to be able to brute-force scale the scheduler as high as you want with just that change alone | |
| 15:54:40 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 15:54:45 | bauzas | scale out you mean ? | |
| 15:54:54 | bauzas | or scale up the load ? | |
| 15:55:02 | dansmith | bauzas: scale it to as high of a throughput as you need by scaling out yes | |
| 15:55:20 | bauzas | k | |
| 15:55:26 | bauzas | I agree | |
| 15:55:28 | gibi | we can easily remove chance from the func test. here is the patch: https://review.openstack.org/491529 | |
| 15:56:13 | bauzas | yeah | |
| 15:56:21 | bauzas | not a big deal | |
| 15:56:35 | bauzas | anyway I need to leave | |
| 15:56:38 | cdent | gibi: nice | |
| 15:56:38 | bauzas | ++ | |
| 15:56:46 | openstackgerrit | Stephen Finucane proposed openstack/nova master: hardware: Flatten functions https://review.openstack.org/367470 | |
| 15:56:47 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Rename '_numa_get_constraints_XXX' functions https://review.openstack.org/385072 | |
| 15:56:47 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Standardize '_get_XXX_constraints' functions https://review.openstack.org/385071 | |
| 15:56:48 | openstackgerrit | Stephen Finucane proposed openstack/nova master: De-duplicate _numa_get_flavor_XXX_map_list https://review.openstack.org/385074 | |
| 15:56:55 | bauzas | FWIW we don't even need RetryFilter | |
| 15:57:01 | bauzas | if we are singlehost | |
| 15:57:53 | sdague | so... is someone going deprecate chance completely? | |
| 15:58:01 | dansmith | sdague: and caching | |
| 15:58:10 | dansmith | we need to deprecate both in pike | |
| 15:58:10 | sdague | sure, that too | |
| 15:58:16 | dansmith | we can remove one or both when appropriate | |
| 15:58:29 | sdague | ok, is that up for review? we're in rc week | |
| 15:58:55 | jaypipes | gibi: why the change from 1 vcpu to 2 vcpus? that's weird.. | |
| 15:59:20 | dansmith | sdague: are you asking if you can do it? | |
| 15:59:24 | gibi | jaypipes: because there is resize to same host where the allocations are doubled up | |
| 15:59:36 | gibi | jaypipes: so at least 2 vcpu is needed | |
| 15:59:41 | dansmith | gibi: yeah | |
| 15:59:47 | jaypipes | gibi: guh, gotcha. | |
| 15:59:49 | dansmith | makes sense, I had to do that in some of the tests as well | |
| 16:00:04 | dansmith | jaypipes: this is the max_unit thing I was talking about with moves | |
| 16:00:05 | dansmith | jaypipes: that is going to piss people off | |
| 16:00:12 | dansmith | because allocation_ratio won't apply to single-host moves | |
| 16:00:46 | dansmith | migration uuid will fix that | |
| 16:00:50 | jaypipes | dansmith: yeah, understood. not sure it's the end of the world, though. | |
| 16:00:59 | cdent | migration uuid will fix everything! | |
| 16:01:02 | dansmith | eff yeah it's the end of the world | |
| 16:01:06 | jaypipes | :) | |
| 16:01:11 | dansmith | I have a single-node cloud at home! | |
| 16:01:26 | gibi | dansmith: only one? ;) | |
| 16:01:52 | dansmith | gibi: many nodes, but only one in the openstack deployment :) | |
| 16:03:17 | mnaser | i'm going over the scheduler code and i've realized that we're pretty much sending instances to our biggest hypervisors and the smaller ones remain empty. the ram weighter by default weighs based on free memory, which means that a hypervisor with 128gb memory and all of it free will have less priority than one that has 384gb with 192gb free | |
| 16:03:39 | mnaser | arguably.. shouldnt it be based on a percentage? | |
| 16:04:03 | mnaser | https://github.com/openstack/nova/blob/master/nova/scheduler/weights/ram.py | |
| 16:04:19 | jaypipes | mnaser: perhaps use the num instances weigher instead? | |