| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 15:48:08 | mriedem | i don't think we can remove the caching scheduler until we've shown that the filter scheduler + claims in the scheduler outperforms caching scheduler at scale | |
| 15:48:20 | bauzas | mriedem: jaypipes: huh, s/caching/chance | |
| 15:48:34 | dansmith | mriedem: that means no removing the compute node claiming | |
| 15:48:42 | dansmith | mriedem: which I think is really terrible | |
| 15:49:08 | bauzas | sorry, was just explaining that we can I guess easily remove chance if we just provide a very standard fixture that would just run filterscheduler without all filters | |
| 15:49:12 | bauzas | (c) dansmith | |
| 15:49:12 | dansmith | all the complexity that is going to make jaypipes' head explode is because we're double-claiming with different rules in the RT | |
| 15:49:23 | bauzas | cachingscheduler is a totally different story | |
| 15:49:25 | mriedem | dansmith: we should be able to test that before removal | |
| 15:49:54 | bauzas | we want to remove caching because we do claims only in filter, right? | |
| 15:50:04 | mriedem | anyway, we're not removing those today, if we want to signal i think that's fine | |
| 15:50:06 | dansmith | mriedem: who is going to test that at scale do you think? | |
| 15:50:09 | bauzas | jaypipes: can't we imagine to have caching scheduler to claim as well ? | |
| 15:50:29 | dansmith | OSIC would have been an option, but tbh, I dunno who is likely to do it by the time we need to know in queens | |
| 15:50:34 | bauzas | jaypipes: you said you were blocked when untangling that, but do you think it's doable tho ? | |
| 15:50:45 | mriedem | dansmith: well, i know mirantis + intel has done some scale testing in the past, i don't know if those resources are still available, otherwise i'd like to try and see if huawei can do that, | |
| 15:50:55 | mriedem | because huawei public cloud is using the caching scheduler | |
| 15:50:57 | jaypipes | bauzas: caching scheduler would need to stop using the cached compute node information it has and instead call ComputeNodeList.get_all_by_uuid(), which would make the caching scheduler === the filter scheduler. | |
| 15:51:00 | dansmith | mriedem: okay well, that'd be cool | |
| 15:51:49 | bauzas | jaypipes: caching has different input than filterscheduler, but we could still claim the result right? | |
| 15:52:00 | dansmith | bauzas: not really | |
| 15:52:13 | dansmith | bauzas: if it's not getting allocation candidates, it is missing some of what it needs | |
| 15:52:21 | dansmith | bauzas: we could make it work, but it would be wonky | |
| 15:52:31 | bauzas | I see | |
| 15:52:34 | jaypipes | dansmith: and not worthwhile, imho. | |
| 15:52:36 | dansmith | bauzas: or we'd have to sift through the results of /ac to find the thing we chose based on our cached data | |
| 15:52:39 | dansmith | jaypipes: agreed | |
| 15:52:48 | bauzas | so, yeah, benchmarks | |
| 15:53:07 | dansmith | IIRC, the reason to keep caching scheduler after we moved to placement was because we weren't claiming and thus still making bad (cached) decisions in high traffic | |
| 15:53:23 | bauzas | I need to disappear for family business, but I will watch out thec convo | |
| 15:53:28 | bauzas | dansmith: you're correct | |
| 15:53:29 | dansmith | so claiming now should resolve that I think, and I would expect the gains of being able to run multiple schedulers would be more win than loss | |
| 15:53:30 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: Add a prelude section for Pike https://review.openstack.org/491424 | |
| 15:53:50 | bauzas | dansmith: right, I just mentioned that in the prelude ^ | |
| 15:54:01 | bauzas | dansmith: we can now run multiple scheduler workers | |
| 15:54:10 | bauzas | THAT is the big improvement | |
| 15:54:30 | dansmith | I would expect to be able to brute-force scale the scheduler as high as you want with just that change alone | |
| 15:54:40 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 15:54:45 | bauzas | scale out you mean ? | |
| 15:54:54 | bauzas | or scale up the load ? | |
| 15:55:02 | dansmith | bauzas: scale it to as high of a throughput as you need by scaling out yes | |
| 15:55:20 | bauzas | k | |
| 15:55:26 | bauzas | I agree | |
| 15:55:28 | gibi | we can easily remove chance from the func test. here is the patch: https://review.openstack.org/491529 | |
| 15:56:13 | bauzas | yeah | |
| 15:56:21 | bauzas | not a big deal | |
| 15:56:35 | bauzas | anyway I need to leave | |
| 15:56:38 | bauzas | ++ | |
| 15:56:38 | cdent | gibi: nice | |
| 15:56:46 | openstackgerrit | Stephen Finucane proposed openstack/nova master: hardware: Flatten functions https://review.openstack.org/367470 | |
| 15:56:47 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Standardize '_get_XXX_constraints' functions https://review.openstack.org/385071 | |
| 15:56:47 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Rename '_numa_get_constraints_XXX' functions https://review.openstack.org/385072 | |
| 15:56:48 | openstackgerrit | Stephen Finucane proposed openstack/nova master: De-duplicate _numa_get_flavor_XXX_map_list https://review.openstack.org/385074 | |
| 15:56:55 | bauzas | FWIW we don't even need RetryFilter | |
| 15:57:01 | bauzas | if we are singlehost | |
| 15:57:53 | sdague | so... is someone going deprecate chance completely? | |
| 15:58:01 | dansmith | sdague: and caching | |
| 15:58:10 | sdague | sure, that too | |
| 15:58:10 | dansmith | we need to deprecate both in pike | |
| 15:58:16 | dansmith | we can remove one or both when appropriate | |
| 15:58:29 | sdague | ok, is that up for review? we're in rc week | |
| 15:58:55 | jaypipes | gibi: why the change from 1 vcpu to 2 vcpus? that's weird.. | |
| 15:59:20 | dansmith | sdague: are you asking if you can do it? | |
| 15:59:24 | gibi | jaypipes: because there is resize to same host where the allocations are doubled up | |
| 15:59:36 | gibi | jaypipes: so at least 2 vcpu is needed | |
| 15:59:41 | dansmith | gibi: yeah | |
| 15:59:47 | jaypipes | gibi: guh, gotcha. | |
| 15:59:49 | dansmith | makes sense, I had to do that in some of the tests as well | |
| 16:00:04 | dansmith | jaypipes: this is the max_unit thing I was talking about with moves | |
| 16:00:05 | dansmith | jaypipes: that is going to piss people off | |
| 16:00:12 | dansmith | because allocation_ratio won't apply to single-host moves | |
| 16:00:46 | dansmith | migration uuid will fix that | |
| 16:00:50 | jaypipes | dansmith: yeah, understood. not sure it's the end of the world, though. | |
| 16:00:59 | cdent | migration uuid will fix everything! | |
| 16:01:02 | dansmith | eff yeah it's the end of the world | |
| 16:01:06 | jaypipes | :) | |
| 16:01:11 | dansmith | I have a single-node cloud at home! | |
| 16:01:26 | gibi | dansmith: only one? ;) | |
| 16:01:52 | dansmith | gibi: many nodes, but only one in the openstack deployment :) | |
| 16:03:17 | mnaser | i'm going over the scheduler code and i've realized that we're pretty much sending instances to our biggest hypervisors and the smaller ones remain empty. the ram weighter by default weighs based on free memory, which means that a hypervisor with 128gb memory and all of it free will have less priority than one that has 384gb with 192gb free | |
| 16:03:39 | mnaser | arguably.. shouldnt it be based on a percentage? | |
| 16:04:03 | mnaser | https://github.com/openstack/nova/blob/master/nova/scheduler/weights/ram.py | |
| 16:04:19 | jaypipes | mnaser: perhaps use the num instances weigher instead? | |
| 16:04:43 | dansmith | or weigh the num_instances higher than the ram one | |
| 16:04:49 | dansmith | mnaser: some people want to pack first | |
| 16:04:50 | jaypipes | mnaser: nm, there is no such thing :( | |
| 16:04:52 | dansmith | because windows licenses | |
| 16:05:10 | mnaser | dansmith: packing would still work with -1.0 | |
| 16:05:13 | dansmith | jaypipes: there's something like that pretty sure | |
| 16:05:29 | mnaser | https://github.com/openstack/nova/tree/master/nova/scheduler/weights | |
| 16:05:35 | mnaser | i dont see a num_instances on | |
| 16:05:38 | mriedem | stephenfin: where does this come from? https://review.openstack.org/#/c/490952/3/doc/source/admin/security-groups.rst | |
| 16:05:54 | mnaser | there's a fitler but not a weighter | |
| 16:05:58 | stephenfin | mriedem: cli-nova-manage-projects-security.rst | |
| 16:06:01 | stephenfin | (I think | |
| 16:06:03 | jaypipes | mnaser: yeah :( | |
| 16:06:03 | mriedem | oh nvm https://github.com/openstack/openstack-manuals/blob/stable/ocata/doc/admin-guide/source/cli-nova-manage-projects-security.rst | |
| 16:06:04 | mriedem | yeah | |
| 16:06:18 | dansmith | mnaser: yeah, that's what I'm thinking of, heh | |
| 16:06:47 | mnaser | cause if it was percentage based, the ability to pack vs distribute would still work, but would be based on usage percentage rather than just absolute free memory | |