Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-07
15:43:43 dansmith jaypipes: I think he's kidding
15:43:43 mriedem note that you can still plug in scheduler drivers
15:43:45 mriedem which won't use placement
15:44:19 jaypipes dansmith: you never know with these South Dakatons.
15:44:24 mriedem so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option
15:44:42 dansmith if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions
15:44:46 dansmith which is fine I guess
15:45:40 mriedem i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those
15:45:41 dansmith like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release
15:46:02 mriedem we removed the ability for the compute manager to be pluggable, or do you mean the virt driver?
15:46:12 dansmith that's my point,
15:46:31 dansmith there's basically zero chance you've written your own compute manager that does everything properly
15:46:55 dansmith and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today
15:47:08 openstackgerrit Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469
15:47:10 mriedem well, you just didn't upgrade from kilo L)
15:47:11 mriedem :)
15:47:19 mriedem which is like 50% of deployments probably
15:47:24 bauzas mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters
15:47:34 bauzas mriedem: I mean, when writing func tests
15:47:41 mriedem huh?
15:47:41 jaypipes bauzas: hmm?
15:47:48 dansmith lol
15:47:48 mriedem bauzas: people use caching scheduler for performance
15:48:03 dansmith he means test writers I think
15:48:04 dansmith I hope.
15:48:05 bauzas mriedem: so if we provide them either a fixture or something easy like dansmith said about just running without filters, that would be sufficient I guess
15:48:08 mriedem i don't think we can remove the caching scheduler until we've shown that the filter scheduler + claims in the scheduler outperforms caching scheduler at scale
15:48:20 bauzas mriedem: jaypipes: huh, s/caching/chance
15:48:34 dansmith mriedem: that means no removing the compute node claiming
15:48:42 dansmith mriedem: which I think is really terrible
15:49:08 bauzas sorry, was just explaining that we can I guess easily remove chance if we just provide a very standard fixture that would just run filterscheduler without all filters
15:49:12 dansmith all the complexity that is going to make jaypipes' head explode is because we're double-claiming with different rules in the RT
15:49:12 bauzas (c) dansmith
15:49:23 bauzas cachingscheduler is a totally different story
15:49:25 mriedem dansmith: we should be able to test that before removal
15:49:54 bauzas we want to remove caching because we do claims only in filter, right?
15:50:04 mriedem anyway, we're not removing those today, if we want to signal i think that's fine
15:50:06 dansmith mriedem: who is going to test that at scale do you think?
15:50:09 bauzas jaypipes: can't we imagine to have caching scheduler to claim as well ?
15:50:29 dansmith OSIC would have been an option, but tbh, I dunno who is likely to do it by the time we need to know in queens
15:50:34 bauzas jaypipes: you said you were blocked when untangling that, but do you think it's doable tho ?
15:50:45 mriedem dansmith: well, i know mirantis + intel has done some scale testing in the past, i don't know if those resources are still available, otherwise i'd like to try and see if huawei can do that,
15:50:55 mriedem because huawei public cloud is using the caching scheduler
15:50:57 jaypipes bauzas: caching scheduler would need to stop using the cached compute node information it has and instead call ComputeNodeList.get_all_by_uuid(), which would make the caching scheduler === the filter scheduler.
15:51:00 dansmith mriedem: okay well, that'd be cool
15:51:49 bauzas jaypipes: caching has different input than filterscheduler, but we could still claim the result right?
15:52:00 dansmith bauzas: not really
15:52:13 dansmith bauzas: if it's not getting allocation candidates, it is missing some of what it needs
15:52:21 dansmith bauzas: we could make it work, but it would be wonky
15:52:31 bauzas I see
15:52:34 jaypipes dansmith: and not worthwhile, imho.
15:52:36 dansmith bauzas: or we'd have to sift through the results of /ac to find the thing we chose based on our cached data
15:52:39 dansmith jaypipes: agreed
15:52:48 bauzas so, yeah, benchmarks
15:53:07 dansmith IIRC, the reason to keep caching scheduler after we moved to placement was because we weren't claiming and thus still making bad (cached) decisions in high traffic
15:53:23 bauzas I need to disappear for family business, but I will watch out thec convo
15:53:28 bauzas dansmith: you're correct
15:53:29 dansmith so claiming now should resolve that I think, and I would expect the gains of being able to run multiple schedulers would be more win than loss
15:53:30 openstackgerrit Sylvain Bauza proposed openstack/nova master: Add a prelude section for Pike https://review.openstack.org/491424
15:53:50 bauzas dansmith: right, I just mentioned that in the prelude ^
15:54:01 bauzas dansmith: we can now run multiple scheduler workers
15:54:10 bauzas THAT is the big improvement
15:54:30 dansmith I would expect to be able to brute-force scale the scheduler as high as you want with just that change alone
15:54:40 openstackgerrit Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529
15:54:45 bauzas scale out you mean ?
15:54:54 bauzas or scale up the load ?
15:55:02 dansmith bauzas: scale it to as high of a throughput as you need by scaling out yes
15:55:20 bauzas k
15:55:26 bauzas I agree
15:55:28 gibi we can easily remove chance from the func test. here is the patch: https://review.openstack.org/491529
15:56:13 bauzas yeah
15:56:21 bauzas not a big deal
15:56:35 bauzas anyway I need to leave
15:56:38 cdent gibi: nice
15:56:38 bauzas ++
15:56:46 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Flatten functions https://review.openstack.org/367470
15:56:47 openstackgerrit Stephen Finucane proposed openstack/nova master: Rename '_numa_get_constraints_XXX' functions https://review.openstack.org/385072
15:56:47 openstackgerrit Stephen Finucane proposed openstack/nova master: Standardize '_get_XXX_constraints' functions https://review.openstack.org/385071
15:56:48 openstackgerrit Stephen Finucane proposed openstack/nova master: De-duplicate _numa_get_flavor_XXX_map_list https://review.openstack.org/385074
15:56:55 bauzas FWIW we don't even need RetryFilter
15:57:01 bauzas if we are singlehost
15:57:53 sdague so... is someone going deprecate chance completely?
15:58:01 dansmith sdague: and caching
15:58:10 dansmith we need to deprecate both in pike
15:58:10 sdague sure, that too
15:58:16 dansmith we can remove one or both when appropriate
15:58:29 sdague ok, is that up for review? we're in rc week
15:58:55 jaypipes gibi: why the change from 1 vcpu to 2 vcpus? that's weird..
15:59:20 dansmith sdague: are you asking if you can do it?
15:59:24 gibi jaypipes: because there is resize to same host where the allocations are doubled up
15:59:36 gibi jaypipes: so at least 2 vcpu is needed
15:59:41 dansmith gibi: yeah
15:59:47 jaypipes gibi: guh, gotcha.
15:59:49 dansmith makes sense, I had to do that in some of the tests as well
16:00:04 dansmith jaypipes: this is the max_unit thing I was talking about with moves
16:00:05 dansmith jaypipes: that is going to piss people off
16:00:12 dansmith because allocation_ratio won't apply to single-host moves
16:00:46 dansmith migration uuid will fix that
16:00:50 jaypipes dansmith: yeah, understood. not sure it's the end of the world, though.
16:00:59 cdent migration uuid will fix everything!

Earlier   Later