| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 15:13:20 | jaypipes | dansmith: agreed? | |
| 15:13:22 | dansmith | I dunno why anyone would think it's supported unless you said it was, but okay | |
| 15:13:36 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Add additional content to admin guide https://review.openstack.org/490952 | |
| 15:16:59 | bauzas | mmm, interesting, generation of relnotes fails if you have two reno files sharing the same explicit RST target | |
| 15:33:25 | jaypipes | dansmith: hey, did you catch my question above about agreeing/disagreeing with this: | |
| 15:33:27 | jaypipes | mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly. | |
| 15:34:23 | dansmith | jaypipes: yeah, but that's not a change in course right? | |
| 15:34:26 | dansmith | I mean, that's been the plan? | |
| 15:34:59 | dansmith | it'd be really nice if we were to remove the compute allocations behavior by removing all the legacy RT stuff in the compute node at the same time | |
| 15:35:30 | dansmith | mriedem: in support of that goal, and circling back to the discussion this morning, shouldn't we deprecate chance and caching in pike? | |
| 15:36:16 | jaypipes | dansmith: correct, that was the plan./ | |
| 15:36:28 | dansmith | jaypipes: okay then, yea agreed :) | |
| 15:37:20 | jaypipes | dansmith: w.r.t. to deprecating the caching and chance scheduler, I'd be a-ok with that, but that's also why I was quite concerned with the earlier conversation between gibi, me and bauzas where it was posited that our func tests were using the chance scheduler :( | |
| 15:37:46 | bauzas | jaypipes: dansmith: looks a good discussion for Queens, nope ? | |
| 15:37:53 | bauzas | at least for CachingScheduler | |
| 15:38:02 | dansmith | bauzas: very specifically we need to deprecate both now, which is what I just said | |
| 15:38:07 | bauzas | because we could ask operators to try out our placement nice features | |
| 15:38:31 | bauzas | dansmith: if that's all but just sending a signal, I'm okay | |
| 15:38:35 | dansmith | bauzas: it's about deprecating things before we break them entirely | |
| 15:38:50 | bauzas | dansmith: if that's implying that we will remove that in Queens, I dunno | |
| 15:39:04 | dansmith | I just said it was about removing this stuff in queens, | |
| 15:39:06 | dansmith | or at least being able to | |
| 15:39:07 | jaypipes | bauzas: it's about deprecating in Pike. | |
| 15:39:13 | jaypipes | bauzas: and removing in Queens. | |
| 15:39:17 | dansmith | yes. | |
| 15:39:27 | bauzas | tbh, I don't really care of chance scheduler | |
| 15:39:39 | bauzas | I just gave explanation about why people use it | |
| 15:39:42 | gibi | jaypipes: I'm changing the func test env to use filter_scheduler as we speak | |
| 15:39:50 | gibi | jaypipes: it doesn't seems terribly hard to do | |
| 15:39:55 | efried | jamielennox yt? | |
| 15:40:01 | jaypipes | gibi: that's good news indeed. | |
| 15:40:09 | dansmith | gibi: chance with one (fake) host should be pretty much the same behavior I'd hope :) | |
| 15:40:12 | bauzas | I'm a little more concerned by cachingscheduler, which is AFAIK heavely used somewhere | |
| 15:40:41 | bauzas | gibi: I don't expect real breakages tbh | |
| 15:40:52 | bauzas | since tests using chance are just single-host | |
| 15:40:57 | gibi | I think I can push that patch today | |
| 15:41:01 | jaypipes | dansmith: well, the problem is chance essentially disables the placement API in scheduler, which disables claims in the scheduler, which means everything relies on allocations done on the compute node. :( | |
| 15:41:31 | dansmith | jaypipes: ....right... did I say something wrong above? | |
| 15:41:47 | dansmith | or you're just saying that changing that could break functional tests in weird ways? | |
| 15:42:11 | dansmith | if the latter, that's why I said "pretty much the same behavior" :) | |
| 15:42:16 | dansmith | ideally, and hopefully | |
| 15:42:52 | jaypipes | dansmith: well, I was just saying that's the reason I was dismayed to hear from gibi that the chance scheduler was used for almost all the functional tests... | |
| 15:43:00 | dansmith | oh, indeed | |
| 15:43:09 | jaypipes | anyway, sounds like we're in violent agreement, so I'll shut it. | |
| 15:43:14 | mriedem | we could always write a scheduler driver just used in tests - we have the fake driver | |
| 15:43:21 | mriedem | we could change the fake driver to be like the chance driver | |
| 15:43:31 | jaypipes | mriedem: I hope you're kidding. :) | |
| 15:43:32 | mriedem | but best to just use stripped down filter scheduler in functional tests | |
| 15:43:43 | mriedem | note that you can still plug in scheduler drivers | |
| 15:43:43 | dansmith | jaypipes: I think he's kidding | |
| 15:43:45 | mriedem | which won't use placement | |
| 15:44:19 | jaypipes | dansmith: you never know with these South Dakatons. | |
| 15:44:24 | mriedem | so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option | |
| 15:44:42 | dansmith | if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions | |
| 15:44:46 | dansmith | which is fine I guess | |
| 15:45:40 | mriedem | i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those | |
| 15:45:41 | dansmith | like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release | |
| 15:46:02 | mriedem | we removed the ability for the compute manager to be pluggable, or do you mean the virt driver? | |
| 15:46:12 | dansmith | that's my point, | |
| 15:46:31 | dansmith | there's basically zero chance you've written your own compute manager that does everything properly | |
| 15:46:55 | dansmith | and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today | |
| 15:47:08 | openstackgerrit | Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469 | |
| 15:47:10 | mriedem | well, you just didn't upgrade from kilo L) | |
| 15:47:11 | mriedem | :) | |
| 15:47:19 | mriedem | which is like 50% of deployments probably | |
| 15:47:24 | bauzas | mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters | |
| 15:47:34 | bauzas | mriedem: I mean, when writing func tests | |
| 15:47:41 | jaypipes | bauzas: hmm? | |
| 15:47:41 | mriedem | huh? | |
| 15:47:48 | mriedem | bauzas: people use caching scheduler for performance | |
| 15:47:48 | dansmith | lol | |
| 15:48:03 | dansmith | he means test writers I think | |
| 15:48:04 | dansmith | I hope. | |
| 15:48:05 | bauzas | mriedem: so if we provide them either a fixture or something easy like dansmith said about just running without filters, that would be sufficient I guess | |
| 15:48:08 | mriedem | i don't think we can remove the caching scheduler until we've shown that the filter scheduler + claims in the scheduler outperforms caching scheduler at scale | |
| 15:48:20 | bauzas | mriedem: jaypipes: huh, s/caching/chance | |
| 15:48:34 | dansmith | mriedem: that means no removing the compute node claiming | |
| 15:48:42 | dansmith | mriedem: which I think is really terrible | |
| 15:49:08 | bauzas | sorry, was just explaining that we can I guess easily remove chance if we just provide a very standard fixture that would just run filterscheduler without all filters | |
| 15:49:12 | bauzas | (c) dansmith | |
| 15:49:12 | dansmith | all the complexity that is going to make jaypipes' head explode is because we're double-claiming with different rules in the RT | |
| 15:49:23 | bauzas | cachingscheduler is a totally different story | |
| 15:49:25 | mriedem | dansmith: we should be able to test that before removal | |
| 15:49:54 | bauzas | we want to remove caching because we do claims only in filter, right? | |
| 15:50:04 | mriedem | anyway, we're not removing those today, if we want to signal i think that's fine | |
| 15:50:06 | dansmith | mriedem: who is going to test that at scale do you think? | |
| 15:50:09 | bauzas | jaypipes: can't we imagine to have caching scheduler to claim as well ? | |
| 15:50:29 | dansmith | OSIC would have been an option, but tbh, I dunno who is likely to do it by the time we need to know in queens | |
| 15:50:34 | bauzas | jaypipes: you said you were blocked when untangling that, but do you think it's doable tho ? | |
| 15:50:45 | mriedem | dansmith: well, i know mirantis + intel has done some scale testing in the past, i don't know if those resources are still available, otherwise i'd like to try and see if huawei can do that, | |
| 15:50:55 | mriedem | because huawei public cloud is using the caching scheduler | |
| 15:50:57 | jaypipes | bauzas: caching scheduler would need to stop using the cached compute node information it has and instead call ComputeNodeList.get_all_by_uuid(), which would make the caching scheduler === the filter scheduler. | |
| 15:51:00 | dansmith | mriedem: okay well, that'd be cool | |
| 15:51:49 | bauzas | jaypipes: caching has different input than filterscheduler, but we could still claim the result right? | |
| 15:52:00 | dansmith | bauzas: not really | |
| 15:52:13 | dansmith | bauzas: if it's not getting allocation candidates, it is missing some of what it needs | |
| 15:52:21 | dansmith | bauzas: we could make it work, but it would be wonky | |
| 15:52:31 | bauzas | I see | |
| 15:52:34 | jaypipes | dansmith: and not worthwhile, imho. | |
| 15:52:36 | dansmith | bauzas: or we'd have to sift through the results of /ac to find the thing we chose based on our cached data | |