| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 15:03:08 | jaypipes | mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :( | |
| 15:03:09 | mriedem | jaypipes: the service version won't change until the actual code is upgraded though | |
| 15:03:18 | mriedem | jaypipes: i agree with that | |
| 15:03:38 | dansmith | is it alternate-ego mondays? why is jaypipes being "Frank" ? | |
| 15:03:45 | jaypipes | heh | |
| 15:05:04 | jaypipes | mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information. | |
| 15:06:11 | jaypipes | mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client. | |
| 15:06:14 | dansmith | jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing? | |
| 15:06:19 | jaypipes | mriedem: choose your poison I think? | |
| 15:06:27 | jaypipes | dansmith: yeah. | |
| 15:06:29 | sdague | mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work | |
| 15:06:42 | dansmith | jaypipes: that would be really unfortunate, IMHO | |
| 15:06:49 | dansmith | jaypipes: some people take a looong time to do that upgrade | |
| 15:06:50 | jaypipes | dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again. | |
| 15:08:11 | mriedem | sdague: i can look in a bit | |
| 15:09:13 | jaypipes | dansmith: do you think we should even attempt to fix the shared storage reporting in Pike? | |
| 15:09:27 | jaypipes | dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1 | |
| 15:09:39 | dansmith | jaypipes: well, tbh, I thought that was out the window a while ago | |
| 15:09:51 | dansmith | the computes have no idea which shared storage uuid a given thing is using anyway right? | |
| 15:10:06 | jaypipes | dansmith: right. and ocatas are just going to pummel the allocations either way. | |
| 15:11:22 | mriedem | jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes | |
| 15:11:49 | jaypipes | dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike. | |
| 15:11:55 | mriedem | kind of weird to say, "this feature doesn't yet work" in a known issues section though | |
| 15:12:07 | mriedem | i don't know of omission is better than clearly saying it's not supported | |
| 15:12:23 | mriedem | i'd rather be clear personally | |
| 15:12:28 | mriedem | of what is supported and what's not | |
| 15:12:30 | jaypipes | mriedem: yeah, agreed | |
| 15:13:11 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517 | |
| 15:13:13 | jaypipes | mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly. | |
| 15:13:20 | jaypipes | dansmith: agreed? | |
| 15:13:22 | dansmith | I dunno why anyone would think it's supported unless you said it was, but okay | |
| 15:13:36 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Add additional content to admin guide https://review.openstack.org/490952 | |
| 15:16:59 | bauzas | mmm, interesting, generation of relnotes fails if you have two reno files sharing the same explicit RST target | |
| 15:33:25 | jaypipes | dansmith: hey, did you catch my question above about agreeing/disagreeing with this: | |
| 15:33:27 | jaypipes | mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly. | |
| 15:34:23 | dansmith | jaypipes: yeah, but that's not a change in course right? | |
| 15:34:26 | dansmith | I mean, that's been the plan? | |
| 15:34:59 | dansmith | it'd be really nice if we were to remove the compute allocations behavior by removing all the legacy RT stuff in the compute node at the same time | |
| 15:35:30 | dansmith | mriedem: in support of that goal, and circling back to the discussion this morning, shouldn't we deprecate chance and caching in pike? | |
| 15:36:16 | jaypipes | dansmith: correct, that was the plan./ | |
| 15:36:28 | dansmith | jaypipes: okay then, yea agreed :) | |
| 15:37:20 | jaypipes | dansmith: w.r.t. to deprecating the caching and chance scheduler, I'd be a-ok with that, but that's also why I was quite concerned with the earlier conversation between gibi, me and bauzas where it was posited that our func tests were using the chance scheduler :( | |
| 15:37:46 | bauzas | jaypipes: dansmith: looks a good discussion for Queens, nope ? | |
| 15:37:53 | bauzas | at least for CachingScheduler | |
| 15:38:02 | dansmith | bauzas: very specifically we need to deprecate both now, which is what I just said | |
| 15:38:07 | bauzas | because we could ask operators to try out our placement nice features | |
| 15:38:31 | bauzas | dansmith: if that's all but just sending a signal, I'm okay | |
| 15:38:35 | dansmith | bauzas: it's about deprecating things before we break them entirely | |
| 15:38:50 | bauzas | dansmith: if that's implying that we will remove that in Queens, I dunno | |
| 15:39:04 | dansmith | I just said it was about removing this stuff in queens, | |
| 15:39:06 | dansmith | or at least being able to | |
| 15:39:07 | jaypipes | bauzas: it's about deprecating in Pike. | |
| 15:39:13 | jaypipes | bauzas: and removing in Queens. | |
| 15:39:17 | dansmith | yes. | |
| 15:39:27 | bauzas | tbh, I don't really care of chance scheduler | |
| 15:39:39 | bauzas | I just gave explanation about why people use it | |
| 15:39:42 | gibi | jaypipes: I'm changing the func test env to use filter_scheduler as we speak | |
| 15:39:50 | gibi | jaypipes: it doesn't seems terribly hard to do | |
| 15:39:55 | efried | jamielennox yt? | |
| 15:40:01 | jaypipes | gibi: that's good news indeed. | |
| 15:40:09 | dansmith | gibi: chance with one (fake) host should be pretty much the same behavior I'd hope :) | |
| 15:40:12 | bauzas | I'm a little more concerned by cachingscheduler, which is AFAIK heavely used somewhere | |
| 15:40:41 | bauzas | gibi: I don't expect real breakages tbh | |
| 15:40:52 | bauzas | since tests using chance are just single-host | |
| 15:40:57 | gibi | I think I can push that patch today | |
| 15:41:01 | jaypipes | dansmith: well, the problem is chance essentially disables the placement API in scheduler, which disables claims in the scheduler, which means everything relies on allocations done on the compute node. :( | |
| 15:41:31 | dansmith | jaypipes: ....right... did I say something wrong above? | |
| 15:41:47 | dansmith | or you're just saying that changing that could break functional tests in weird ways? | |
| 15:42:11 | dansmith | if the latter, that's why I said "pretty much the same behavior" :) | |
| 15:42:16 | dansmith | ideally, and hopefully | |
| 15:42:52 | jaypipes | dansmith: well, I was just saying that's the reason I was dismayed to hear from gibi that the chance scheduler was used for almost all the functional tests... | |
| 15:43:00 | dansmith | oh, indeed | |
| 15:43:09 | jaypipes | anyway, sounds like we're in violent agreement, so I'll shut it. | |
| 15:43:14 | mriedem | we could always write a scheduler driver just used in tests - we have the fake driver | |
| 15:43:21 | mriedem | we could change the fake driver to be like the chance driver | |
| 15:43:31 | jaypipes | mriedem: I hope you're kidding. :) | |
| 15:43:32 | mriedem | but best to just use stripped down filter scheduler in functional tests | |
| 15:43:43 | dansmith | jaypipes: I think he's kidding | |
| 15:43:43 | mriedem | note that you can still plug in scheduler drivers | |
| 15:43:45 | mriedem | which won't use placement | |
| 15:44:19 | jaypipes | dansmith: you never know with these South Dakatons. | |
| 15:44:24 | mriedem | so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option | |
| 15:44:42 | dansmith | if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions | |
| 15:44:46 | dansmith | which is fine I guess | |
| 15:45:40 | mriedem | i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those | |
| 15:45:41 | dansmith | like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release | |
| 15:46:02 | mriedem | we removed the ability for the compute manager to be pluggable, or do you mean the virt driver? | |
| 15:46:12 | dansmith | that's my point, | |
| 15:46:31 | dansmith | there's basically zero chance you've written your own compute manager that does everything properly | |
| 15:46:55 | dansmith | and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today | |
| 15:47:08 | openstackgerrit | Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469 | |
| 15:47:10 | mriedem | well, you just didn't upgrade from kilo L) | |
| 15:47:11 | mriedem | :) | |
| 15:47:19 | mriedem | which is like 50% of deployments probably | |
| 15:47:24 | bauzas | mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters | |
| 15:47:34 | bauzas | mriedem: I mean, when writing func tests | |
| 15:47:41 | mriedem | huh? | |
| 15:47:41 | jaypipes | bauzas: hmm? | |
| 15:47:48 | dansmith | lol | |
| 15:47:48 | mriedem | bauzas: people use caching scheduler for performance | |