Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-07
15:00:25 jaypipes mriedem: another thing I've been thinking of is instead of further mucking with the resource tracker and/or making the scheduler report client that we do a nova-status or nova-manage audit function that cleans up allocations based on the cell DB states.
15:01:28 jaypipes mriedem: and once done, we have all computes on service version 22 (or whatever) and after that point, we remove all allocation stuff from the compute node entirely (other than the aforementioned delete allocations pieces)
15:02:37 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove files from 'tools' https://review.openstack.org/491444
15:02:37 openstackgerrit Stephen Finucane proposed openstack/nova master: trivial: Remove "vif" script https://review.openstack.org/491443
15:03:05 lennyb dansmith, thanks, I will recheck it.
15:03:08 jaypipes mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :(
15:03:09 mriedem jaypipes: the service version won't change until the actual code is upgraded though
15:03:18 mriedem jaypipes: i agree with that
15:03:38 dansmith is it alternate-ego mondays? why is jaypipes being "Frank" ?
15:03:45 jaypipes heh
15:05:04 jaypipes mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information.
15:06:11 jaypipes mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client.
15:06:14 dansmith jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing?
15:06:19 jaypipes mriedem: choose your poison I think?
15:06:27 jaypipes dansmith: yeah.
15:06:29 sdague mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work
15:06:42 dansmith jaypipes: that would be really unfortunate, IMHO
15:06:49 dansmith jaypipes: some people take a looong time to do that upgrade
15:06:50 jaypipes dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again.
15:08:11 mriedem sdague: i can look in a bit
15:09:13 jaypipes dansmith: do you think we should even attempt to fix the shared storage reporting in Pike?
15:09:27 jaypipes dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1
15:09:39 dansmith jaypipes: well, tbh, I thought that was out the window a while ago
15:09:51 dansmith the computes have no idea which shared storage uuid a given thing is using anyway right?
15:10:06 jaypipes dansmith: right. and ocatas are just going to pummel the allocations either way.
15:11:22 mriedem jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes
15:11:49 jaypipes dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike.
15:11:55 mriedem kind of weird to say, "this feature doesn't yet work" in a known issues section though
15:12:07 mriedem i don't know of omission is better than clearly saying it's not supported
15:12:23 mriedem i'd rather be clear personally
15:12:28 mriedem of what is supported and what's not
15:12:30 jaypipes mriedem: yeah, agreed
15:13:11 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517
15:13:13 jaypipes mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly.
15:13:20 jaypipes dansmith: agreed?
15:13:22 dansmith I dunno why anyone would think it's supported unless you said it was, but okay
15:13:36 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add additional content to admin guide https://review.openstack.org/490952
15:16:59 bauzas mmm, interesting, generation of relnotes fails if you have two reno files sharing the same explicit RST target
15:33:25 jaypipes dansmith: hey, did you catch my question above about agreeing/disagreeing with this:
15:33:27 jaypipes mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly.
15:34:23 dansmith jaypipes: yeah, but that's not a change in course right?
15:34:26 dansmith I mean, that's been the plan?
15:34:59 dansmith it'd be really nice if we were to remove the compute allocations behavior by removing all the legacy RT stuff in the compute node at the same time
15:35:30 dansmith mriedem: in support of that goal, and circling back to the discussion this morning, shouldn't we deprecate chance and caching in pike?
15:36:16 jaypipes dansmith: correct, that was the plan./
15:36:28 dansmith jaypipes: okay then, yea agreed :)
15:37:20 jaypipes dansmith: w.r.t. to deprecating the caching and chance scheduler, I'd be a-ok with that, but that's also why I was quite concerned with the earlier conversation between gibi, me and bauzas where it was posited that our func tests were using the chance scheduler :(
15:37:46 bauzas jaypipes: dansmith: looks a good discussion for Queens, nope ?
15:37:53 bauzas at least for CachingScheduler
15:38:02 dansmith bauzas: very specifically we need to deprecate both now, which is what I just said
15:38:07 bauzas because we could ask operators to try out our placement nice features
15:38:31 bauzas dansmith: if that's all but just sending a signal, I'm okay
15:38:35 dansmith bauzas: it's about deprecating things before we break them entirely
15:38:50 bauzas dansmith: if that's implying that we will remove that in Queens, I dunno
15:39:04 dansmith I just said it was about removing this stuff in queens,
15:39:06 dansmith or at least being able to
15:39:07 jaypipes bauzas: it's about deprecating in Pike.
15:39:13 jaypipes bauzas: and removing in Queens.
15:39:17 dansmith yes.
15:39:27 bauzas tbh, I don't really care of chance scheduler
15:39:39 bauzas I just gave explanation about why people use it
15:39:42 gibi jaypipes: I'm changing the func test env to use filter_scheduler as we speak
15:39:50 gibi jaypipes: it doesn't seems terribly hard to do
15:39:55 efried jamielennox yt?
15:40:01 jaypipes gibi: that's good news indeed.
15:40:09 dansmith gibi: chance with one (fake) host should be pretty much the same behavior I'd hope :)
15:40:12 bauzas I'm a little more concerned by cachingscheduler, which is AFAIK heavely used somewhere
15:40:41 bauzas gibi: I don't expect real breakages tbh
15:40:52 bauzas since tests using chance are just single-host
15:40:57 gibi I think I can push that patch today
15:41:01 jaypipes dansmith: well, the problem is chance essentially disables the placement API in scheduler, which disables claims in the scheduler, which means everything relies on allocations done on the compute node. :(
15:41:31 dansmith jaypipes: ....right... did I say something wrong above?
15:41:47 dansmith or you're just saying that changing that could break functional tests in weird ways?
15:42:11 dansmith if the latter, that's why I said "pretty much the same behavior" :)
15:42:16 dansmith ideally, and hopefully
15:42:52 jaypipes dansmith: well, I was just saying that's the reason I was dismayed to hear from gibi that the chance scheduler was used for almost all the functional tests...
15:43:00 dansmith oh, indeed
15:43:09 jaypipes anyway, sounds like we're in violent agreement, so I'll shut it.
15:43:14 mriedem we could always write a scheduler driver just used in tests - we have the fake driver
15:43:21 mriedem we could change the fake driver to be like the chance driver
15:43:31 jaypipes mriedem: I hope you're kidding. :)
15:43:32 mriedem but best to just use stripped down filter scheduler in functional tests
15:43:43 mriedem note that you can still plug in scheduler drivers
15:43:43 dansmith jaypipes: I think he's kidding
15:43:45 mriedem which won't use placement
15:44:19 jaypipes dansmith: you never know with these South Dakatons.
15:44:24 mriedem so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option
15:44:42 dansmith if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions
15:44:46 dansmith which is fine I guess
15:45:40 mriedem i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those
15:45:41 dansmith like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release
15:46:02 mriedem we removed the ability for the compute manager to be pluggable, or do you mean the virt driver?
15:46:12 dansmith that's my point,
15:46:31 dansmith there's basically zero chance you've written your own compute manager that does everything properly
15:46:55 dansmith and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today
15:47:08 openstackgerrit Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469
15:47:10 mriedem well, you just didn't upgrade from kilo L)
15:47:11 mriedem :)
15:47:19 mriedem which is like 50% of deployments probably
15:47:24 bauzas mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters

Earlier   Later