Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-07
15:05:04 jaypipes mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information.
15:06:11 jaypipes mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client.
15:06:14 dansmith jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing?
15:06:19 jaypipes mriedem: choose your poison I think?
15:06:27 jaypipes dansmith: yeah.
15:06:29 sdague mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work
15:06:42 dansmith jaypipes: that would be really unfortunate, IMHO
15:06:49 dansmith jaypipes: some people take a looong time to do that upgrade
15:06:50 jaypipes dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again.
15:08:11 mriedem sdague: i can look in a bit
15:09:13 jaypipes dansmith: do you think we should even attempt to fix the shared storage reporting in Pike?
15:09:27 jaypipes dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1
15:09:39 dansmith jaypipes: well, tbh, I thought that was out the window a while ago
15:09:51 dansmith the computes have no idea which shared storage uuid a given thing is using anyway right?
15:10:06 jaypipes dansmith: right. and ocatas are just going to pummel the allocations either way.
15:11:22 mriedem jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes
15:11:49 jaypipes dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike.
15:11:55 mriedem kind of weird to say, "this feature doesn't yet work" in a known issues section though
15:12:07 mriedem i don't know of omission is better than clearly saying it's not supported
15:12:23 mriedem i'd rather be clear personally
15:12:28 mriedem of what is supported and what's not
15:12:30 jaypipes mriedem: yeah, agreed
15:13:11 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517
15:13:13 jaypipes mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly.
15:13:20 jaypipes dansmith: agreed?
15:13:22 dansmith I dunno why anyone would think it's supported unless you said it was, but okay
15:13:36 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add additional content to admin guide https://review.openstack.org/490952
15:16:59 bauzas mmm, interesting, generation of relnotes fails if you have two reno files sharing the same explicit RST target
15:33:25 jaypipes dansmith: hey, did you catch my question above about agreeing/disagreeing with this:
15:33:27 jaypipes mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly.
15:34:23 dansmith jaypipes: yeah, but that's not a change in course right?
15:34:26 dansmith I mean, that's been the plan?
15:34:59 dansmith it'd be really nice if we were to remove the compute allocations behavior by removing all the legacy RT stuff in the compute node at the same time
15:35:30 dansmith mriedem: in support of that goal, and circling back to the discussion this morning, shouldn't we deprecate chance and caching in pike?
15:36:16 jaypipes dansmith: correct, that was the plan./
15:36:28 dansmith jaypipes: okay then, yea agreed :)
15:37:20 jaypipes dansmith: w.r.t. to deprecating the caching and chance scheduler, I'd be a-ok with that, but that's also why I was quite concerned with the earlier conversation between gibi, me and bauzas where it was posited that our func tests were using the chance scheduler :(
15:37:46 bauzas jaypipes: dansmith: looks a good discussion for Queens, nope ?
15:37:53 bauzas at least for CachingScheduler
15:38:02 dansmith bauzas: very specifically we need to deprecate both now, which is what I just said
15:38:07 bauzas because we could ask operators to try out our placement nice features
15:38:31 bauzas dansmith: if that's all but just sending a signal, I'm okay
15:38:35 dansmith bauzas: it's about deprecating things before we break them entirely
15:38:50 bauzas dansmith: if that's implying that we will remove that in Queens, I dunno
15:39:04 dansmith I just said it was about removing this stuff in queens,
15:39:06 dansmith or at least being able to
15:39:07 jaypipes bauzas: it's about deprecating in Pike.
15:39:13 jaypipes bauzas: and removing in Queens.
15:39:17 dansmith yes.
15:39:27 bauzas tbh, I don't really care of chance scheduler
15:39:39 bauzas I just gave explanation about why people use it
15:39:42 gibi jaypipes: I'm changing the func test env to use filter_scheduler as we speak
15:39:50 gibi jaypipes: it doesn't seems terribly hard to do
15:39:55 efried jamielennox yt?
15:40:01 jaypipes gibi: that's good news indeed.
15:40:09 dansmith gibi: chance with one (fake) host should be pretty much the same behavior I'd hope :)
15:40:12 bauzas I'm a little more concerned by cachingscheduler, which is AFAIK heavely used somewhere
15:40:41 bauzas gibi: I don't expect real breakages tbh
15:40:52 bauzas since tests using chance are just single-host
15:40:57 gibi I think I can push that patch today
15:41:01 jaypipes dansmith: well, the problem is chance essentially disables the placement API in scheduler, which disables claims in the scheduler, which means everything relies on allocations done on the compute node. :(
15:41:31 dansmith jaypipes: ....right... did I say something wrong above?
15:41:47 dansmith or you're just saying that changing that could break functional tests in weird ways?
15:42:11 dansmith if the latter, that's why I said "pretty much the same behavior" :)
15:42:16 dansmith ideally, and hopefully
15:42:52 jaypipes dansmith: well, I was just saying that's the reason I was dismayed to hear from gibi that the chance scheduler was used for almost all the functional tests...
15:43:00 dansmith oh, indeed
15:43:09 jaypipes anyway, sounds like we're in violent agreement, so I'll shut it.
15:43:14 mriedem we could always write a scheduler driver just used in tests - we have the fake driver
15:43:21 mriedem we could change the fake driver to be like the chance driver
15:43:31 jaypipes mriedem: I hope you're kidding. :)
15:43:32 mriedem but best to just use stripped down filter scheduler in functional tests
15:43:43 mriedem note that you can still plug in scheduler drivers
15:43:43 dansmith jaypipes: I think he's kidding
15:43:45 mriedem which won't use placement
15:44:19 jaypipes dansmith: you never know with these South Dakatons.
15:44:24 mriedem so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option
15:44:42 dansmith if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions
15:44:46 dansmith which is fine I guess
15:45:40 mriedem i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those
15:45:41 dansmith like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release
15:46:02 mriedem we removed the ability for the compute manager to be pluggable, or do you mean the virt driver?
15:46:12 dansmith that's my point,
15:46:31 dansmith there's basically zero chance you've written your own compute manager that does everything properly
15:46:55 dansmith and same with scheduler, as we continue to move responsibility and apis around, there's no chance you wrote a scheduler in kilo and are still using it successfully today
15:47:08 openstackgerrit Stephen Finucane proposed openstack/nova master: conf: Move availability zones opts to a group https://review.openstack.org/462469
15:47:10 mriedem well, you just didn't upgrade from kilo L)
15:47:11 mriedem :)
15:47:19 mriedem which is like 50% of deployments probably
15:47:24 bauzas mriedem: by default, we use filter scheduler, but people tend to use cachingscheduler because they just want to not run all filters
15:47:34 bauzas mriedem: I mean, when writing func tests
15:47:41 jaypipes bauzas: hmm?
15:47:41 mriedem huh?
15:47:48 mriedem bauzas: people use caching scheduler for performance
15:47:48 dansmith lol
15:48:03 dansmith he means test writers I think
15:48:04 dansmith I hope.
15:48:05 bauzas mriedem: so if we provide them either a fixture or something easy like dansmith said about just running without filters, that would be sufficient I guess
15:48:08 mriedem i don't think we can remove the caching scheduler until we've shown that the filter scheduler + claims in the scheduler outperforms caching scheduler at scale
15:48:20 bauzas mriedem: jaypipes: huh, s/caching/chance

Earlier   Later