| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-07 | |||
| 14:58:04 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove dead function, variable https://review.openstack.org/491512 | |
| 14:59:00 | jaypipes | mriedem: so, w.r.t. the problem of shared storage, I think instead of adding the type of code you added in your patch, maybe we should just wait for Queens when we can assume no compute hosts are attempting to heal allocations in the compute node. | |
| 14:59:00 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove some single use function from utils https://review.openstack.org/491513 | |
| 14:59:12 | jaypipes | mriedem: I'd like dansmith's thoughts on that, though. | |
| 14:59:40 | mriedem | jaypipes: i don't think we can claim support for shared storage at this point regardless | |
| 14:59:45 | mriedem | we're basically implementing a feature now | |
| 14:59:56 | mriedem | which is why i un-completed the shared storage bp | |
| 15:00:07 | dansmith | lennyb: that's not a lost connection to the cell database, that's a lack of configuration of some node for the api database. are you running conductors on your second (non-allinone) node? | |
| 15:00:25 | jaypipes | mriedem: another thing I've been thinking of is instead of further mucking with the resource tracker and/or making the scheduler report client that we do a nova-status or nova-manage audit function that cleans up allocations based on the cell DB states. | |
| 15:01:28 | jaypipes | mriedem: and once done, we have all computes on service version 22 (or whatever) and after that point, we remove all allocation stuff from the compute node entirely (other than the aforementioned delete allocations pieces) | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove files from 'tools' https://review.openstack.org/491444 | |
| 15:02:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Remove "vif" script https://review.openstack.org/491443 | |
| 15:03:05 | lennyb | dansmith, thanks, I will recheck it. | |
| 15:03:08 | jaypipes | mriedem: because, if I'm being frank, the report client is starting to look like the mess of spaghetti code in the resource tracker that placement was supposed to simplify. :( | |
| 15:03:09 | mriedem | jaypipes: the service version won't change until the actual code is upgraded though | |
| 15:03:18 | mriedem | jaypipes: i agree with that | |
| 15:03:38 | dansmith | is it alternate-ego mondays? why is jaypipes being "Frank" ? | |
| 15:03:45 | jaypipes | heh | |
| 15:05:04 | jaypipes | mriedem: no, I understand that about the service version. what I'm suggesting is that we add a patch that does the "reconcile allocations by looking at cell DB state", we remove all allocation healing from the nova-computes (with a service version bump) and then we have the upgrade process be a run of the nova-manage/nova-status command that audits allocation information. | |
| 15:06:11 | jaypipes | mriedem: the alternative is all the mess of conditionals and such that we need to add to the resource tracker and report client. | |
| 15:06:14 | dansmith | jaypipes: so that means allocations are just all out of whack until you finish upgrading all your ocata computes and then run this fixer-upper thing? | |
| 15:06:19 | jaypipes | mriedem: choose your poison I think? | |
| 15:06:27 | jaypipes | dansmith: yeah. | |
| 15:06:29 | sdague | mriedem (or others) anyone want to approve these new docs sub pages - https://review.openstack.org/#/c/490994/ ? I'd like to get those 2 in before doing more to prevent wasted work | |
| 15:06:42 | dansmith | jaypipes: that would be really unfortunate, IMHO | |
| 15:06:49 | dansmith | jaypipes: some people take a looong time to do that upgrade | |
| 15:06:50 | jaypipes | dansmith: as opposed to being out of whack repeatedly with pike computes fixing ocata data over and over again. | |
| 15:08:11 | mriedem | sdague: i can look in a bit | |
| 15:09:13 | jaypipes | dansmith: do you think we should even attempt to fix the shared storage reporting in Pike? | |
| 15:09:27 | jaypipes | dansmith: aka mriedem's https://review.openstack.org/#/c/491098/1 | |
| 15:09:39 | dansmith | jaypipes: well, tbh, I thought that was out the window a while ago | |
| 15:09:51 | dansmith | the computes have no idea which shared storage uuid a given thing is using anyway right? | |
| 15:10:06 | jaypipes | dansmith: right. and ocatas are just going to pummel the allocations either way. | |
| 15:11:22 | mriedem | jaypipes: dansmith: i'm resigned to give up on shared storage for pike, i'll put a known issue in the release notes | |
| 15:11:49 | jaypipes | dansmith, mriedem: so maybe the best we can hope for in Pike is to land https://review.openstack.org/#/c/491012/ (the "stop healing allocations if all Pike" patch), just accept shared storage doesn't work and tell operators not to create shared providers in Pike. | |
| 15:11:55 | mriedem | kind of weird to say, "this feature doesn't yet work" in a known issues section though | |
| 15:12:07 | mriedem | i don't know of omission is better than clearly saying it's not supported | |
| 15:12:23 | mriedem | i'd rather be clear personally | |
| 15:12:28 | mriedem | of what is supported and what's not | |
| 15:12:30 | jaypipes | mriedem: yeah, agreed | |
| 15:13:11 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Address review comments for contributor index https://review.openstack.org/491517 | |
| 15:13:13 | jaypipes | mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly. | |
| 15:13:20 | jaypipes | dansmith: agreed? | |
| 15:13:22 | dansmith | I dunno why anyone would think it's supported unless you said it was, but okay | |
| 15:13:36 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Add additional content to admin guide https://review.openstack.org/490952 | |
| 15:16:59 | bauzas | mmm, interesting, generation of relnotes fails if you have two reno files sharing the same explicit RST target | |
| 15:33:25 | jaypipes | dansmith: hey, did you catch my question above about agreeing/disagreeing with this: | |
| 15:33:27 | jaypipes | mriedem: and at least, first thing in Queens we can get rid of the auto-healing ocata code, assume all computes are no longer auto-healing allocations and fix shared providers properly. | |
| 15:34:23 | dansmith | jaypipes: yeah, but that's not a change in course right? | |
| 15:34:26 | dansmith | I mean, that's been the plan? | |
| 15:34:59 | dansmith | it'd be really nice if we were to remove the compute allocations behavior by removing all the legacy RT stuff in the compute node at the same time | |
| 15:35:30 | dansmith | mriedem: in support of that goal, and circling back to the discussion this morning, shouldn't we deprecate chance and caching in pike? | |
| 15:36:16 | jaypipes | dansmith: correct, that was the plan./ | |
| 15:36:28 | dansmith | jaypipes: okay then, yea agreed :) | |
| 15:37:20 | jaypipes | dansmith: w.r.t. to deprecating the caching and chance scheduler, I'd be a-ok with that, but that's also why I was quite concerned with the earlier conversation between gibi, me and bauzas where it was posited that our func tests were using the chance scheduler :( | |
| 15:37:46 | bauzas | jaypipes: dansmith: looks a good discussion for Queens, nope ? | |
| 15:37:53 | bauzas | at least for CachingScheduler | |
| 15:38:02 | dansmith | bauzas: very specifically we need to deprecate both now, which is what I just said | |
| 15:38:07 | bauzas | because we could ask operators to try out our placement nice features | |
| 15:38:31 | bauzas | dansmith: if that's all but just sending a signal, I'm okay | |
| 15:38:35 | dansmith | bauzas: it's about deprecating things before we break them entirely | |
| 15:38:50 | bauzas | dansmith: if that's implying that we will remove that in Queens, I dunno | |
| 15:39:04 | dansmith | I just said it was about removing this stuff in queens, | |
| 15:39:06 | dansmith | or at least being able to | |
| 15:39:07 | jaypipes | bauzas: it's about deprecating in Pike. | |
| 15:39:13 | jaypipes | bauzas: and removing in Queens. | |
| 15:39:17 | dansmith | yes. | |
| 15:39:27 | bauzas | tbh, I don't really care of chance scheduler | |
| 15:39:39 | bauzas | I just gave explanation about why people use it | |
| 15:39:42 | gibi | jaypipes: I'm changing the func test env to use filter_scheduler as we speak | |
| 15:39:50 | gibi | jaypipes: it doesn't seems terribly hard to do | |
| 15:39:55 | efried | jamielennox yt? | |
| 15:40:01 | jaypipes | gibi: that's good news indeed. | |
| 15:40:09 | dansmith | gibi: chance with one (fake) host should be pretty much the same behavior I'd hope :) | |
| 15:40:12 | bauzas | I'm a little more concerned by cachingscheduler, which is AFAIK heavely used somewhere | |
| 15:40:41 | bauzas | gibi: I don't expect real breakages tbh | |
| 15:40:52 | bauzas | since tests using chance are just single-host | |
| 15:40:57 | gibi | I think I can push that patch today | |
| 15:41:01 | jaypipes | dansmith: well, the problem is chance essentially disables the placement API in scheduler, which disables claims in the scheduler, which means everything relies on allocations done on the compute node. :( | |
| 15:41:31 | dansmith | jaypipes: ....right... did I say something wrong above? | |
| 15:41:47 | dansmith | or you're just saying that changing that could break functional tests in weird ways? | |
| 15:42:11 | dansmith | if the latter, that's why I said "pretty much the same behavior" :) | |
| 15:42:16 | dansmith | ideally, and hopefully | |
| 15:42:52 | jaypipes | dansmith: well, I was just saying that's the reason I was dismayed to hear from gibi that the chance scheduler was used for almost all the functional tests... | |
| 15:43:00 | dansmith | oh, indeed | |
| 15:43:09 | jaypipes | anyway, sounds like we're in violent agreement, so I'll shut it. | |
| 15:43:14 | mriedem | we could always write a scheduler driver just used in tests - we have the fake driver | |
| 15:43:21 | mriedem | we could change the fake driver to be like the chance driver | |
| 15:43:31 | jaypipes | mriedem: I hope you're kidding. :) | |
| 15:43:32 | mriedem | but best to just use stripped down filter scheduler in functional tests | |
| 15:43:43 | mriedem | note that you can still plug in scheduler drivers | |
| 15:43:43 | dansmith | jaypipes: I think he's kidding | |
| 15:43:45 | mriedem | which won't use placement | |
| 15:44:19 | jaypipes | dansmith: you never know with these South Dakatons. | |
| 15:44:24 | mriedem | so if we deprecated caching and chance in pike b/c of removing allocations in the computes in queens, we should also deprecate the option | |
| 15:44:42 | dansmith | if we're moving the responsibility of resource tracking and claiming to the scheduler, then that just means a driver that doesn't use placement should use something else to make sure it's making reasonable decisions | |
| 15:44:46 | dansmith | which is fine I guess | |
| 15:45:40 | mriedem | i think it would be ok to deprecate caching and chance as a signaling mechanism that new deployments shouldn't use those | |
| 15:45:41 | dansmith | like compute manager being pluggable, also think it's unlikely anyone has written their own scheduler that doesn't break every release | |
| 15:46:02 | mriedem | we removed the ability for the compute manager to be pluggable, or do you mean the virt driver? | |