| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-01-29 | |||
| 15:59:04 | gibi | mriedem: OK. Then I'm going to rebase the above patch | |
| 15:59:19 | gibi | mriedem: the backport impossibility was agreed earlier so that is clear | |
| 15:59:28 | jaypipes | maciejjozefczyk: ok. with your operator hat on, what do you think about my proposal, considering it involves a technically not backwards-compatible behaviour change? | |
| 15:59:59 | efried | cdent I think I get that logically, conductor < cluster <=> n-cpu < resource pool < compute hosts (right?) and from that perspective you want to be able to "model" that structure as a tree. | |
| 16:00:32 | cdent | efried, jaypipes: I'd be much better at this conversation if I wasn't still fighting off this cold, but it's not just the desired to model, there's also a desire to do some accounting | |
| 16:00:57 | efried | cdent 1 < many and 1 <=> 1 | |
| 16:01:30 | maciejjozefczyk | jaypipes: From my perspective it would not change anyting, I set allocation_ratio per host not per aggregate, so my opinion seems not be helpful | |
| 16:01:46 | jaypipes | mgagne: hey! :) | |
| 16:02:00 | mgagne | jaypipes: hi! | |
| 16:02:07 | jaypipes | mgagne: we're talking about your favorite new anti-feature regarding host aggregates and allocation ratios :) | |
| 16:02:21 | mgagne | jaypipes: I just saw your email =) | |
| 16:02:39 | jaypipes | mgagne: cool. I'm hoping to commandeer maciejjozefczyk's patch to add this fix for you. | |
| 16:03:01 | jaypipes | of course, I'm trying to convince maciejjozefczyk that this would be of benefit ;) | |
| 16:03:04 | mgagne | jaypipes: I have yet to properly analyze your suggestion and related bug report. | |
| 16:03:22 | jaypipes | mgagne: no worries, would be good to get your feedback some time today. | |
| 16:03:39 | jaypipes | mgagne: I'm going to ping melwitt on it when she's up, too. | |
| 16:03:45 | mgagne | jaypipes: sure | |
| 16:04:00 | sauloaugusto | Is it possible to setup pike without cells ? | |
| 16:04:11 | openstackgerrit | Sen Yang proposed openstack/python-novaclient master: Implement hypervisor hostname exact patten match for server cold migrate https://review.openstack.org/520187 | |
| 16:04:14 | jaypipes | sauloaugusto: nope. | |
| 16:04:56 | efried | cdent So I think where we're getting to is this: You could have your nice satisfying model, and have to hack the inventory distribution to make it work; or you could have a flatter model that's less closely representative of the "real world", but your inventorying is simple, and accounting flows automatically from scheduling/allocations. | |
| 16:05:25 | efried | cdent ...except for the part where vcenter is gonna move an instance from one host to another. Yeah, that gets a tad hairy without having done a real live migration or whatever. | |
| 16:05:40 | sauloaugusto | jaypipes: Is there any clue about the error when you migrate from a no cell env to pike and you can only list instances from admin tenant ? | |
| 16:06:39 | dansmith | sauloaugusto: that doesn't sound like a cells-related thing | |
| 16:07:30 | dansmith | (the cells code doesn't do anything special with users/tenants) | |
| 16:07:33 | cdent | efried: having allocations against physical hosts (instead of the cluster) cannot work because instances move and we've declared the virt driver is not allowed to manipulate allocations | |
| 16:07:56 | efried | cdent Yeah, I see where that's a problem. I don't have an answer for that. | |
| 16:08:08 | cdent | thus inventory on resource pools | |
| 16:08:17 | efried | and resultant hackage. | |
| 16:08:25 | efried | which is unfortunate. | |
| 16:08:32 | maciejjozefczyk | jaypipes mgagne: Need to go. Let me know in bug report about what you decided okey ;) I'll update the patch then | |
| 16:08:33 | cdent | because that is the level at which accounting is desired | |
| 16:08:43 | jaypipes | maciejjozefczyk: will do, thanks mate :) | |
| 16:09:49 | efried | cdent Well, at the very least the resource pools should not be children in the same tree. They should be roots. | |
| 16:10:15 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Rollback instance.image_ref on failed rebuild https://review.openstack.org/538961 | |
| 16:10:16 | mriedem | artom: here is the fix ^ let me know when you have an LP bug number so i can update the references in here | |
| 16:10:33 | artom | mriedem, yep, writing it up as we speak | |
| 16:11:02 | cdent | efried: remind me (because I'm half asleep) of the mechanism to go from a root (when in nova-scheduler) to the compute node where the spawn can be called? | |
| 16:11:14 | sauloaugusto | dansmith: I think that is related , because the error start after I do themigration steps from nova . After that I start to see the Can't upgrade a READER transaction to a WRITER mid-transaction error on nova-api.log . | |
| 16:11:20 | cdent | efried: if the root is in fact not that compute-monde | |
| 16:11:26 | artom | mriedem, so we're rolling back the image but not other things that could change, like keypair? | |
| 16:11:43 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add regression test for bug 1735407 https://review.openstack.org/526095 | |
| 16:11:45 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add late server group policy check to rebuild https://review.openstack.org/525242 | |
| 16:11:45 | openstack | bug 1735407 in OpenStack Compute (nova) "[Nova] Evacuation doesn't respect anti-affinity rules" [Medium,In progress] https://launchpad.net/bugs/1735407 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:11:48 | mriedem | artom: yeah - which was my original complaint in the earlier fix | |
| 16:12:03 | dansmith | sauloaugusto: the migration steps from no cells to one cell were required before pike | |
| 16:12:09 | efried | cdent Not sure we quite have that mechanism yet. | |
| 16:12:13 | mriedem | artom: we can't really roll everything back with what we have today in the conductor code | |
| 16:12:35 | mriedem | we'd either have to change the rpc cast to a call, or move the code that changes the instance from the api to conductor | |
| 16:12:37 | cdent | efried: indeed, thus why I've been modelling it the way I've been | |
| 16:12:40 | efried | cdent Yeah, a way to tie these brother-roots back to the compute host for purposes of allocation and deploy. | |
| 16:12:44 | cdent | if it is a child you know where to go | |
| 16:12:50 | artom | mriedem, yeah... I guess it's fine as a bandaid for now? Ideally we need to come back to this and fix it properly, maybe like what Andrey was proposing | |
| 16:12:50 | mriedem | hell, maybe we should have just moved the instance.save() from api to conductor... | |
| 16:12:52 | cdent | because | |
| 16:12:54 | cdent | it is a child | |
| 16:13:04 | mriedem | artom: yes this is a "tactical fix" | |
| 16:13:13 | cdent | a resource provider withing a compute-node is _obviously_ a child | |
| 16:13:22 | efried | cdent But if it's a child, the scheduler will happily deploy an instance with VCPU from pool1, MEMORY_MB from pool2, and DISK_GB from pool3. | |
| 16:13:24 | cdent | and it seems to me that the ironic model might be the thing that's weird | |
| 16:13:25 | artom | mriedem, heh, code for "someone else can deal with it once I'm retired" ;) | |
| 16:13:38 | efried | cdent Unless you *always* use a numbered granular group. Which doesn't seem like the right answer. | |
| 16:13:43 | mriedem | maybe - it's just something one of my old managers at ibm always said, | |
| 16:13:48 | cfriesen | jaypipes: for https://bugs.launchpad.net/nova/+bug/1742747 do we look at the cpu_allocation_ratio of the first host aggregate, or do we look for the first host aggregate that has cpu_allocation_ratio set? | |
| 16:13:49 | openstack | Launchpad bug 1742747 in OpenStack Compute (nova) "RT overrides default allocation_ratios for ram cpu and disk" [Undecided,In progress] - Assigned to Maciej Jozefczyk (maciej.jozefczyk) | |
| 16:13:54 | mriedem | "do we have a tactical fix while we work on the long-term strategic fix" | |
| 16:14:16 | artom | mriedem, not a bad way of thinking | |
| 16:14:26 | efried | I was doing a word puzzle the other day where the clue was "strategic" and the answer was "tactical". I thought how the IBM ppt-jockeys would flip out at that. | |
| 16:14:38 | artom | awaugama, you around? Since you're the cause of all this, if we build you our proposed tactical fix, you want to try and break it again? | |
| 16:14:39 | sauloaugusto | dansmith: Yes I did that, and I get get list o admin instances , and also create new instances at that . The problem that I can not do nothing at all other tenants . | |
| 16:14:53 | artom | Maybe if we let QE loose on it *before* merging it, we'll avoid the pain? | |
| 16:15:03 | awaugama | artom: I can do some sanity checks on it | |
| 16:15:09 | mriedem | artom: that never seems to happen | |
| 16:15:13 | mriedem | until it's in product | |
| 16:15:24 | awaugama | artom, and *Technically* I'm just the warning light, i'm not the cause of the issue | |
| 16:15:47 | artom | awaugama, hah, sorry, should have added a ;) or /s in there somewher e:) | |
| 16:15:52 | awaugama | ;) | |
| 16:15:58 | cdent | efried: these issues are why I pressed rado to start talking about and experimenting with this stuff. I think the sort of gravitational mechanics of nested providers is going to have lots of weird | |
| 16:16:23 | cdent | and I'd much prefer to see us working from concrete situations then not | |
| 16:16:48 | efried | cdent Agree. At the moment I'm actually leaning towards the thing I originally said was the worst idea: always using a single numbered request group. | |
| 16:17:59 | Spazmotic | Always good when the best idea is your worst idea | |
| 16:21:07 | cdent | efried: I guess we have time to experiement... | |
| 16:25:35 | artom | mriedem, https://bugs.launchpad.net/nova/+bug/1746032 | |
| 16:25:36 | openstack | Launchpad bug 1746032 in OpenStack Compute (nova) "By rebuilding twice with the same "forbidden" image one can circumvent scheduler rebuild restrictions" [Undecided,New] | |
| 16:27:34 | mriedem | artom: L145 https://etherpad.openstack.org/p/nova-ptg-rocky | |
| 16:28:14 | artom | mriedem, oh hey and I'll actually be there this time :) | |
| 16:28:32 | mriedem | we can finally meet, face to face, and settle all the scores | |
| 16:28:38 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Reproduce bug 1724172 in the functional test env https://review.openstack.org/512553 | |
| 16:28:38 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Enhance service restart in functional env https://review.openstack.org/512552 | |
| 16:28:39 | mriedem | there can be only one | |
| 16:28:39 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: cleanup evacuated instances not on hypervisor https://review.openstack.org/512623 | |
| 16:28:39 | openstack | bug 1724172 in OpenStack Compute (nova) "Allocation of an evacuated instance is not cleaned on the source host if instance is not defined on the hypervisor" [Undecided,In progress] https://launchpad.net/bugs/1724172 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:28:49 | artom | ... one of *what*? | |
| 16:29:00 | mriedem | idk, i'm just thinking about the quickening | |
| 16:29:04 | mriedem | even though that was scotland | |
| 16:29:29 | artom | Man, your references are out of control! | |
| 16:31:40 | edleafe | cdent: I've been following along, but don't have a clear picture of how things are arranged. Is there a diagram or something that would illustrate this? | |
| 16:32:22 | cdent | edleafe: a) which "this"?, b) I wish I never showed up today because I simply don't have the capacity to actual engage well with this topic | |
| 16:32:31 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Rollback instance.image_ref on failed rebuild https://review.openstack.org/538961 | |
| 16:32:32 | mriedem | artom: ok here we go, single patch ^ | |