| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-01-29 | |||
| 15:50:03 | cdent | efried: it is the instance | |
| 15:50:37 | efried | okay. "move" where, inter-pool or just intra-pool? | |
| 15:50:46 | efried | not that it actually matters. | |
| 15:51:19 | maciejjozefczyk | jaypipes: so we should also check aggregation metadatas if there is already definition of allocation ratio defined? | |
| 15:51:42 | maciejjozefczyk | if zeros are in config? | |
| 15:51:50 | jaypipes | maciejjozefczyk: yeah. check out that operators mailing list thread I linked in there... | |
| 15:52:07 | cdent | efried: intra-cluster. | |
| 15:52:24 | efried | cdent So the thing is, we didn't have a way to model multiple providers before. Now we do. But it sounds to me like nested is *not* the multiple-provider model we want to use here. It just doesn't make sense in light of recently-merged capabilities for a single provider *or* tree to "own" inventory spanning more than one actual compute host. (Unless, as jaypipes implies, there's a memory sharing thing - we can talk abou | |
| 15:52:24 | efried | t that later.) | |
| 15:52:47 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Add regression test wrinkle for rebuilding twice with invalid image https://review.openstack.org/538955 | |
| 15:52:49 | mriedem | artom: ^ | |
| 15:54:14 | efried | cdent If you do it in a single provider OR a tree, you'll always have to reconcile a request for some subset of the total resources that you can't actually satisfy from any one host. | |
| 15:54:16 | artom | mriedem, I think you'd have to get the server out of ERROR first | |
| 15:55:02 | efried | cdent With the single-provider model, which I gather is how it has needed to be done to this point, it sounds like some of that reconciliation was that max-unit hack. | |
| 15:55:20 | jaypipes | efried: yeah. I think it's a much better solution for vCenter to model their resource tracking in the vCenter virt driver in a similar fashion to how Ironic's virt driver works. | |
| 15:55:56 | cdent | efried: I have no issue with the inventory for hardware things being pushed all the way down the real hardware, but I _do_ think that representing the hierarchical relationships between things using a tool we've created that has hierarchy makes sense (in this case cluster->resource-pool->physical host) | |
| 15:56:27 | jaypipes | efried: that wouldn't *preclude* using nested providers if, say, certain compute nodes had PCI devices or NUMA nodes, etc. But I would use a new root provider for each compute node in the cluster and represent the cluster as a placement aggregate, not as a child node under a "fake root" | |
| 15:56:42 | efried | jaypipes ++ exactly | |
| 15:57:01 | jaypipes | cdent: I think we're mostly agreeing. just disagreeing on the "single root vs multiple roots" thing. | |
| 15:57:09 | gibi | mriedem: quick question: There is a bugfix https://review.openstack.org/#/c/525242 that needs rpc version bump. Is it something that needs to wait for Rocky? | |
| 15:57:24 | efried | cdent A placement tree is more than just a way to represent a hierarchy. It is also allows resource requests to span a tree. Which doesn't make sense in the vcenter model as I'm undestanding it. | |
| 15:57:52 | mriedem | artom: you can rebuild an ERROR server | |
| 15:58:05 | cdent | efried: then I think it probably hasn't been explained well enough | |
| 15:58:14 | artom | mriedem, hah, that makes it even easier to hit | |
| 15:58:29 | mriedem | gibi: that doesn't necessarily mean it has to wait for rocky, | |
| 15:58:35 | mriedem | gibi: but it does mean we wouldn't backport it | |
| 15:58:39 | maciejjozefczyk | jaypipes: I'm ok with the idea to check aggregation metadatas first | |
| 15:59:04 | gibi | mriedem: OK. Then I'm going to rebase the above patch | |
| 15:59:19 | gibi | mriedem: the backport impossibility was agreed earlier so that is clear | |
| 15:59:28 | jaypipes | maciejjozefczyk: ok. with your operator hat on, what do you think about my proposal, considering it involves a technically not backwards-compatible behaviour change? | |
| 15:59:59 | efried | cdent I think I get that logically, conductor < cluster <=> n-cpu < resource pool < compute hosts (right?) and from that perspective you want to be able to "model" that structure as a tree. | |
| 16:00:32 | cdent | efried, jaypipes: I'd be much better at this conversation if I wasn't still fighting off this cold, but it's not just the desired to model, there's also a desire to do some accounting | |
| 16:00:57 | efried | cdent 1 < many and 1 <=> 1 | |
| 16:01:30 | maciejjozefczyk | jaypipes: From my perspective it would not change anyting, I set allocation_ratio per host not per aggregate, so my opinion seems not be helpful | |
| 16:01:46 | jaypipes | mgagne: hey! :) | |
| 16:02:00 | mgagne | jaypipes: hi! | |
| 16:02:07 | jaypipes | mgagne: we're talking about your favorite new anti-feature regarding host aggregates and allocation ratios :) | |
| 16:02:21 | mgagne | jaypipes: I just saw your email =) | |
| 16:02:39 | jaypipes | mgagne: cool. I'm hoping to commandeer maciejjozefczyk's patch to add this fix for you. | |
| 16:03:01 | jaypipes | of course, I'm trying to convince maciejjozefczyk that this would be of benefit ;) | |
| 16:03:04 | mgagne | jaypipes: I have yet to properly analyze your suggestion and related bug report. | |
| 16:03:22 | jaypipes | mgagne: no worries, would be good to get your feedback some time today. | |
| 16:03:39 | jaypipes | mgagne: I'm going to ping melwitt on it when she's up, too. | |
| 16:03:45 | mgagne | jaypipes: sure | |
| 16:04:00 | sauloaugusto | Is it possible to setup pike without cells ? | |
| 16:04:11 | openstackgerrit | Sen Yang proposed openstack/python-novaclient master: Implement hypervisor hostname exact patten match for server cold migrate https://review.openstack.org/520187 | |
| 16:04:14 | jaypipes | sauloaugusto: nope. | |
| 16:04:56 | efried | cdent So I think where we're getting to is this: You could have your nice satisfying model, and have to hack the inventory distribution to make it work; or you could have a flatter model that's less closely representative of the "real world", but your inventorying is simple, and accounting flows automatically from scheduling/allocations. | |
| 16:05:25 | efried | cdent ...except for the part where vcenter is gonna move an instance from one host to another. Yeah, that gets a tad hairy without having done a real live migration or whatever. | |
| 16:05:40 | sauloaugusto | jaypipes: Is there any clue about the error when you migrate from a no cell env to pike and you can only list instances from admin tenant ? | |
| 16:06:39 | dansmith | sauloaugusto: that doesn't sound like a cells-related thing | |
| 16:07:30 | dansmith | (the cells code doesn't do anything special with users/tenants) | |
| 16:07:33 | cdent | efried: having allocations against physical hosts (instead of the cluster) cannot work because instances move and we've declared the virt driver is not allowed to manipulate allocations | |
| 16:07:56 | efried | cdent Yeah, I see where that's a problem. I don't have an answer for that. | |
| 16:08:08 | cdent | thus inventory on resource pools | |
| 16:08:17 | efried | and resultant hackage. | |
| 16:08:25 | efried | which is unfortunate. | |
| 16:08:32 | maciejjozefczyk | jaypipes mgagne: Need to go. Let me know in bug report about what you decided okey ;) I'll update the patch then | |
| 16:08:33 | cdent | because that is the level at which accounting is desired | |
| 16:08:43 | jaypipes | maciejjozefczyk: will do, thanks mate :) | |
| 16:09:49 | efried | cdent Well, at the very least the resource pools should not be children in the same tree. They should be roots. | |
| 16:10:15 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Rollback instance.image_ref on failed rebuild https://review.openstack.org/538961 | |
| 16:10:16 | mriedem | artom: here is the fix ^ let me know when you have an LP bug number so i can update the references in here | |
| 16:10:33 | artom | mriedem, yep, writing it up as we speak | |
| 16:11:02 | cdent | efried: remind me (because I'm half asleep) of the mechanism to go from a root (when in nova-scheduler) to the compute node where the spawn can be called? | |
| 16:11:14 | sauloaugusto | dansmith: I think that is related , because the error start after I do themigration steps from nova . After that I start to see the Can't upgrade a READER transaction to a WRITER mid-transaction error on nova-api.log . | |
| 16:11:20 | cdent | efried: if the root is in fact not that compute-monde | |
| 16:11:26 | artom | mriedem, so we're rolling back the image but not other things that could change, like keypair? | |
| 16:11:43 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add regression test for bug 1735407 https://review.openstack.org/526095 | |
| 16:11:45 | openstack | bug 1735407 in OpenStack Compute (nova) "[Nova] Evacuation doesn't respect anti-affinity rules" [Medium,In progress] https://launchpad.net/bugs/1735407 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:11:45 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add late server group policy check to rebuild https://review.openstack.org/525242 | |
| 16:11:48 | mriedem | artom: yeah - which was my original complaint in the earlier fix | |
| 16:12:03 | dansmith | sauloaugusto: the migration steps from no cells to one cell were required before pike | |
| 16:12:09 | efried | cdent Not sure we quite have that mechanism yet. | |
| 16:12:13 | mriedem | artom: we can't really roll everything back with what we have today in the conductor code | |
| 16:12:35 | mriedem | we'd either have to change the rpc cast to a call, or move the code that changes the instance from the api to conductor | |
| 16:12:37 | cdent | efried: indeed, thus why I've been modelling it the way I've been | |
| 16:12:40 | efried | cdent Yeah, a way to tie these brother-roots back to the compute host for purposes of allocation and deploy. | |
| 16:12:44 | cdent | if it is a child you know where to go | |
| 16:12:50 | mriedem | hell, maybe we should have just moved the instance.save() from api to conductor... | |
| 16:12:50 | artom | mriedem, yeah... I guess it's fine as a bandaid for now? Ideally we need to come back to this and fix it properly, maybe like what Andrey was proposing | |
| 16:12:52 | cdent | because | |
| 16:12:54 | cdent | it is a child | |
| 16:13:04 | mriedem | artom: yes this is a "tactical fix" | |
| 16:13:13 | cdent | a resource provider withing a compute-node is _obviously_ a child | |
| 16:13:22 | efried | cdent But if it's a child, the scheduler will happily deploy an instance with VCPU from pool1, MEMORY_MB from pool2, and DISK_GB from pool3. | |
| 16:13:24 | cdent | and it seems to me that the ironic model might be the thing that's weird | |
| 16:13:25 | artom | mriedem, heh, code for "someone else can deal with it once I'm retired" ;) | |
| 16:13:38 | efried | cdent Unless you *always* use a numbered granular group. Which doesn't seem like the right answer. | |
| 16:13:43 | mriedem | maybe - it's just something one of my old managers at ibm always said, | |
| 16:13:48 | cfriesen | jaypipes: for https://bugs.launchpad.net/nova/+bug/1742747 do we look at the cpu_allocation_ratio of the first host aggregate, or do we look for the first host aggregate that has cpu_allocation_ratio set? | |
| 16:13:49 | openstack | Launchpad bug 1742747 in OpenStack Compute (nova) "RT overrides default allocation_ratios for ram cpu and disk" [Undecided,In progress] - Assigned to Maciej Jozefczyk (maciej.jozefczyk) | |
| 16:13:54 | mriedem | "do we have a tactical fix while we work on the long-term strategic fix" | |
| 16:14:16 | artom | mriedem, not a bad way of thinking | |
| 16:14:26 | efried | I was doing a word puzzle the other day where the clue was "strategic" and the answer was "tactical". I thought how the IBM ppt-jockeys would flip out at that. | |
| 16:14:38 | artom | awaugama, you around? Since you're the cause of all this, if we build you our proposed tactical fix, you want to try and break it again? | |
| 16:14:39 | sauloaugusto | dansmith: Yes I did that, and I get get list o admin instances , and also create new instances at that . The problem that I can not do nothing at all other tenants . | |
| 16:14:53 | artom | Maybe if we let QE loose on it *before* merging it, we'll avoid the pain? | |
| 16:15:03 | awaugama | artom: I can do some sanity checks on it | |
| 16:15:09 | mriedem | artom: that never seems to happen | |