Earlier  
Posted Nick Remark
#openstack-nova - 2018-01-29
15:13:42 cdent the intent is to deploy to a specific resource pool
15:14:12 cdent if you have an aggregate for each specific resource pool, that might be workabe
15:14:30 efried But you still can't deploy to a specific resource pool with just that.
15:14:42 cdent brb
15:14:44 efried You would have to have some other mechanism in place.
15:14:45 mriedem can you add a custom filter and custom scheduler hint which is just 'make it go directly here'
15:15:10 jaypipes efried: uhm, yeah you could request a specific aggregate...
15:15:19 efried jaypipes How?
15:15:28 efried With some mechanism yet to be designed?
15:15:46 jaypipes efried: we already have the member_of param. we can work that in to alloc candidates very easily.
15:16:08 efried okay, agree, but not yet designed or started.
15:16:23 jaypipes efried: of course, that would require the users knowing about the aggregate UUID/names...
15:16:31 jaypipes efried: which generally we have not encouraged.
15:16:41 mriedem stephenfin: got it
15:16:42 mriedem thanks
15:17:14 mriedem there is no aggregate name https://developer.openstack.org/api-ref/placement/#list-resource-provider-aggregates
15:17:33 jaypipes mriedem: sorry, UUID...
15:17:58 jaypipes efried: still, what is the root of the problem that is trying to be solved?
15:18:32 cdent It remains unclear to me why (eventually) nested isn't the ideal solution for this stuff? A compute-node is an rp, it cotains resource pool rps and maybe even some exsi rp, a hybrd representing the stuff managed by the conpute node
15:18:39 jaypipes efried: I mean, are we doing scheduling here or are we just doing resource tracking and vcenter basically takes over the scheduling/placement decisions (that's what I *thought* the situation was currently...)
15:18:59 mriedem sounds like the latter
15:19:13 stephenfin mriedem: Is this something that should wait for a few weeks? https://review.openstack.org/#/c/534371/
15:19:38 mriedem stephenfin: is that just cleanup? if so, yes
15:19:45 efried sorry guys, pulled into a quick meeting, back in a bit.
15:19:52 mriedem stephenfin: we're in bug fix only mode at this point
15:19:53 stephenfin mriedem: (y)
15:19:55 mriedem or tests, docs, etc
15:22:04 cdent jaypipes: the underlying feature that rado wants to expose here are resource pools and he wants to do it in a way that allows placement to be aware of them
15:22:52 jaypipes cdent: will the *end user* be capable of a) understanding what these resource pools are and b) asking for a particular resource pool to be used by an instance?
15:23:23 cdent jaypipes: yes. "tenant virtual data center"
15:23:43 stephenfin lyarwood (or another stable-core): fancy hitting these? https://review.openstack.org/#/c/538939/ https://review.openstack.org/#/c/538938/
15:24:02 jaypipes cdent: why not have multiple nova-compute services, each managing a single "resource pool"?
15:26:44 cdent because the clusters behind a nova-compute service are intentionally dynamic: it's desirable to be able to create destroy and change resource pools in the cluster without messing with the openstack control plane much
15:27:08 efried I'm back.
15:27:56 efried cdent Do these clusters / resource pools (are those different things) contain single compute-ish nodes to which instances get deployed, OR do they contain some general pool of resources from which an instance could be deployed using some resource from one member and some from another?
15:28:02 jaypipes cdent: that's problematic for many reasons, as you know.
15:29:44 cdent jaypipes: yes, but problematic is how we push the ball forward, no? efried cluster and resource pools are different, yes. logically an instance is deployed to a resource pool and not a specific host (although physically they are, of course)
15:30:01 cdent a cluster contains 1 or more resource pools
15:30:29 cdent a resource pool has access to some (dynamic) subsection of a clusters inventory
15:31:24 efried Are physical hosts the underlying members of a resource pool?
15:32:25 cdent efried: that's not really the way it is modeled, no, but I'm not sure what you're actually asking
15:32:42 efried cdent I'm just not understanding where the physical resources are here.
15:32:44 jaypipes cdent: what does the vmware virt driver *currently* do? does it report the entire amount of VCPU, memory, etc for all resource pools and all clusters that vCenter knows about?
15:33:27 jaypipes cdent: and the scheduler always picks the single nova-compute service daemon and basically all scheduling/placement is simply deferred to the vCenter API?
15:33:35 cdent right now: the vcenter represents one cluster, and reports all vcpu, memory, etc that it knows abouts
15:34:32 cdent but there are a multiple nova-copute services, each for a given cluster
15:34:50 jaypipes cdent: I'm not seeing the benefit of modeling things using placement's resource providers data modeling if deployments using vCenter won't actually end up using any of that information when doing placement/scheduling decisions anyway.
15:37:40 cdent jaypipes: two reasons a) if you have multiple clusters, and thus multiple nova-compues, and those clusters differet in various ways (for example by having different tenant virtual data centers) you need to scheduler differently, b) despite our troubles, we've made a very good resource/provider/inventory/etc modelling system with placement and it is a good place to track and select where to place things
15:38:35 cdent "this cluster over there is the one with all the GPUs"
15:38:43 efried cdent It'll help me a lot to understand what is underneath a resource pool. If a resource pool reports a total of 8 VCPUs available, does that mean I am guaranteed to be able to satisfy a deploy requesting 8 VCPUs? (All other things being available)
15:39:27 efried Cause I got the impression that those 8 VCPUs might be spread across say four physical hosts, and an instance could only go to one of those, so the biggest one you could handle might only be 2 VCPUs or whatever.
15:40:20 jaypipes don't see the point.
15:40:20 jaypipes cdent: for a) if the end user is just saying "I'm in tenant X, so launch an instance on virtual data center X", that is a stupid-simple scheduling decision that needs nothing other than a single trait (CUSTOM_TENANT_X decorates a single resource provider representing the nova-compute service that manages that tenant virtual data center). For b) again, if vCenter isn't going to use that information in its intra-cluster scheduling decisions, I
15:42:02 jaypipes cdent: another question for you..
15:42:22 cdent jaypipes: can I pause you for a minute while I answer efried, then your current question?
15:43:11 jaypipes cdent: even if placement ended up getting proper inventory information for these vCenter clusters, would VMWare OpenStack users actually be able looking at that inventory information? Or would they be using VMWare GUIs to look at inventory information for their specific tenant virtual datacenter?
15:43:19 jaypipes cdent: of course, no worries
15:44:42 cdent efried: a while back there were some adjustment made to the vmware driver to make sure that the max_unit was no greater than what any individual host could provide. so your second statement about "the biggest one" is an issue in some cases, but it's squishy because of the DRS (dynamic resource scheduler thing which can do unattended live migrations and what not, magic stuff I haven't dug into)
15:46:24 efried So it is in fact the case that there are some number of physical hosts behind a resource pool, and an instance can only be deployed to one, not spread out across more than one. True statement?
15:46:26 cdent efried: so, yes, there is a fair bit of complexity which is _why_ being able to manipulate the entire provider tree from within the virt driver has been of such interest to rado (and me).
15:46:49 jaypipes maciejjozefczyk: you around? wanted to chat with you about https://bugs.launchpad.net/nova/+bug/1742747
15:46:50 openstack Launchpad bug 1742747 in OpenStack Compute (nova) "RT overrides default allocation_ratios for ram cpu and disk" [Undecided,In progress] - Assigned to Maciej Jozefczyk (maciej.jozefczyk)
15:46:56 cdent efried: it can move and you can never know that it did
15:47:52 efried "it" the physical host? "can move" to another resource pool? and "you" the virt driver or you the operator?
15:48:10 jaypipes cdent: with the shared-memory solution, can't VMs literally be on multiple hypervisor hosts at once?
15:48:11 maciejjozefczyk jaypipes: yes, hey!
15:48:33 jaypipes maciejjozefczyk: you see my note on that bug?
15:49:25 maciejjozefczyk jaypipes: reading
15:49:40 cdent jaypipes: I've done my damnedest thus far to stay out of the guts of what vcenter gets up to, but there's plenty of magical knobs, but I'm trying to avoid knowing about them. there's already too many things to know
15:49:57 jaypipes cdent: ack
15:50:03 cdent efried: it is the instance
15:50:37 efried okay. "move" where, inter-pool or just intra-pool?
15:50:46 efried not that it actually matters.
15:51:19 maciejjozefczyk jaypipes: so we should also check aggregation metadatas if there is already definition of allocation ratio defined?
15:51:42 maciejjozefczyk if zeros are in config?
15:51:50 jaypipes maciejjozefczyk: yeah. check out that operators mailing list thread I linked in there...
15:52:07 cdent efried: intra-cluster.
15:52:24 efried t that later.)
15:52:24 efried cdent So the thing is, we didn't have a way to model multiple providers before. Now we do. But it sounds to me like nested is *not* the multiple-provider model we want to use here. It just doesn't make sense in light of recently-merged capabilities for a single provider *or* tree to "own" inventory spanning more than one actual compute host. (Unless, as jaypipes implies, there's a memory sharing thing - we can talk abou
15:52:47 openstackgerrit Matt Riedemann proposed openstack/nova master: WIP: Add regression test wrinkle for rebuilding twice with invalid image https://review.openstack.org/538955
15:52:49 mriedem artom: ^
15:54:14 efried cdent If you do it in a single provider OR a tree, you'll always have to reconcile a request for some subset of the total resources that you can't actually satisfy from any one host.
15:54:16 artom mriedem, I think you'd have to get the server out of ERROR first
15:55:02 efried cdent With the single-provider model, which I gather is how it has needed to be done to this point, it sounds like some of that reconciliation was that max-unit hack.
15:55:20 jaypipes efried: yeah. I think it's a much better solution for vCenter to model their resource tracking in the vCenter virt driver in a similar fashion to how Ironic's virt driver works.
15:55:56 cdent efried: I have no issue with the inventory for hardware things being pushed all the way down the real hardware, but I _do_ think that representing the hierarchical relationships between things using a tool we've created that has hierarchy makes sense (in this case cluster->resource-pool->physical host)
15:56:27 jaypipes efried: that wouldn't *preclude* using nested providers if, say, certain compute nodes had PCI devices or NUMA nodes, etc. But I would use a new root provider for each compute node in the cluster and represent the cluster as a placement aggregate, not as a child node under a "fake root"
15:56:42 efried jaypipes ++ exactly
15:57:01 jaypipes cdent: I think we're mostly agreeing. just disagreeing on the "single root vs multiple roots" thing.
15:57:09 gibi mriedem: quick question: There is a bugfix https://review.openstack.org/#/c/525242 that needs rpc version bump. Is it something that needs to wait for Rocky?
15:57:24 efried cdent A placement tree is more than just a way to represent a hierarchy. It is also allows resource requests to span a tree. Which doesn't make sense in the vcenter model as I'm undestanding it.
15:57:52 mriedem artom: you can rebuild an ERROR server
15:58:05 cdent efried: then I think it probably hasn't been explained well enough
15:58:14 artom mriedem, hah, that makes it even easier to hit
15:58:29 mriedem gibi: that doesn't necessarily mean it has to wait for rocky,
15:58:35 mriedem gibi: but it does mean we wouldn't backport it
15:58:39 maciejjozefczyk jaypipes: I'm ok with the idea to check aggregation metadatas first
15:59:04 gibi mriedem: OK. Then I'm going to rebase the above patch
15:59:19 gibi mriedem: the backport impossibility was agreed earlier so that is clear
15:59:28 jaypipes maciejjozefczyk: ok. with your operator hat on, what do you think about my proposal, considering it involves a technically not backwards-compatible behaviour change?

Earlier   Later