| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-03-26 | |||
| 16:26:42 | bauzas | that's all about traits and step_size | |
| 16:26:45 | bauzas | but | |
| 16:26:59 | bauzas | if we go decide to provide a tree of NUMA nodes | |
| 16:27:35 | bauzas | but we don't support yet hugepages, then we have a problem because flavors asking for hugepages would get NoValidHosts | |
| 16:28:01 | bauzas | we somehow still need to accept compute nodes to report their memory the old way | |
| 16:28:09 | efried | what's a hugepage? | |
| 16:28:17 | efried | (use little words) | |
| 16:28:26 | bauzas | https://docs.openstack.org/nova/latest/admin/huge-pages.html | |
| 16:28:39 | bauzas | it's a defined memory page size | |
| 16:28:48 | bauzas | but that's just an example | |
| 16:28:57 | bauzas | for the moment, placement only checks the total memory of the host, right? | |
| 16:29:12 | bauzas | then the scheduler filter does magic things to find a destination | |
| 16:29:40 | efried | Well, yeah, so this is kind of in line with what we need to do wrt bandwidth. | |
| 16:29:52 | efried | We represent the number of huge pages as inventory alongside the MEMORY_MB inventory. | |
| 16:29:53 | efried | but | |
| 16:30:12 | efried | if the consumer is going to include hugepages as part of their request, they *always* need to do so. | |
| 16:30:27 | bauzas | so, my concern is that if I'm beginning to shard the memory between NUMA nodes, then it requires the flavors to be updated to explicitly ask for a NUMA node, whereas huge pages are totally NUMA unrelated | |
| 16:30:59 | efried | because what we can't (or shouldn't) do is try to convert a request for MEMORY_MB:4096 into MEMORY_MB:4096,HUGEPAGES:4 (or whatever) | |
| 16:31:26 | bauzas | efried: I'm not trying to design now how to make placement queries for hugepages | |
| 16:31:29 | efried | bauzas: But does a huge page come from the same place as a MEMORY_MB or is it a separate thing? | |
| 16:31:46 | bauzas | efried: what I'm trying is to make sure we keep a compatible behaviour for the existing feature | |
| 16:32:05 | bauzas | others, later, will try to solve that design and use placement resources for that | |
| 16:32:13 | bauzas | like the PCPU spec | |
| 16:33:06 | bauzas | but again, what I want is to make sure that if operators enable reporting of NUMA nodes using NRPs, then it can still be possible to use hugepages flavors for finding a destination | |
| 16:33:06 | efried | bauzas: okay, maybe we back up and I just answer your original question :) | |
| 16:34:51 | efried | You can specify multiple resources of different classes in a request group. A numbered request group will get *all* of those resources from the *one* resource provider. The un-numbered request group will get the resources from any provider in a tree or associated sharing providers. However, even in the latter case, all resources of a specific resource *class* will still come from a single provider. | |
| 16:39:51 | bauzas | efried: I see, thanks | |
| 16:40:07 | bauzas | so that could work | |
| 16:40:59 | bauzas | if I'm providing a NUMA topology through nested RPs, placement will give me a resource provider that supports that memory | |
| 16:42:17 | bauzas | efried: from a scheduler perspective, when it finds a nested resource provider as a destination when calling Placement API, I guess it uses the root RP for passing it down to the filters ? | |
| 16:43:42 | efried | bauzas: We haven't fully closed the switch on that yet, but yes, even if zero resource comes from the root RP, it'll still be the thing used as the "destination host". Not sure if that's a full answer to your question. Because filtering might need more info than that. | |
| 16:43:58 | efried | I'm guessing the entire allocation_request will need to be considered for some filters. | |
| 16:44:23 | bauzas | efried: that's the problem I see with NUMA filter | |
| 16:44:59 | bauzas | efried: because say placement finds a NUMA node, then it will return the child to the scheduler on a classic call | |
| 16:45:13 | bauzas | eg. a regular flavor | |
| 16:45:22 | bauzas | so we need to pass down the root RP | |
| 16:45:43 | bauzas | but then, the NUMA filter could try to find another NUMA node instead of using the one allocated | |
| 16:46:13 | efried | bauzas: Correlating the allocation_request with the provider_summary ought to allow you to figure out which NUMA node the resources were allocated from. | |
| 16:46:31 | efried | But yeah, without further invention, only the virt driver will know which RP UUID corresponds to which NUMA node. | |
| 16:47:10 | bauzas | that's not really the problme | |
| 16:47:16 | efried | ...which is kind of appropriate, because "identifying a NUMA node" is a virt-specific thing. | |
| 16:47:26 | efried | I.e. libvirt is gonna do it a different way than hyperv or whatever. | |
| 16:48:12 | bauzas | the problem is, say you ask for 2GB of memory within a NUMA node, then placement gives you host A with 2 NUMA nodes but only one NUMA node for host B | |
| 16:48:24 | bauzas | because the other NUMA node of host B is full | |
| 16:48:45 | bauzas | then, we need to pass to the scheduler filters the root RP | |
| 16:49:10 | bauzas | in theory, when it goes on NUMA filter for host B, it could consider the second NUMA node for host B as legit | |
| 16:49:17 | bauzas | there be dragons | |
| 16:50:10 | efried | how could it? | |
| 16:50:27 | efried | There's no candidate with allocations in that second NUMA node on host B. | |
| 16:50:36 | sean-k-mooney | bauzas: there might be dragons but the host state object for host b should also know that the second numa node is fully used and ignore it | |
| 16:51:14 | sean-k-mooney | bauzas: what is an issue if both numa nodes are valid and the filter chooses the other one form placement | |
| 16:51:49 | sean-k-mooney | e.g. placement decremetes the inventor that corresponds to node 0 but the numa topology filter decrements node 1 | |
| 16:51:58 | bauzas | okay, then I'm maybe overthinking | |
| 16:52:38 | sean-k-mooney | bauzas: there is an edge case here but its for host A with 2 NUMA not host B with 1 | |
| 16:53:52 | sean-k-mooney | for host b the resouce tracker will have updted the numatoplogy blob to also show the second numa nodes as full but in the case of host A both are valid from its point of view | |
| 16:55:02 | bauzas | I guess my fears are coming from the fact we litterally try to draw something out of nowhere, and without good testing for making sure we don't trample folks | |
| 16:55:14 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: DiskAdapter parent class https://review.openstack.org/549053 | |
| 16:55:28 | bauzas | if I was able to just test what I write, I wouldn't be trying to consider all the edge cases | |
| 16:55:55 | openstackgerrit | Eric Berglund proposed openstack/nova master: WIP: PowerVM Driver: Localdisk https://review.openstack.org/549300 | |
| 16:56:05 | bauzas | and I just feel I'm just trying to sink all the ocean's water | |
| 16:57:36 | sean-k-mooney | bauzas: i think you have raised a valid issue here. the virt driver will likely need to tag the resouce providres with a trait or aggregat to allow it to map its internal view(in the resouce tracker/numa toploygy bob) to the view it gets back in the allocation canditates | |
| 16:58:25 | bauzas | sean-k-mooney: how do you see that ? | |
| 16:59:40 | sean-k-mooney | how would you do it? when the virt driver create the RP for the numanode in the provider tree update it would include a CUSTOM_HOST_NUMA_ID_X trait where X is its internal identify for the numa node e.g. 0 or 1 | |
| 17:00:11 | sean-k-mooney | then in the allocation canditates resoponce the numa topology filter can use that trait to map the the correct cell in the numa topology blob | |
| 17:00:43 | sean-k-mooney | you could also use an agregate but that would be harder to map to the topoploy blob unless we add teh aggregate uuid to the blob | |
| 17:01:32 | bauzas | sean-k-mooney: ouch. | |
| 17:01:37 | sean-k-mooney | that trait/aggreage would be uses soly by the virtdriver/filter and never passed in any request to placement | |
| 17:02:12 | bauzas | sean-k-mooney: the problem is that the virt.hardware module is a pleasure to modify | |
| 17:02:39 | bauzas | I'd really want to avoid any subsequent modification | |
| 17:03:26 | sean-k-mooney | bauzas: yes well if we use a trait then we dont need to modify it | |
| 17:03:43 | sean-k-mooney | we just need to modify the update provider tree stuff to also include the cell id | |
| 17:03:57 | sean-k-mooney | as a trait on the numa node | |
| 17:04:23 | bauzas | not sure I'm getting you | |
| 17:04:48 | sean-k-mooney | i have not looked but im assumeing the numa patches where going to use the info from the numa topology blob to create teh RPs for the NUMA node and the sub resouces of that node | |
| 17:04:56 | bauzas | because the virt.hardware module gets a topology from both the hoststate and the instance proposed topolgy | |
| 17:05:33 | bauzas | here, we would need to hack the module to look at the resource providers, right ? | |
| 17:05:42 | bauzas | instead of the host state | |
| 17:06:22 | bauzas | anyway, I'm running out of fuel for my brain | |
| 17:06:49 | sean-k-mooney | bauzas: i think we would need to pass in the allocation candiates to the filter yes and then pass that down into the fit_instance_to_host fucntion or whatever it is called so that it could make a descission based on the allcoation candiate | |
| 17:07:10 | bauzas | right, that's what I meant | |
| 17:07:41 | sean-k-mooney | ya that fuction is a pain to modify or debug but its going to need to be scoped to the allocation candiate to work correct when numa is in placement | |
| 17:09:54 | jaypipes | bauzas, stephenfin, sean-k-mooney, efried: sorry, done with call now. | |
| 17:10:13 | bauzas | jaypipes: I'm just rat-holing | |
| 17:11:44 | bauzas | my concern is, how to make sure we can still have all the NUMA features be workable in a world with nested RPs albeit all things solved in the future++ with placement resources | |
| 17:12:00 | sean-k-mooney | jaypipes: the issue is basically how to correlate placement RPs with the compute node resouce tracker so that when the the numa topology filter or pining code runs we only look at the resouce selected by placement in the allocation candidate and not all numa nodes for example. | |
| 17:20:22 | jaypipes | efried: iota? | |
| 17:33:24 | jaypipes | host() is run again. If that picks a different NUMA node than what is in the allocation_request that is sent along with the build request, then we raise an exception and just retry the scheduling. | |
| 17:33:24 | jaypipes | sean-k-mooney: I don't think it's really a big issue. Basically, let the NUMA topology filter just run as-is. It will "pick" a NUMA node to pin the instance to (and then promptly forget about its pick). The scheduler will claim resources against one of the NUMA nodes on the host (via the normal allocation request claim_resources() process). The build request gets to the compute host. During the instance_claim() process, the numa_fit_instance_to_ | |
| 17:35:11 | dansmith | efried: so, questions about L371 here: https://review.openstack.org/#/c/547990/10/nova/scheduler/client/report.py | |
| 17:35:16 | dansmith | efried: what 406 are you talking about? | |
| 17:35:38 | dansmith | the only one I know of is if the version isn't supported that we need for member_of | |
| 17:36:10 | dansmith | efried: and, I'm only running the intersection and setting of the member_of if aggregates is non-empty, which is what you're saying I'll need to do | |
| 17:36:15 | sean-k-mooney | jaypipes: thats one option be se should really not have to retry here | |
| 17:36:32 | sean-k-mooney | jaypipes: we should be able to just look at the cell that was selected by placement | |
| 17:38:00 | sean-k-mooney | also when numa_fit_instance_to_host runs how to you tell if it picked a different node to the one in the allocation_request if you cant correlate between them | |
| 17:39:44 | efried | dansmith: The functionality you want is to be able to do the set logic on aggregates. We plan to allow placement to do that via some new syntax (your pending spec delta). Once that happens, there'll be a microversion for that. And at that time, if you try to use that new microversion, you'll also have to handle the case where placement is downlevel, just like we do everywhere else where we might be straddling versions. | |
| 17:40:25 | edleafe | efried: is there a spec/bp/bug for adding the consumer generation? | |
| 17:40:40 | efried | dansmith: And what I'm saying is that the 406 branch will have to do the placement query without taking member_of into account (or using the intersection thing as a prefilter) and then do the set logic on the candidates that come back for that too-broad query. | |
| 17:40:53 | dansmith | efried: why? the last couple times we've bumped that version we have't supported an older placement | |
| 17:41:04 | efried | Whoah. Yes, we do, every time. | |