| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-03-26 | |||
| 16:24:21 | bauzas | (tbh) | |
| 16:24:39 | bauzas | I'm just hardly trying to not lock all the world | |
| 16:25:03 | efried | IMO we should not try to do the thing where we have e.g. 2048MB of memory in each NUMA node and then represent 4096MB on the compute node which isn't really there. | |
| 16:25:05 | bauzas | like I'd like to provide a mechanism to select a NUMA node but huge pages could still be checked by the virt driver | |
| 16:26:02 | bauzas | efried: I got your point, my problem is more about trying hard to not draw a house of cards | |
| 16:26:15 | efried | bauzas: is there a fixed ratio between memory MB and number of huge pages? | |
| 16:26:32 | bauzas | efried: the solution is simple in my mind | |
| 16:26:42 | bauzas | that's all about traits and step_size | |
| 16:26:45 | bauzas | but | |
| 16:26:59 | bauzas | if we go decide to provide a tree of NUMA nodes | |
| 16:27:35 | bauzas | but we don't support yet hugepages, then we have a problem because flavors asking for hugepages would get NoValidHosts | |
| 16:28:01 | bauzas | we somehow still need to accept compute nodes to report their memory the old way | |
| 16:28:09 | efried | what's a hugepage? | |
| 16:28:17 | efried | (use little words) | |
| 16:28:26 | bauzas | https://docs.openstack.org/nova/latest/admin/huge-pages.html | |
| 16:28:39 | bauzas | it's a defined memory page size | |
| 16:28:48 | bauzas | but that's just an example | |
| 16:28:57 | bauzas | for the moment, placement only checks the total memory of the host, right? | |
| 16:29:12 | bauzas | then the scheduler filter does magic things to find a destination | |
| 16:29:40 | efried | Well, yeah, so this is kind of in line with what we need to do wrt bandwidth. | |
| 16:29:52 | efried | We represent the number of huge pages as inventory alongside the MEMORY_MB inventory. | |
| 16:29:53 | efried | but | |
| 16:30:12 | efried | if the consumer is going to include hugepages as part of their request, they *always* need to do so. | |
| 16:30:27 | bauzas | so, my concern is that if I'm beginning to shard the memory between NUMA nodes, then it requires the flavors to be updated to explicitly ask for a NUMA node, whereas huge pages are totally NUMA unrelated | |
| 16:30:59 | efried | because what we can't (or shouldn't) do is try to convert a request for MEMORY_MB:4096 into MEMORY_MB:4096,HUGEPAGES:4 (or whatever) | |
| 16:31:26 | bauzas | efried: I'm not trying to design now how to make placement queries for hugepages | |
| 16:31:29 | efried | bauzas: But does a huge page come from the same place as a MEMORY_MB or is it a separate thing? | |
| 16:31:46 | bauzas | efried: what I'm trying is to make sure we keep a compatible behaviour for the existing feature | |
| 16:32:05 | bauzas | others, later, will try to solve that design and use placement resources for that | |
| 16:32:13 | bauzas | like the PCPU spec | |
| 16:33:06 | efried | bauzas: okay, maybe we back up and I just answer your original question :) | |
| 16:33:06 | bauzas | but again, what I want is to make sure that if operators enable reporting of NUMA nodes using NRPs, then it can still be possible to use hugepages flavors for finding a destination | |
| 16:34:51 | efried | You can specify multiple resources of different classes in a request group. A numbered request group will get *all* of those resources from the *one* resource provider. The un-numbered request group will get the resources from any provider in a tree or associated sharing providers. However, even in the latter case, all resources of a specific resource *class* will still come from a single provider. | |
| 16:39:51 | bauzas | efried: I see, thanks | |
| 16:40:07 | bauzas | so that could work | |
| 16:40:59 | bauzas | if I'm providing a NUMA topology through nested RPs, placement will give me a resource provider that supports that memory | |
| 16:42:17 | bauzas | efried: from a scheduler perspective, when it finds a nested resource provider as a destination when calling Placement API, I guess it uses the root RP for passing it down to the filters ? | |
| 16:43:42 | efried | bauzas: We haven't fully closed the switch on that yet, but yes, even if zero resource comes from the root RP, it'll still be the thing used as the "destination host". Not sure if that's a full answer to your question. Because filtering might need more info than that. | |
| 16:43:58 | efried | I'm guessing the entire allocation_request will need to be considered for some filters. | |
| 16:44:23 | bauzas | efried: that's the problem I see with NUMA filter | |
| 16:44:59 | bauzas | efried: because say placement finds a NUMA node, then it will return the child to the scheduler on a classic call | |
| 16:45:13 | bauzas | eg. a regular flavor | |
| 16:45:22 | bauzas | so we need to pass down the root RP | |
| 16:45:43 | bauzas | but then, the NUMA filter could try to find another NUMA node instead of using the one allocated | |
| 16:46:13 | efried | bauzas: Correlating the allocation_request with the provider_summary ought to allow you to figure out which NUMA node the resources were allocated from. | |
| 16:46:31 | efried | But yeah, without further invention, only the virt driver will know which RP UUID corresponds to which NUMA node. | |
| 16:47:10 | bauzas | that's not really the problme | |
| 16:47:16 | efried | ...which is kind of appropriate, because "identifying a NUMA node" is a virt-specific thing. | |
| 16:47:26 | efried | I.e. libvirt is gonna do it a different way than hyperv or whatever. | |
| 16:48:12 | bauzas | the problem is, say you ask for 2GB of memory within a NUMA node, then placement gives you host A with 2 NUMA nodes but only one NUMA node for host B | |
| 16:48:24 | bauzas | because the other NUMA node of host B is full | |
| 16:48:45 | bauzas | then, we need to pass to the scheduler filters the root RP | |
| 16:49:10 | bauzas | in theory, when it goes on NUMA filter for host B, it could consider the second NUMA node for host B as legit | |
| 16:49:17 | bauzas | there be dragons | |
| 16:50:10 | efried | how could it? | |
| 16:50:27 | efried | There's no candidate with allocations in that second NUMA node on host B. | |
| 16:50:36 | sean-k-mooney | bauzas: there might be dragons but the host state object for host b should also know that the second numa node is fully used and ignore it | |
| 16:51:14 | sean-k-mooney | bauzas: what is an issue if both numa nodes are valid and the filter chooses the other one form placement | |
| 16:51:49 | sean-k-mooney | e.g. placement decremetes the inventor that corresponds to node 0 but the numa topology filter decrements node 1 | |
| 16:51:58 | bauzas | okay, then I'm maybe overthinking | |
| 16:52:38 | sean-k-mooney | bauzas: there is an edge case here but its for host A with 2 NUMA not host B with 1 | |
| 16:53:52 | sean-k-mooney | for host b the resouce tracker will have updted the numatoplogy blob to also show the second numa nodes as full but in the case of host A both are valid from its point of view | |
| 16:55:02 | bauzas | I guess my fears are coming from the fact we litterally try to draw something out of nowhere, and without good testing for making sure we don't trample folks | |
| 16:55:14 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: DiskAdapter parent class https://review.openstack.org/549053 | |
| 16:55:28 | bauzas | if I was able to just test what I write, I wouldn't be trying to consider all the edge cases | |
| 16:55:55 | openstackgerrit | Eric Berglund proposed openstack/nova master: WIP: PowerVM Driver: Localdisk https://review.openstack.org/549300 | |
| 16:56:05 | bauzas | and I just feel I'm just trying to sink all the ocean's water | |
| 16:57:36 | sean-k-mooney | bauzas: i think you have raised a valid issue here. the virt driver will likely need to tag the resouce providres with a trait or aggregat to allow it to map its internal view(in the resouce tracker/numa toploygy bob) to the view it gets back in the allocation canditates | |
| 16:58:25 | bauzas | sean-k-mooney: how do you see that ? | |
| 16:59:40 | sean-k-mooney | how would you do it? when the virt driver create the RP for the numanode in the provider tree update it would include a CUSTOM_HOST_NUMA_ID_X trait where X is its internal identify for the numa node e.g. 0 or 1 | |
| 17:00:11 | sean-k-mooney | then in the allocation canditates resoponce the numa topology filter can use that trait to map the the correct cell in the numa topology blob | |
| 17:00:43 | sean-k-mooney | you could also use an agregate but that would be harder to map to the topoploy blob unless we add teh aggregate uuid to the blob | |
| 17:01:32 | bauzas | sean-k-mooney: ouch. | |
| 17:01:37 | sean-k-mooney | that trait/aggreage would be uses soly by the virtdriver/filter and never passed in any request to placement | |
| 17:02:12 | bauzas | sean-k-mooney: the problem is that the virt.hardware module is a pleasure to modify | |
| 17:02:39 | bauzas | I'd really want to avoid any subsequent modification | |
| 17:03:26 | sean-k-mooney | bauzas: yes well if we use a trait then we dont need to modify it | |
| 17:03:43 | sean-k-mooney | we just need to modify the update provider tree stuff to also include the cell id | |
| 17:03:57 | sean-k-mooney | as a trait on the numa node | |
| 17:04:23 | bauzas | not sure I'm getting you | |
| 17:04:48 | sean-k-mooney | i have not looked but im assumeing the numa patches where going to use the info from the numa topology blob to create teh RPs for the NUMA node and the sub resouces of that node | |
| 17:04:56 | bauzas | because the virt.hardware module gets a topology from both the hoststate and the instance proposed topolgy | |
| 17:05:33 | bauzas | here, we would need to hack the module to look at the resource providers, right ? | |
| 17:05:42 | bauzas | instead of the host state | |
| 17:06:22 | bauzas | anyway, I'm running out of fuel for my brain | |
| 17:06:49 | sean-k-mooney | bauzas: i think we would need to pass in the allocation candiates to the filter yes and then pass that down into the fit_instance_to_host fucntion or whatever it is called so that it could make a descission based on the allcoation candiate | |
| 17:07:10 | bauzas | right, that's what I meant | |
| 17:07:41 | sean-k-mooney | ya that fuction is a pain to modify or debug but its going to need to be scoped to the allocation candiate to work correct when numa is in placement | |
| 17:09:54 | jaypipes | bauzas, stephenfin, sean-k-mooney, efried: sorry, done with call now. | |
| 17:10:13 | bauzas | jaypipes: I'm just rat-holing | |
| 17:11:44 | bauzas | my concern is, how to make sure we can still have all the NUMA features be workable in a world with nested RPs albeit all things solved in the future++ with placement resources | |
| 17:12:00 | sean-k-mooney | jaypipes: the issue is basically how to correlate placement RPs with the compute node resouce tracker so that when the the numa topology filter or pining code runs we only look at the resouce selected by placement in the allocation candidate and not all numa nodes for example. | |
| 17:20:22 | jaypipes | efried: iota? | |
| 17:33:24 | jaypipes | sean-k-mooney: I don't think it's really a big issue. Basically, let the NUMA topology filter just run as-is. It will "pick" a NUMA node to pin the instance to (and then promptly forget about its pick). The scheduler will claim resources against one of the NUMA nodes on the host (via the normal allocation request claim_resources() process). The build request gets to the compute host. During the instance_claim() process, the numa_fit_instance_to_ | |
| 17:33:24 | jaypipes | host() is run again. If that picks a different NUMA node than what is in the allocation_request that is sent along with the build request, then we raise an exception and just retry the scheduling. | |
| 17:35:11 | dansmith | efried: so, questions about L371 here: https://review.openstack.org/#/c/547990/10/nova/scheduler/client/report.py | |
| 17:35:16 | dansmith | efried: what 406 are you talking about? | |
| 17:35:38 | dansmith | the only one I know of is if the version isn't supported that we need for member_of | |
| 17:36:10 | dansmith | efried: and, I'm only running the intersection and setting of the member_of if aggregates is non-empty, which is what you're saying I'll need to do | |
| 17:36:15 | sean-k-mooney | jaypipes: thats one option be se should really not have to retry here | |