| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-02-04 | |||
| 15:05:34 | stephenfin | efried: You can but only if you don't configure anything with NUMA | |
| 15:05:42 | sean-k-mooney | got to join a call | |
| 15:05:46 | stephenfin | darn, ninja'd by sean-k-mooney | |
| 15:05:49 | efried | Feel like that came all the way back in a circle. | |
| 15:06:42 | stephenfin | efried: the original question was why can't we always report NUMA to placement, yeah? | |
| 15:07:45 | efried | But now we're considering removing that first possibility by forcing all hosts to be NUMA-aware. | |
| 15:07:45 | efried | And the proposal we started the morning with would allow you to do the same. | |
| 15:07:45 | efried | IOW to boot a huge instance today, you can artificially give it a multi-numa topo, or you can say nothing about NUMA. | |
| 15:07:57 | efried | sorry, removing the *second* possibility | |
| 15:08:31 | stephenfin | I don't think it's possible to force all hosts to be NUMA-aware | |
| 15:09:26 | bauzas | that's the crux of the problem | |
| 15:09:51 | efried | But dansmith has been arguing against that. | |
| 15:09:51 | efried | And you could only boot flavors with hw:numa*isms into the former; and you could only boot flavors *without* numa*isms into the latter. | |
| 15:09:51 | efried | stephenfin: I thought we had this all sussed out. I thought we were going to segregate the cloud into NUMA-aware (placement resources split along NUMA RPs) and non-NUMA-aware (what it looks like today, with all proc/mem on the root RP) hosts | |
| 15:09:54 | bauzas | are we able nowadays to express a placement query for a large VM with a NUMA-aware host ? | |
| 15:10:25 | bauzas | given the new proposal you made in the etherpad | |
| 15:10:28 | stephenfin | Agree on the first point | |
| 15:10:30 | dansmith | efried: dude, can you let up a bit? I think multiple people are saying it would be nice to not have this restriction, no? | |
| 15:10:47 | stephenfin | Don't recall agreeing to the latter | |
| 15:10:58 | bauzas | like, "I want 8 VCPUs" can it be satisfied with 4 CPUs on each NUMA node ? | |
| 15:10:59 | dansmith | efried: I'm not demanding anything, I'm just saying I don't think that expecting to need to segregate the fleet forever is the best long term plan | |
| 15:11:32 | efried | It would be nice, but (and again I may be misremembering the discussions) I thought we decided to compromise because that would be too hard to do. | |
| 15:11:38 | bauzas | like, could we assume a specific query attribute to placement unless others are expressed ? | |
| 15:11:45 | efried | bauzas: no, that doesn't work. | |
| 15:11:53 | bauzas | efried: I'm just challenging this idea | |
| 15:12:19 | efried | bauzas: we've been down that road before -- that was the thing where you would have to ask for individual MB of memory in a zillion granular groups with group_policy=none, remember? | |
| 15:12:56 | bauzas | I see | |
| 15:13:02 | efried | We also proposed can_split to help with that, but abandoned the idea for reasons. | |
| 15:13:20 | bauzas | yeah, I was considering can_spit | |
| 15:13:23 | bauzas | split heh | |
| 15:13:43 | efried | But the other reason was that we had decided on the above architecture and designed same_subtree etc. to accommodate it. | |
| 15:13:43 | efried | one reason was that it was going to be really hard to make the syntax work properly | |
| 15:14:07 | bauzas | like, until we somehow have a placement construction that allows us to 'spread a query across multiple RPs', then the option is mandatory :( | |
| 15:14:12 | stephenfin | dansmith: Yeah, no reason this opt-out of NUMA behavior couldn't be phased out over multiple releases | |
| 15:14:25 | dansmith | stephenfin: that's specifically what I'm saying | |
| 15:14:49 | efried | stephenfin: so in order to fit my large VM, I would have to specify multiple NUMA nodes? | |
| 15:15:00 | efried | ...in that future release? | |
| 15:15:05 | stephenfin | efried: Yup | |
| 15:15:26 | stephenfin | Or turn off NUMA on your host | |
| 15:15:39 | stephenfin | It's usually tucked away in the BIOS | |
| 15:15:57 | dansmith | from the user's perspective, | |
| 15:16:00 | efried | okay, I didn't know that was even an option. So the driver would report effectively a single NUMA node in that case | |
| 15:16:06 | stephenfin | correct | |
| 15:16:10 | efried | well shit | |
| 15:16:20 | stephenfin | fwiw, this is the same decision we made with the thread policies | |
| 15:16:49 | dansmith | if we had a hw:numa_hodes_min= thing, then they could express in the flavor whether they *need* two numa nodes, or are willing to *tolerate* multiple nodes, the current thing being the upper limit if both specified, right? | |
| 15:16:53 | bauzas | I need to come to a conclusion because parenting taxi duties | |
| 15:17:08 | dansmith | the fact that we've backed ourselves into a corner with placement and expressivity notwithstanding | |
| 15:17:15 | efried | dansmith: So yeah, I was going to address that. The problem is that we have no way to translate that into placement... yeah. | |
| 15:17:16 | bauzas | A/ we are about to propose a flag for allowing NUMA architecture | |
| 15:17:25 | stephenfin | dansmith: the biggest issues with that is that we've to make multiple requests to placement at the moment | |
| 15:17:31 | bauzas | B/ we're not intending to remove this flag in a foreseenable future | |
| 15:18:02 | dansmith | efried: and I'm saying that sucks for the users, and why they're opting into the dumb behavior (whether through nova or bios) because _we_ can't figure out how to organize our own data | |
| 15:18:06 | bauzas | C/ we explicity ask our operators to turn this flag on to allow them to boot NUMA-aware guests on such hosts | |
| 15:18:10 | dansmith | stephenfin: yep, understand | |
| 15:18:16 | stephenfin | i.e. give me a host that can fit all N in 1 NUMA node, else give me one that fit them in 2 nodes, ... | |
| 15:18:33 | stephenfin | *all N instance cores | |
| 15:18:40 | efried | stephenfin: but if in 2 nodes, we can split evenly or asymmetrically... | |
| 15:18:42 | bauzas | D/ I state in the alternatives section that this whole plan sucks because we miss placement expressivity | |
| 15:18:49 | bauzas | WFY folks ? | |
| 15:19:34 | dansmith | efried: right, and we'd want some "minimum split is 70/30" type expression too.. I get that it's not something we can do today, and will be harder to land that across now two separate projects | |
| 15:19:51 | dansmith | I'm just saying, the user sees this as a rock and a hard place, with no real-life justification for it | |
| 15:20:29 | bauzas | ttyl later folks and will scroll back | |
| 15:20:30 | dansmith | I guess if you don't care about a specific topology, you want the same percentage of memory on each node as cpus | |
| 15:20:33 | gibi | efried: thanks the summary about group_policy, it works fo me | |
| 15:20:33 | sean-k-mooney | ok back | |
| 15:20:40 | stephenfin | efried: yeah, correct, otherwise you force people to use those awful 'hw:numa.cpu{N}=<cpumap>' extra specs | |
| 15:21:08 | stephenfin | dansmith: sure, though it's a bit of weird one, implementation wise | |
| 15:22:05 | stephenfin | because placement will give us e.g. a 70/30 split on cores, but we wouldn't really be reflecting this in the pinning of the instance to the host | |
| 15:22:33 | stephenfin | so we'll have to be careful not to do strict NUMA memory affinity in that case | |
| 15:22:38 | dansmith | yeah | |
| 15:23:04 | dansmith | presumably there's a middle ground between fully constrained and unconstrained, | |
| 15:23:20 | dansmith | where your vcpus are constrained to only run on cores that are on the numa node they represent, right? | |
| 15:23:40 | stephenfin | Correct. That's what happens when you turn on a NUMA topology without pinning at the moment | |
| 15:23:51 | dansmith | you balance between cores looser than pinning, but... right okay | |
| 15:24:01 | stephenfin | i.e. use 'hw:numa_nodes' or 'hw:mem_page_size' | |
| 15:24:22 | dansmith | that seems fine to me then | |
| 15:24:25 | stephenfin | We "pin" to the whole range of enabled cores from N NUMA nodes | |
| 15:24:55 | stephenfin | rn there's no way to say give me N $resource from adjacent/child providers, right? That's what the 'can_split' thing was supposed to do? | |
| 15:25:08 | dansmith | if placement were able to cough up a topology for 1-2 numa nodes (i.e. i don't care) and then I get vcpus loosely pinned to cores on the right numa node according to the memory split... | |
| 15:25:24 | stephenfin | yeah, ideal | |
| 15:25:35 | dansmith | aye | |
| 15:26:02 | stephenfin | I guess to retain the current behavior, you'd cough up a topology for *all* NUMA nodes on the host | |
| 15:26:27 | stephenfin | but we can't do that since it would break e.g. a 1 core instance on a 2 node host | |
| 15:26:41 | dansmith | that's why it has to be a range I think | |
| 15:26:50 | stephenfin | yup | |
| 15:27:28 | sean-k-mooney | you know we can totally allcoate memroy from multiple host numa nodes and expose it as one gues numa node with qemu right | |
| 15:27:33 | sean-k-mooney | same with cores | |
| 15:27:44 | dansmith | sure, but that's not helpful | |
| 15:27:47 | efried | isn't that what we do today for a don't-care-about-NUMA guest? | |
| 15:28:09 | efried | and is the exact flexibility we're talking about getting rid of? | |
| 15:28:21 | dansmith | nobody is asking to be lied to :) | |
| 15:28:44 | sean-k-mooney | dansmith: well im jsut saying we can always use the resources that correspond to the placmenet allcaotion and expose a different virtual numa toplogy if we were willing to not require the 1:1 mapping unless you said you cared about numa | |
| 15:28:54 | dansmith | efried: no, that's not flexibility | |
| 15:29:10 | dansmith | efried: nobody is asking for "show me one numa node even though that's not the truth" | |
| 15:29:11 | sean-k-mooney | efried: lie to it yes | |
| 15:29:19 | mriedem | ignorance is bliss | |
| 15:29:31 | efried | If I ask for a kosher sausage, I'm going to be upset if it's pork. | |
| 15:29:31 | efried | If I ask for a sausage, I'm going to be fine if the sausage is beef or pork. | |
| 15:29:33 | dansmith | sean-k-mooney: gotcha | |