| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-02-11 | |||
| 18:44:45 | efried | sean-k-mooney, bauzas, gibi: I think I would actually prefer the approach dansmith suggests, even if we only do it at the coarsest level, rather than effectively disable the topology modeling in ussuri. | |
| 18:49:49 | efried | - Flavors without hw:numa*-isms get looped for $n in range(0, $max_sane_number_of_numa_nodes_we_expect_a_host_to_ever_have) to behave as if they had specified hw:numa_nodes=$n, stopping as soon as we get a hit. | |
| 18:49:49 | efried | - Flavors with hw:numa*-isms get translated as specced. | |
| 18:49:49 | efried | - Model with NUMA topology by default. Provide the [workaround] option to *disable* (and un-reshape) for situations where the following is just sh*tting all over itself and nothing is landing. | |
| 18:49:49 | efried | in other words: | |
| 18:50:58 | efried | This is without changing anything in placement. | |
| 18:52:44 | efried | For uneven splits, we just get as close as we can, but make no attempt to be "fuzzy" (like "up to 60/40" or anything like that). So like, for $n=2 and VCPU=10, we try 10, then 5/5, then 3/3/4, then 2/3/2/3. | |
| 18:56:04 | dansmith | there has to be some sort of policy tunable for how wild we're able to get I think | |
| 18:56:29 | efried | max_numa_nodes_guessed ? | |
| 18:57:23 | efried | Are there really hosts out there with more than, say, 8 NUMA cells? | |
| 18:57:27 | dansmith | no, I mean how small of a footprint on a given node we're going to allow | |
| 18:58:34 | efried | To be clear, I'm talking about splitting as evenly as possible, always. | |
| 18:58:45 | efried | You're saying that sometimes that would result in unreasonably small footprint on one numa node anyway? | |
| 18:59:05 | efried | obv we stop trying to split if any of the cells are going to get 0 of anything. | |
| 18:59:21 | efried | so if you ask for 2 vcpu, we go to a max of hw:numa_nodes=2 | |
| 18:59:45 | dansmith | sure, but I think we need to not be willing to split memory at 1MB and 8191MB | |
| 18:59:55 | dansmith | or 31 cpus and 1 | |
| 19:00:05 | dansmith | and the ratio needs to be close/equal for cpus and memory | |
| 19:00:08 | efried | Right, I'm saying that happens implicitly | |
| 19:00:25 | dansmith | how? | |
| 19:00:50 | efried | each iteration of the loop simply behaves as if you said hw:numa_nodes=$n. | |
| 19:01:07 | efried | which simply splits your CPU and mem "evenly" across $n nodes. | |
| 19:01:15 | dansmith | that just seems to naive to me | |
| 19:01:19 | efried | yes | |
| 19:01:20 | efried | it is | |
| 19:01:32 | efried | but so is "I don't care about NUMA"? | |
| 19:01:41 | efried | And it ought to work 99% of the time | |
| 19:01:52 | efried | and if it doesn't, switch on your [workaround] for a couple of hosts. | |
| 19:02:07 | dansmith | I want the workaround to go away, remember | |
| 19:02:23 | efried | Yes, the workaround goes away completely once we beef up placement to understand can_split | |
| 19:02:53 | dansmith | okay your even splitting is just the first pass at sanity? then that's fine | |
| 19:04:27 | efried | oh, yeah, eventually we want the utopia where all of this happens in one call with can_split, whose ratios are tunable (via placement side conf? via nova conf fed into placement qparams?) | |
| 19:04:43 | dansmith | has to be nova-side | |
| 19:04:52 | dansmith | communicated to placement via the query | |
| 19:05:15 | efried | then we won't need the workaround anymore, because everything will be able to land (and if it can't, it's because it *shouldn't*). | |
| 19:05:43 | efried | In practice, we may even find that nobody needs the workaround. But we'll see. | |
| 19:09:39 | efried | Did we decide how we're going to deal with "control plane is updated but some computes are not"? Does the control plane wait to start using the new query style until all the computes are updated? We have that capability, right? | |
| 19:11:05 | efried | okay, I see that discussed in the spec. | |
| 19:18:58 | sean-k-mooney | efried: i think if we go that route in ussuri we wont be able to land it in time | |
| 19:19:36 | efried | why not? | |
| 19:19:47 | efried | we're not talking about trying to implement can_split in any form in ussuri | |
| 19:20:04 | efried | are you concerned that the progressive-splitting algorithm is too complicated? | |
| 19:20:28 | sean-k-mooney | just multiple queries wehre we try to progressively split | |
| 19:20:36 | sean-k-mooney | efried: yes | |
| 19:21:01 | efried | meh, I don't see how it's any worse than the proposed fallback. | |
| 19:21:28 | sean-k-mooney | efried: it will have to take into accoung numa, native request groups in teh flavor, external requests form cyborng and netorn ports and not suck form a performnce point of view | |
| 19:21:58 | sean-k-mooney | the proposed fallback is two queries, the native nuam one followed by the query we do today | |
| 19:22:08 | sean-k-mooney | and it only impacts the perfomce of numa instances | |
| 19:22:20 | sean-k-mooney | the other way imacpts the perfomce of all non numa instnaces | |
| 19:22:37 | efried | sean-k-mooney: I don't think it has to take all that stuff into account at all. | |
| 19:23:07 | efried | sean-k-mooney: It should behave *exactly* as if you said hw:numa_nodes=$n with no other hw:numa*-isms. | |
| 19:23:40 | efried | (except I think we said we would bounce if we couldn't split evenly; that restriction would have to be lifted for this case.) | |
| 19:24:12 | sean-k-mooney | so we would create a fake flaovr where we overrid that and pass it to the current get numa constratis funct | |
| 19:24:19 | sean-k-mooney | *function | |
| 19:24:25 | efried | if you like | |
| 19:24:33 | efried | that would be the spirit, anyway. | |
| 19:24:36 | sean-k-mooney | i gues that makes it simpler | |
| 19:25:14 | sean-k-mooney | so the get_numa_constraits function will reject any invalid toplogy with regard to even spliting with an excetion | |
| 19:25:32 | sean-k-mooney | so we would just loop and contiue if an excption is raised up to the limit | |
| 19:26:09 | efried | or we relax the constraint to split as close to evenly as possible. Or do that split first. | |
| 19:26:33 | efried | Implementation detail. Point is, it shouldn't be super hard to figure out. | |
| 19:26:49 | sean-k-mooney | we could use the asemetric numa modeling support that is there yes | |
| 19:27:19 | sean-k-mooney | e.g. if you had a 9 core vm and we are on numa=2 do 4cpus+5cpus | |
| 19:28:01 | sean-k-mooney | instead of going to 3 numa nodes with 3 cpus | |
| 19:28:14 | efried | right | |
| 19:28:36 | sean-k-mooney | im not sure which is more likely to cause fragmentation of the top of my head | |
| 19:28:50 | sean-k-mooney | we should tell people to just use powers of 2 | |
| 19:28:51 | efried | example I gave above was with 10 VCPUs, we would try 10, then 5/5, then 3/3/4, then 2/3/2/3. | |
| 19:28:55 | efried | no | |
| 19:29:01 | sean-k-mooney | i was jokeing | |
| 19:29:10 | efried | we should tell people to use real numa specs if they care. | |
| 19:29:19 | sean-k-mooney | it does make life easier when they do but ya i can see that working | |
| 19:30:51 | sean-k-mooney | ok if we caluate the split in the tempory flavor we pass to the numa constratis function then it would populate the instance numa toplogy object as if the user had set it manuyaly in teh falvor | |
| 19:31:07 | sean-k-mooney | then if we save that in the instnace we could ensure we dont break live migration by changing it | |
| 19:31:08 | efried | just so. | |
| 19:31:32 | efried | well, I would expect we shouldn't save it in the flavor, because we want the instance to be able to morph to fit somewhere else if it needs to. | |
| 19:31:46 | sean-k-mooney | right | |
| 19:31:57 | efried | but I don't know how that works, do you break an instance if you "change" its topo from under it? | |
| 19:31:58 | sean-k-mooney | i ment save the instance_numa_toplogy object | |
| 19:32:01 | sean-k-mooney | not the falavor | |
| 19:32:21 | sean-k-mooney | efried: during live migration you would | |
| 19:32:38 | sean-k-mooney | cold migration it might mess up some manula config but it should not break it in general | |
| 19:32:58 | sean-k-mooney | i was suggesting once we select a toplogy we store it in the request_spec and instnace | |
| 19:32:58 | efried | hm, well that's a bummer. So how do we migrate numa-agnostic instances today? | |
| 19:33:11 | sean-k-mooney | so that it stays the same for the liftime of the instance unless you resize | |
| 19:33:14 | sean-k-mooney | or rebuild | |
| 19:33:34 | sean-k-mooney | am today | |
| 19:33:44 | sean-k-mooney | non numa isntace are alwasy exposed as 1 numa node | |
| 19:33:54 | sean-k-mooney | so it never chagnes form the gest point of view | |
| 19:34:02 | sean-k-mooney | so we just lie to the guest | |
| 19:34:10 | efried | couldn't we continue lying to the guest? | |
| 19:34:21 | sean-k-mooney | we can yes | |
| 19:34:39 | efried | does the guest do things differently if it knows CPU x is affined to memory y? | |
| 19:34:51 | sean-k-mooney | so we would do the progessive spliting to select the host resouces and present it as 1 numa node to the guest | |
| 19:35:04 | sean-k-mooney | but that will have worse performace then telling it its actull toplogy | |
| 19:35:10 | sean-k-mooney | yes | |
| 19:35:22 | sean-k-mooney | the kernel will take that into account wehn allocation memroy for a proces | |
| 19:35:36 | sean-k-mooney | trying to use numa local memory ahead of remote numa memory | |
| 19:35:44 | spatel | sean-k-mooney: hey | |
| 19:35:50 | efried | well, I guess this is a problem we would have eventually anyway, right dansmith? | |