Earlier  
Posted Nick Remark
#openstack-nova - 2020-02-11
19:01:07 efried which simply splits your CPU and mem "evenly" across $n nodes.
19:01:15 dansmith that just seems to naive to me
19:01:19 efried yes
19:01:20 efried it is
19:01:32 efried but so is "I don't care about NUMA"?
19:01:41 efried And it ought to work 99% of the time
19:01:52 efried and if it doesn't, switch on your [workaround] for a couple of hosts.
19:02:07 dansmith I want the workaround to go away, remember
19:02:23 efried Yes, the workaround goes away completely once we beef up placement to understand can_split
19:02:53 dansmith okay your even splitting is just the first pass at sanity? then that's fine
19:04:27 efried oh, yeah, eventually we want the utopia where all of this happens in one call with can_split, whose ratios are tunable (via placement side conf? via nova conf fed into placement qparams?)
19:04:43 dansmith has to be nova-side
19:04:52 dansmith communicated to placement via the query
19:05:15 efried then we won't need the workaround anymore, because everything will be able to land (and if it can't, it's because it *shouldn't*).
19:05:43 efried In practice, we may even find that nobody needs the workaround. But we'll see.
19:09:39 efried Did we decide how we're going to deal with "control plane is updated but some computes are not"? Does the control plane wait to start using the new query style until all the computes are updated? We have that capability, right?
19:11:05 efried okay, I see that discussed in the spec.
19:18:58 sean-k-mooney efried: i think if we go that route in ussuri we wont be able to land it in time
19:19:36 efried why not?
19:19:47 efried we're not talking about trying to implement can_split in any form in ussuri
19:20:04 efried are you concerned that the progressive-splitting algorithm is too complicated?
19:20:28 sean-k-mooney just multiple queries wehre we try to progressively split
19:20:36 sean-k-mooney efried: yes
19:21:01 efried meh, I don't see how it's any worse than the proposed fallback.
19:21:28 sean-k-mooney efried: it will have to take into accoung numa, native request groups in teh flavor, external requests form cyborng and netorn ports and not suck form a performnce point of view
19:21:58 sean-k-mooney the proposed fallback is two queries, the native nuam one followed by the query we do today
19:22:08 sean-k-mooney and it only impacts the perfomce of numa instances
19:22:20 sean-k-mooney the other way imacpts the perfomce of all non numa instnaces
19:22:37 efried sean-k-mooney: I don't think it has to take all that stuff into account at all.
19:23:07 efried sean-k-mooney: It should behave *exactly* as if you said hw:numa_nodes=$n with no other hw:numa*-isms.
19:23:40 efried (except I think we said we would bounce if we couldn't split evenly; that restriction would have to be lifted for this case.)
19:24:12 sean-k-mooney so we would create a fake flaovr where we overrid that and pass it to the current get numa constratis funct
19:24:19 sean-k-mooney *function
19:24:25 efried if you like
19:24:33 efried that would be the spirit, anyway.
19:24:36 sean-k-mooney i gues that makes it simpler
19:25:14 sean-k-mooney so the get_numa_constraits function will reject any invalid toplogy with regard to even spliting with an excetion
19:25:32 sean-k-mooney so we would just loop and contiue if an excption is raised up to the limit
19:26:09 efried or we relax the constraint to split as close to evenly as possible. Or do that split first.
19:26:33 efried Implementation detail. Point is, it shouldn't be super hard to figure out.
19:26:49 sean-k-mooney we could use the asemetric numa modeling support that is there yes
19:27:19 sean-k-mooney e.g. if you had a 9 core vm and we are on numa=2 do 4cpus+5cpus
19:28:01 sean-k-mooney instead of going to 3 numa nodes with 3 cpus
19:28:14 efried right
19:28:36 sean-k-mooney im not sure which is more likely to cause fragmentation of the top of my head
19:28:50 sean-k-mooney we should tell people to just use powers of 2
19:28:51 efried example I gave above was with 10 VCPUs, we would try 10, then 5/5, then 3/3/4, then 2/3/2/3.
19:28:55 efried no
19:29:01 sean-k-mooney i was jokeing
19:29:10 efried we should tell people to use real numa specs if they care.
19:29:19 sean-k-mooney it does make life easier when they do but ya i can see that working
19:30:51 sean-k-mooney ok if we caluate the split in the tempory flavor we pass to the numa constratis function then it would populate the instance numa toplogy object as if the user had set it manuyaly in teh falvor
19:31:07 sean-k-mooney then if we save that in the instnace we could ensure we dont break live migration by changing it
19:31:08 efried just so.
19:31:32 efried well, I would expect we shouldn't save it in the flavor, because we want the instance to be able to morph to fit somewhere else if it needs to.
19:31:46 sean-k-mooney right
19:31:57 efried but I don't know how that works, do you break an instance if you "change" its topo from under it?
19:31:58 sean-k-mooney i ment save the instance_numa_toplogy object
19:32:01 sean-k-mooney not the falavor
19:32:21 sean-k-mooney efried: during live migration you would
19:32:38 sean-k-mooney cold migration it might mess up some manula config but it should not break it in general
19:32:58 sean-k-mooney i was suggesting once we select a toplogy we store it in the request_spec and instnace
19:32:58 efried hm, well that's a bummer. So how do we migrate numa-agnostic instances today?
19:33:11 sean-k-mooney so that it stays the same for the liftime of the instance unless you resize
19:33:14 sean-k-mooney or rebuild
19:33:34 sean-k-mooney am today
19:33:44 sean-k-mooney non numa isntace are alwasy exposed as 1 numa node
19:33:54 sean-k-mooney so it never chagnes form the gest point of view
19:34:02 sean-k-mooney so we just lie to the guest
19:34:10 efried couldn't we continue lying to the guest?
19:34:21 sean-k-mooney we can yes
19:34:39 efried does the guest do things differently if it knows CPU x is affined to memory y?
19:34:51 sean-k-mooney so we would do the progessive spliting to select the host resouces and present it as 1 numa node to the guest
19:35:04 sean-k-mooney but that will have worse performace then telling it its actull toplogy
19:35:10 sean-k-mooney yes
19:35:22 sean-k-mooney the kernel will take that into account wehn allocation memroy for a proces
19:35:36 sean-k-mooney trying to use numa local memory ahead of remote numa memory
19:35:44 spatel sean-k-mooney: hey
19:35:50 efried well, I guess this is a problem we would have eventually anyway, right dansmith?
19:35:52 sean-k-mooney if we dotn expose the toplogy to the vm the vm kernel wont know how to optimise
19:36:01 spatel did you see my last mesg?
19:36:43 dansmith efried: what? needing everything to support guests with numa?
19:37:01 spatel running vm on single numa0 give very high performance compare to running on both numa node (with cpu_socket=2 and cpu_threads=2 option)
19:37:32 efried Which will then limit where you can fit it on migration.
19:37:32 efried If you migrate that instance, you have to preserve that topo on the target; you can't just munge it to a new shape.
19:37:32 efried dansmith: TL;DR: In the world where you say you don't care about NUMA, we give you an instance that's NUMA-ified, but with whatever split we were able to fit.
19:37:36 sean-k-mooney spatel: i was away but just saw it now. fundemtally i think erlang is jut not optimising correctly. im not really sure how to help other then suggestin run more small vms if you can scale out the application horrizontally instead of vertically
19:37:44 dansmith sean-k-mooney: presumably if the flavor on the instance is not numa-aware then we can just find a new host during cold migration and give it a new topology when it moves
19:37:59 sean-k-mooney dansmith: yes we could
19:38:00 dansmith efried: only on cold migration
19:38:02 dansmith sorry
19:38:05 dansmith efried: only on live migration
19:38:28 dansmith efried: on cold migration it could change because you're rebooting and there shouldn't be anything in the guest that persistently cares what the topology is..
19:38:29 efried Okay, that would be a new limitation.
19:38:32 spatel sean-k-mooney: totally understand but i am going to loose 16 cpu in that case. but anyway i can live with that
19:38:45 dansmith efried: we have that limitation today and it's handled by the numa live migration stuff, AFAIK
19:39:34 efried dansmith: oh, my understanding was that, by lying to the guest and saying it only has one NUMA node, regardless of how many are under the covers, we can change the under-the-covers on live migration without "affecting" the guest. It would still have crappy performance on both sides, but it wouldn't notice that anything had changed.
19:39:38 sean-k-mooney efried: well the limitation would be dont change the view of hardwar the vm sees in live migration
19:39:46 dansmith efried: and selecting a host in that situation should be the same as selecting a host for a new instance boot where the flavor cares deeply about the topology
19:39:47 sean-k-mooney when framed that way its what we do today

Earlier   Later