Earlier  
Posted Nick Remark
#openstack-nova - 2021-03-04
16:49:19 sean-k-mooney yep its the current head of master and the main delta is just fixing the lower-constratists-job and a deprecation wraning
16:49:46 sean-k-mooney normlaly i would say that it could be a bugfix release but we always do feature version bump for the end of a cycle
16:49:49 sean-k-mooney so this looks correct
16:50:37 claudiub Also, speaking of NUMA, I was wondering if you know if an instance placed in a single NUMA node can be live-migrated to another node in a different NUMA node, or it has to be in the same NUMA node? Trying to gauge the severity of that bug that places all the instances in the same NUMA node.
16:51:00 sean-k-mooney claudiub: that can be done but only from train
16:51:23 sean-k-mooney claudiub: artom added numa live migration in the train release whre we can regenerate teh xml as part of the migration
16:51:31 claudiub live-migrating instances with numa topologies, right? I've seen that bit in code
16:52:04 artom claudiub, except the user can't specify which NUMA node - which I think is what you're getting at
16:52:04 sean-k-mooney what bug are you triaging?
16:52:10 artom Nova just finds a "free" one
16:52:52 claudiub artom: that's perfect then. :)
16:53:09 sean-k-mooney claudiub: please do not file a new bug for numa blanacing
16:53:41 claudiub sean-k-mooney: I was hitting an issue where all the created instances were created on the same numa node, so I was asking about that. :)
16:53:58 sean-k-mooney claudiub: yep that is by design
16:54:19 sean-k-mooney its somethign we could chagne but when i brough tit up i was told it was a featur not a bug
16:54:57 sean-k-mooney we talked about in the last ptg and when numa was first being added many years ago
16:55:06 claudiub well, one thing's for certain: the docs and the code don't match. :) The docs say the chosen NUMA node is random.
16:55:28 sean-k-mooney from a user perspecitve it is
16:55:41 claudiub Wondering if we could at least have a config option to randomize the selected numa node (ofc, if it can be randomized. not talking about PCI devices)
16:55:43 sean-k-mooney its not actully random form an api point of view it underfied
16:56:13 sean-k-mooney i have a much better solution for this just trying to fine the bug i already had filed
16:56:34 sean-k-mooney https://bugs.launchpad.net/nova/+bug/1893121
16:56:35 openstack Launchpad bug 1893121 in OpenStack Compute (nova) "nova does not balance vm across numa node or prefer numa node with pci device when one is requested" [Undecided,Confirmed] - Assigned to sean mooney (sean-k-mooney)
16:57:23 sean-k-mooney claudiub: this is what determins the order
16:57:25 sean-k-mooney https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/virt/hardware.py#L2268-L2277
16:57:42 sean-k-mooney as an end user you are not allowed to rely on the detail of that implemenation
16:58:05 sean-k-mooney what i am proposing and had plannd on woking on is doing muliple sorts
16:58:21 sean-k-mooney so that we will blance vm plamcne based on avaiable resouces on a host
16:59:20 sean-k-mooney claudiub: line 693 is the ptg dicussion form the last ptg on the topic https://etherpad.opendev.org/p/nova-wallaby-ptg
17:00:28 sean-k-mooney this is one of the topic i need to bring up internally but i would like to adress this next cycle if i can make time but if you have time to work on it then that woudl be good too.
17:01:25 sean-k-mooney numa blancing is the non invaisive way to optimise better then we do today.
17:02:01 sean-k-mooney long term we should be doing this more abstractly. e.g. computeing a cost metic for any give plamcnet based on a number of factors and then minimsiing that.
17:02:22 sean-k-mooney kind of like how the wehers work but placment complictates that.
17:03:52 sean-k-mooney claudiub: the workaround for now is to follow the advice we always gave. try to create flavors that aproximate the host toplogy. e.g. if you hosts all have 2 numa nodes then default to createign flavors with hw:numa_nodes=2
17:04:40 openstackgerrit Balazs Gibizer proposed openstack/nova master: Replace blind retry with libvirt event waiting in detach https://review.opendev.org/c/openstack/nova/+/770246
17:08:57 claudiub sean-k-mooney: Hmm, I see. but correct me if I'm wrong, but in that code snippet, the host cells are not randomized or sorted in any other way if there are no pci_requests and no pci_stats. It's the same host_cells athat are set in host_topology.cells (not sure if the order ever changes here). I could spend some time on it, will read the PTG notes as well.
17:09:55 sean-k-mooney if pci devices are not requested we sort the numa nodes to prefer the onces without numa nodes
17:09:56 claudiub But in any case, setting hw_numa_nodes=2 doesn't help in our scenario, especially since we also have nodes with just 1 NUMA node. :) Additionally, setting the instances on 2 numa nodes could affect the guest performance as well.
17:10:26 claudiub sean-k-mooney: indeed, agreed there.
17:10:27 sean-k-mooney otherwise itertools.permutations iterates over them in a determisict maner
17:10:51 sean-k-mooney which for singel numa nodes instace is acending order form numa node 0
17:11:04 sean-k-mooney then we bail out if it fits and dont try any other nodes
17:11:10 sean-k-mooney so that packs numa0
17:11:14 claudiub yep
17:11:39 sean-k-mooney this is has alwasy been the case since numa was firsts added
17:12:01 sean-k-mooney the specific behavior is implemation defiend and is not garenteed by the api
17:12:02 artom stephenfin, D: how did you not jump on https://review.opendev.org/c/openstack/nova/+/774240/12/nova/virt/libvirt/driver.py#697 with a -2?
17:12:04 artom ;)
17:12:35 sean-k-mooney so it can be modified but the current beahivor optimises for being able to spawn large vms
17:12:56 sean-k-mooney if we balance the vms between numa ndoes it will improve performance in genreal but pessimise spawning large vms
17:13:11 sean-k-mooney so there is a tradeoff between packeing and spreading/blancing
17:13:17 claudiub i agree there, and I am aware of that. :)
17:13:36 sean-k-mooney yep so this is why we said this cant be change in a bug and need a spec
17:13:44 sean-k-mooney which means upstream at least not backportable
17:14:03 sean-k-mooney downstream we likely would backport it but not chagne the behavior by default
17:14:13 stephenfin artom: My yoga guy said I needed to be more chill about these things
17:14:17 stephenfin I'll fire him in the morning
17:14:20 sean-k-mooney upstream if accpetd i would like to change the default but in the cycle after its added
17:14:43 stephenfin <sean-k-mooney> so it can be modified but the current beahivor optimises for being able to spawn large vms
17:14:51 stephenfin are you talking about unpinned instances?
17:15:01 sean-k-mooney stephenfin: no all numa instances
17:15:16 sean-k-mooney stephenfin: by packing numa nodes before moving on
17:15:22 sean-k-mooney we keep the rest free
17:15:33 sean-k-mooney so vms that need all teh ram or cpus on a numa node can boot
17:15:44 stephenfin ah, yeah I don't think that's a good argument for the unpinned case
17:15:45 sean-k-mooney if we spread then that makes large vms less likely to fit
17:15:56 stephenfin spreading makes sense there IMO
17:15:59 sean-k-mooney unpinned float so ya not an issue
17:16:03 stephenfin yup
17:16:16 claudiub in our scenario, large vms is not a concern, since we have pretty large hosts, so in our case, it would work better for a spread-out approach. But indeed, not everyone is the same. Could this be a config option then?
17:16:20 stephenfin if you don't spread, you'll basically never end up on node 1, as claudiub is seeing
17:16:26 stephenfin or node N > 0
17:16:43 sean-k-mooney claudiub: yep it could be a config option or it coudl be done via aggreate metadata
17:17:00 stephenfin I don't think that's needed or unpinned
17:17:03 stephenfin *for
17:17:27 sean-k-mooney stephenfin: well this is needed for all numa guests pinned or unpinned
17:17:33 claudiub hmm, aggregate metadata also sounds interesting. it could have best of both worlds
17:17:35 sean-k-mooney but its not needed for non numa instnaces
17:17:56 sean-k-mooney claudiub: my concern is if we codify this as a feature we need to supprot it with placment in the future
17:17:58 stephenfin pinned guests can already use NUMA nodes > 0
17:18:16 sean-k-mooney claudiub: we can do that but that means we need to do the same behavior by sorting the allocation candiates
17:18:48 sean-k-mooney yep they can but that not really the issue
17:19:03 claudiub hm, I am a bit outside the loop with the placement api, but doesn't the NUMAPlacementFilter also use the numa_fit_instance_to_host function?
17:19:30 sean-k-mooney claudiub: yes currently we are not usign placmnet for numa and wont be for a few release
17:19:30 claudiub or whatever the fitler name was. :)
17:19:41 claudiub oh ok, gotcha.
17:19:46 stephenfin can we back up and say why any of this needs a aggregate metadata filter
17:19:56 stephenfin we're not trying to change stack/spread behavior for pinned instances, right?
17:20:00 sean-k-mooney claudiub: the concern is the more featers the filter has the more we need to port to the plamcent version
17:20:10 sean-k-mooney stephenfin: no new aggreate filter
17:20:18 stephenfin only unpinned NUMA instances, because those are all landing on NUMA node 0
17:20:30 stephenfin sorry, an aggregate metadata key
17:20:40 sean-k-mooney stephenfin: nope hugepages and all other numa instnace land there too
17:20:40 stephenfin <sean-k-mooney> claudiub: yep it could be a config option or it coudl be done via aggreate metadata
17:20:42 stephenfin ^ that
17:21:14 stephenfin a page with hugepages and no pinning *is* an unpinned NUMA instance
17:21:16 sean-k-mooney stephenfin: if we do it per host it has to be accounted for on live migration and we have the same proble we had wtih PCPUs and hyperthreading
17:21:41 stephenfin if we do what per hose?
17:21:41 stephenfin *host
17:21:41 claudiub Hm, I'm wondering why the current implementation is: "Hey HostState, I'm a request spec, pls fit me", rather than: "Hey HostState, Placement told me to sit in your X numa node."

Earlier   Later