| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-06-14 | |||
| 11:33:59 | sean-k-mooney | would we consider that backportable if it was opt in and both implemations could be in the code in parallel even fi only one would be used on a compute host at a time | |
| 11:34:26 | gibi | why we need both? does yours produce a different result than the current? | |
| 11:34:41 | sean-k-mooney | mine does not do everythin the current one does | |
| 11:34:59 | gibi | yepp, then we might need a config flag | |
| 11:35:04 | sean-k-mooney | so im wondering if we coudl make it incremental | |
| 11:35:40 | sean-k-mooney | and also have the option to run both as filters | |
| 11:35:55 | sean-k-mooney | so mine currently only handeles hugepages and cpu pinning | |
| 11:36:17 | sean-k-mooney | its missing pci devices and mixed cpus and maybe pmem | |
| 11:36:38 | sean-k-mooney | pmem enabels a numa toplogy but i dont think we provide affintiy for pmem | |
| 11:36:57 | sean-k-mooney | so its reall just mixed mode and pci devices that woudl initally be missing | |
| 11:37:38 | sean-k-mooney | my tought is you coudl enable both filters and have the fast one run first so you woudl only validate the slow hosts if we knew the cpu and ram request were valid | |
| 11:38:07 | sean-k-mooney | s/slow hosts/hosts with the slow version/ | |
| 11:38:32 | gibi | overall I'm OK to make it incremental or even selectable, but I'm not sure stable cores will like it | |
| 11:38:51 | gibi | this will include a new bounch of code to stable branches | |
| 11:38:54 | sean-k-mooney | so basicaly you woudl do pci_passthroughfilter,numa_v2,numa_toplogy_filter | |
| 11:39:01 | gibi | that is liability | |
| 11:39:09 | sean-k-mooney | yep | |
| 11:39:12 | sean-k-mooney | it is | |
| 11:39:34 | gibi | what if we just add what we have to master as selectable, then we improve on it later on master | |
| 11:39:50 | sean-k-mooney | we could yes | |
| 11:40:07 | sean-k-mooney | im just trying to think if there is anything we can do to adress the bug report on older branches too | |
| 11:40:52 | sean-k-mooney | i dont think there is anythign trivial that can be fixed on the older brances to make the perfromance accpetabel. | |
| 11:41:39 | gibi | I don't know. maybe we can look at the reproducer in a profiler | |
| 11:42:11 | sean-k-mooney | i think its to do with our usage so itertools permuations | |
| 11:42:33 | sean-k-mooney | we do a liniar loop over all permuations of numa nodes when trying to fit the instance | |
| 11:43:06 | sean-k-mooney | we early out once the instance fits but that is really not efficent with more then a hand full of numa nodes | |
| 11:44:37 | sean-k-mooney | gibi: i can try an take a look at it breifly | |
| 11:45:11 | gibi | sean-k-mooney: only if you want :) I'm not asking to take this | |
| 11:45:29 | gibi | I'm not promising either that I can fire up a profiler today | |
| 11:45:36 | sean-k-mooney | hehe | |
| 11:45:44 | sean-k-mooney | this is not new | |
| 11:46:01 | sean-k-mooney | so i dont think it supper high priority but i was half way thorugh fixing up my vdpa patches | |
| 11:46:07 | sean-k-mooney | so i want to finish those today | |
| 11:46:30 | sean-k-mooney | so im somehwat interested in is there a minimal fix we can do to make ti better | |
| 11:46:39 | sean-k-mooney | i just dont want to spend all day looking at it | |
| 11:46:48 | sean-k-mooney | so ill give it tilll the top of the hour | |
| 11:47:07 | sean-k-mooney | we can take another look later in the week or next week | |
| 11:47:36 | gibi | ack | |
| 11:57:21 | opendevreview | Merged openstack/nova stable/wallaby: Add service version check workaround for FFU https://review.opendev.org/c/openstack/nova/+/844202 | |
| 11:59:05 | gibi | :) | |
| 11:59:41 | sean-k-mooney | its amazing how much simpler nova code becomes without eventlets | |
| 12:01:11 | sean-k-mooney | actully there might be a small tweak we can do | |
| 12:02:08 | sean-k-mooney | we currently do this | |
| 12:02:10 | sean-k-mooney | for host_cell_perm in itertools.permutations( | |
| 12:02:12 | sean-k-mooney | host_cells, len(instance_topology) | |
| 12:02:14 | sean-k-mooney | ): | |
| 12:02:26 | sean-k-mooney | so we get the next perumation for the full pinning | |
| 12:02:47 | sean-k-mooney | i wonder if we could do this one numa node at a time | |
| 12:03:10 | sean-k-mooney | so loop over the instnace numa cells | |
| 12:03:33 | sean-k-mooney | and try to pin them one at a time | |
| 12:03:36 | sean-k-mooney | then try to pin the rest | |
| 12:04:04 | sean-k-mooney | i think that would be a lot faster | |
| 12:04:36 | sean-k-mooney | as currently if the first 7 numa nodes are full we have to try all permuations for the 7 numa nodes before we will try the 8th i think | |
| 12:05:37 | sean-k-mooney | if i refactor this we will do 7 test for the first guest numa node it will match on the 8th numa node and then we will try fiting the next numa node | |
| 12:05:48 | gibi | would that just do the permutation generation with loops? sure the order of the search would be different | |
| 12:06:10 | gibi | so it might optimize for the current csae | |
| 12:06:11 | gibi | case | |
| 12:06:31 | sean-k-mooney | the worst case performace woudl still be the same but i think it would imporve best and average case | |
| 12:11:03 | gibi | so we optimize based on some heuristics that the first numa nodes are more likely to be filled than the later numa nodes | |
| 12:11:24 | sean-k-mooney | well until master it did a liniar search | |
| 12:11:37 | sean-k-mooney | so yes the first numa nodes were always filled first deterministically | |
| 12:11:51 | sean-k-mooney | we recently added numa node blancing | |
| 12:12:22 | gibi | true, so using the old linear fill, it make sense to add a heuristics to the search to | |
| 12:12:25 | gibi | o | |
| 12:12:44 | sean-k-mooney | also my entire system just hard locked up while i was debuging that even with it paused | |
| 12:12:51 | gibi | wondering if just simply using permutations(reversed(host_cells)) would be enough to optimize too | |
| 12:12:53 | sean-k-mooney | perhaps i should close some broser tabs | |
| 12:13:10 | sean-k-mooney | well we are not sorting the host_cells | |
| 12:13:25 | gibi | the other day I run out of memory on my laptop, so now I added some swap | |
| 12:13:29 | sean-k-mooney | based on aviable ram disk pci devices and if you asked for them | |
| 12:14:00 | sean-k-mooney | i ran out of swap but still had 8GB of ram free | |
| 12:14:31 | sean-k-mooney | i do have a 2 node devstack and openshift running in 3 8G vms currently too | |
| 12:16:17 | sean-k-mooney | oh my email client is only using 4G of ram today that is nice of it. it was using 15 last week... | |
| 12:17:34 | sean-k-mooney | 1910527.379325] Out of memory: Killed process 1338765 (.qemu-system-x8) total-vm:13382508kB, anon-rss:7832544kB, file-rss:0kB, shmem-rss:4kB, UID:0 pgtables:17324kB oom_score_adj:0 | |
| 12:17:42 | sean-k-mooney | yep i ran out of memory ill stop the vms for now | |
| 12:18:44 | sean-k-mooney | thats totally a good sign for the effiency fo this code right | |
| 12:29:30 | sean-k-mooney | gibi: :) | |
| 12:29:33 | sean-k-mooney | import nova.conf | |
| 12:29:35 | sean-k-mooney | CONF = nova.conf.CONF | |
| 12:29:37 | sean-k-mooney | CONF.compute.packing_host_numa_cells_allocation_strategy = False | |
| 12:29:39 | sean-k-mooney | that fixes the issue | |
| 12:29:58 | sean-k-mooney | we disable the numa blancing by defualt for "backwards compatiblity" | |
| 12:30:31 | sean-k-mooney | gibi: if you enable it when we sort by free memory all the used nodes go to the end | |
| 12:30:41 | sean-k-mooney | so the first permuation fits | |
| 12:31:17 | sean-k-mooney | [13:30:55]➜ time python t.py | |
| 12:31:18 | sean-k-mooney | InstanceNUMATopology(cells=[InstanceNUMACell(8),InstanceNUMACell(9),InstanceNUMACell(10),InstanceNUMACell(11),InstanceNUMACell(12),InstanceNUMACell(13),InstanceNUMACell(14)],emulator_threads_policy=None,id=<?>,instance_uuid=<?>) | |
| 12:31:20 | sean-k-mooney | real 0m1.489s | |
| 12:31:22 | sean-k-mooney | user 0m1.373s | |
| 12:31:24 | sean-k-mooney | sys 0m0.095s | |
| 12:31:44 | sean-k-mooney | gibi: not as fast as my out of tree version but pretty close | |
| 12:31:51 | gibi | ahh | |
| 12:31:53 | sean-k-mooney | and fully feature complete | |
| 12:32:02 | gibi | that is an easy workaround for the particular case | |
| 12:32:43 | sean-k-mooney | well the spread approch is generaly beter provided you dont need really large vms | |
| 12:32:56 | sean-k-mooney | that depend on fully filling the numa nodes to spawn | |
| 12:32:58 | sean-k-mooney | but yes | |
| 12:38:31 | sean-k-mooney | gibi: ok added a comment https://bugs.launchpad.net/nova/+bug/1978372/comments/6 | |
| 12:38:59 | sean-k-mooney | gibi: we have already started backporting the pack/spread behavior to xena they reported it on wallaby | |
| 12:39:22 | sean-k-mooney | so if we just continue the backport of the sorting behavior we could close it as a dupe i guess | |
| 12:39:56 | gibi | sean-k-mooney: thanks | |