| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-10 | |||
| 16:02:55 | bauzas | melwitt: could you please remove your -2 on https://review.opendev.org/c/openstack/nova/+/863177 (the ironic-vnc-console blueprint got accepted for the cycle) | |
| 16:06:41 | bauzas | gibi: I think we rounded on all the possible blueprints we have | |
| 16:07:13 | gibi | bauzas: yeah, I don't have brainpower to look at the ironic vnc one | |
| 16:07:16 | bauzas | we can try to take a look at https://review.opendev.org/c/openstack/nova/+/863177 possibly but sounds a bit optimistic | |
| 16:07:29 | gibi | yeah | |
| 16:08:14 | bauzas | gmann: planning to progress on https://review.opendev.org/c/openstack/nova/+/864594 ? | |
| 16:08:26 | gibi | I think I won't start anything big any more today but I'm still around for a bit if specific review is needed | |
| 16:08:34 | bauzas | me too, I'm done for today | |
| 16:08:50 | bauzas | I'll start reviewing the maxphysnet series today, but I'm not an expert in this | |
| 16:09:01 | bauzas | s//maxphysaddr | |
| 16:10:05 | gibi | ohh that one I can take a look | |
| 16:10:08 | gibi | I reviewd the spec there | |
| 16:12:43 | spatel | sean-k-mooney Hi, I have a question related HugePages, currently i am running SRIOV + CPU pinning + HugePages | |
| 16:13:27 | spatel | Lets say i don't want to use HugePages. Does that possible? | |
| 16:13:41 | bauzas | spatel: sean-k-mooney is on PTO until wed | |
| 16:13:49 | bauzas | lemme try to answer you | |
| 16:14:15 | bauzas | spatel: yes, it's possible to have CPU pinning without huge pages | |
| 16:14:26 | bauzas | but I guess you have running workloads ? | |
| 16:15:00 | spatel | Yes | |
| 16:15:39 | spatel | I am deploying new cloud and planning to not use HugePage. We had some incident in past related memory cause strange issue. | |
| 16:15:43 | bauzas | ok, so, do you want to tune off hugepages for all your computes but a subset ? | |
| 16:15:53 | bauzas | ah | |
| 16:16:00 | spatel | New cloud with no HugePage at all.. | |
| 16:16:17 | bauzas | ok, then you just need to use flavors that don't request hugepages | |
| 16:16:37 | spatel | This is what i have currently in my cloud - intel_iommu=on iommu=pt hugepagesz=2M hugepages=30000 transparent_hugepage=never | |
| 16:16:57 | spatel | Thinking to remove hugepages and change flavor | |
| 16:17:46 | bauzas | https://docs.openstack.org/nova/latest/admin/cpu-topologies.html | |
| 16:18:40 | bauzas | I need to verify one bit, sec | |
| 16:18:49 | spatel | sure! | |
| 16:19:18 | bauzas | https://docs.openstack.org/nova/latest/admin/huge-pages.html | |
| 16:19:53 | bauzas | so, say you no longer ask for hugepages, it won't request a NUMA topolgy | |
| 16:20:18 | spatel | That was my next question.. How numa play with HugePages? | |
| 16:20:38 | spatel | We want our workload schedule in single NUMA zone | |
| 16:21:54 | bauzas | so you want NUMA without hugepages | |
| 16:22:48 | spatel | Yes | |
| 16:23:27 | spatel | I believe openstack automatically schedule workload according NUNA correct? | |
| 16:23:33 | bauzas | see that doc https://docs.openstack.org/nova/latest/admin/cpu-topologies.html#customizing-instance-numa-placement-policies | |
| 16:23:40 | gibi | bauzas: the max_phy_address patch is just the start of the series. we will need more patches there. | |
| 16:24:00 | bauzas | either you explicitly specific a NUMA topology for your guest or you make it implicit with cpu pinning or hugepages flavor extra specs | |
| 16:24:09 | bauzas | gibi: ack | |
| 16:24:23 | spatel | This is my flavor properties - hw:cpu_policy='dedicated', hw:cpu_sockets='2', hw:cpu_threads='2', hw:mem_page_size='large', hw:pci_numa_affinity_policy='preferred', sriov='true' | |
| 16:24:57 | bauzas | spatel: so, you'd just remove the mention of using large pages | |
| 16:24:58 | spatel | I didn't tell my workload about where to schedule but it always put my VMs on single NUMA zone | |
| 16:25:27 | spatel | Yep! that is what i am thinking, remove from grub and flavor and it should be fine. | |
| 16:26:15 | bauzas | the fact it goes to the same NUMA cell is because of the policy https://docs.openstack.org/nova/latest/configuration/extra-specs.html#hw:pci_numa_affinity_policy | |
| 16:27:14 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Fix logging in MemEncryption-related checks https://review.opendev.org/c/openstack/nova/+/873388 | |
| 16:28:17 | spatel | bauzas worth running some test.. i will pick one compute and try to play and see how it goes | |
| 16:28:37 | spatel | I love numa but it has some downside... | |
| 16:28:39 | bauzas | spatel: you'll need to modify the nova config, not only grub | |
| 16:29:01 | spatel | Yes..i will start with fresh compute node.. i am not going to touch existing one. | |
| 16:29:29 | bauzas | actually I'm wrong | |
| 16:29:41 | spatel | Ouch!! now what? | |
| 16:29:44 | bauzas | no nova conf is required for page management | |
| 16:29:54 | bauzas | it just gets it from what we have | |
| 16:29:54 | spatel | oh! | |
| 16:30:17 | bauzas | spatel: read the docs I gave to you | |
| 16:30:17 | spatel | I can't modify existing VM correct? | |
| 16:30:23 | bauzas | spatel: no, you can't | |
| 16:30:31 | spatel | Perfect! now i got it what to do. | |
| 16:30:32 | bauzas | flavor is embedded into the instance data | |
| 16:30:59 | bauzas | if you modify a flavor, the instances that booted from that flavor won't magically update | |
| 16:31:00 | spatel | I will add fresh compute nodes with no HugePage and create new flavor without Pages | |
| 16:31:14 | bauzas | you'll be required to resize with another flavor | |
| 16:31:31 | spatel | bauzas I totally understand.. you want just change flavor and it will work magically :) | |
| 16:32:17 | spatel | This is what happened last week, one of memory module die which crash my whole compute nodes because of HugePage requirement :( | |
| 16:34:06 | bauzas | hah, that's a common failure | |
| 16:34:22 | bauzas | and yeah, relying on RAM can be dangerous | |
| 16:35:03 | spatel | Yes.. because of that crash it created loop in my switch (I don't know how but it lock up my switch because of STP) | |
| 16:35:31 | spatel | Just trying to re-produce this issue with multiple variable to see if i can re-create | |
| 18:32:06 | sean-k-mooney[m] | spatel you might want to look at the numa blancing config option | |
| 18:32:55 | sean-k-mooney[m] | packing_host_numa_cells_allocation_strategy | |
| 18:34:10 | sean-k-mooney[m] | spatel by the way if you are using cpu pinning but not hugepages you should set hw:mem_page_size=small | |
| 18:34:35 | sean-k-mooney[m] | if you dont then the vms will randomly get kill due to OOM events | |
| 18:36:58 | spatel | hmm is that a new option packing_host_numa_cells_allocation_strategy ? | |
| 18:37:32 | spatel | sean-k-mooney[m] this is interesting - by the way if you are using cpu pinning but not hugepages you should set hw:mem_page_size=small | |
| 18:37:48 | spatel | does it going to work if i don't configure hugepage in grub? | |
| 18:38:43 | sean-k-mooney[m] | packing_host_numa_cells_allocation_strategy is new and we backported it | |
| 18:39:31 | sean-k-mooney[m] | i think it was added in zed or yoga we changed the default to spread this cycle or last | |
| 18:39:58 | sean-k-mooney[m] | packing_host_numa_cells_allocation_strategy goes in the compute section of the nova.conf i belvie | |
| 18:40:16 | sean-k-mooney[m] | https://docs.openstack.org/nova/latest/configuration/config.html#compute.packing_host_numa_cells_allocation_strategy | |
| 18:41:16 | sean-k-mooney[m] | spatel: if you are using cpu pinnign hw:mem_page_size need to be set to some valid value to turn on numa aware memory allocation | |
| 18:41:42 | sean-k-mooney[m] | if you dont we will scudle based on the gloabl not numa local memory | |
| 18:42:12 | sean-k-mooney[m] | the OOM reaper in the kernel operates at the numa level | |
| 18:42:43 | spatel | ohhhh | |
| 18:43:27 | spatel | I know what you saying.. to run workload in NUMA we need to set hw:mem_page_size | |
| 18:43:46 | sean-k-mooney[m] | so the end result of not setting it is we will overcommit the numa node since we are only schduling based on the cpu in that case | |
| 18:44:01 | sean-k-mooney[m] | ya basically | |
| 18:44:50 | sean-k-mooney[m] | i have wanted to enforce this for a while but there were concerns that operators are depending on the incorrect behavior | |
| 18:45:36 | sean-k-mooney[m] | i have wanted to make hw:mem_page_size=any the default if you have a numa toplogy in the guest and dont set anything | |
| 18:45:37 | sean-k-mooney[m] | any is the same as small excpet it allows you to override it in the image | |
| 18:47:48 | spatel | hmmm | |
| 18:51:34 | sean-k-mooney[m] | the simple way to think about it is if its a numa vm you should set a mem_page_size as well | |
| 18:52:43 | sean-k-mooney[m] | well or use file backed memory but that is not a configuration that many people use | |
| 18:56:37 | spatel | I will do it.. | |
| 19:14:55 | melwitt | bauzas: done, thanks for reminding | |
| 22:00:47 | gmann | bauzas: yes, I am planning to progress on 864594 but let's see if i can push it before FF | |
| #openstack-nova - 2023-02-11 | |||
| 10:13:09 | opendevreview | Takashi Natsume proposed openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489 | |
| 11:59:19 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Fix logging in MemEncryption-related checks https://review.opendev.org/c/openstack/nova/+/873388 | |
| #openstack-nova - 2023-02-12 | |||
| 04:57:13 | opendevreview | Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594 | |
| 20:28:52 | opendevreview | Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594 | |