Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-10
16:30:17 spatel I can't modify existing VM correct?
16:30:23 bauzas spatel: no, you can't
16:30:31 spatel Perfect! now i got it what to do.
16:30:32 bauzas flavor is embedded into the instance data
16:30:59 bauzas if you modify a flavor, the instances that booted from that flavor won't magically update
16:31:00 spatel I will add fresh compute nodes with no HugePage and create new flavor without Pages
16:31:14 bauzas you'll be required to resize with another flavor
16:31:31 spatel bauzas I totally understand.. you want just change flavor and it will work magically :)
16:32:17 spatel This is what happened last week, one of memory module die which crash my whole compute nodes because of HugePage requirement :(
16:34:06 bauzas hah, that's a common failure
16:34:22 bauzas and yeah, relying on RAM can be dangerous
16:35:03 spatel Yes.. because of that crash it created loop in my switch (I don't know how but it lock up my switch because of STP)
16:35:31 spatel Just trying to re-produce this issue with multiple variable to see if i can re-create
18:32:06 sean-k-mooney[m] spatel you might want to look at the numa blancing config option
18:32:55 sean-k-mooney[m] packing_host_numa_cells_allocation_strategy
18:34:10 sean-k-mooney[m] spatel by the way if you are using cpu pinning but not hugepages you should set hw:mem_page_size=small
18:34:35 sean-k-mooney[m] if you dont then the vms will randomly get kill due to OOM events
18:36:58 spatel hmm is that a new option packing_host_numa_cells_allocation_strategy ?
18:37:32 spatel sean-k-mooney[m] this is interesting - by the way if you are using cpu pinning but not hugepages you should set hw:mem_page_size=small
18:37:48 spatel does it going to work if i don't configure hugepage in grub?
18:38:43 sean-k-mooney[m] packing_host_numa_cells_allocation_strategy is new and we backported it
18:39:31 sean-k-mooney[m] i think it was added in zed or yoga we changed the default to spread this cycle or last
18:39:58 sean-k-mooney[m] packing_host_numa_cells_allocation_strategy goes in the compute section of the nova.conf i belvie
18:40:16 sean-k-mooney[m] https://docs.openstack.org/nova/latest/configuration/config.html#compute.packing_host_numa_cells_allocation_strategy
18:41:16 sean-k-mooney[m] spatel: if you are using cpu pinnign hw:mem_page_size need to be set to some valid value to turn on numa aware memory allocation
18:41:42 sean-k-mooney[m] if you dont we will scudle based on the gloabl not numa local memory
18:42:12 sean-k-mooney[m] the OOM reaper in the kernel operates at the numa level
18:42:43 spatel ohhhh
18:43:27 spatel I know what you saying.. to run workload in NUMA we need to set hw:mem_page_size
18:43:46 sean-k-mooney[m] so the end result of not setting it is we will overcommit the numa node since we are only schduling based on the cpu in that case
18:44:01 sean-k-mooney[m] ya basically
18:44:50 sean-k-mooney[m] i have wanted to enforce this for a while but there were concerns that operators are depending on the incorrect behavior
18:45:36 sean-k-mooney[m] i have wanted to make hw:mem_page_size=any the default if you have a numa toplogy in the guest and dont set anything
18:45:37 sean-k-mooney[m] any is the same as small excpet it allows you to override it in the image
18:47:48 spatel hmmm
18:51:34 sean-k-mooney[m] the simple way to think about it is if its a numa vm you should set a mem_page_size as well
18:52:43 sean-k-mooney[m] well or use file backed memory but that is not a configuration that many people use
18:56:37 spatel I will do it..
19:14:55 melwitt bauzas: done, thanks for reminding
22:00:47 gmann bauzas: yes, I am planning to progress on 864594 but let's see if i can push it before FF
#openstack-nova - 2023-02-11
10:13:09 opendevreview Takashi Natsume proposed openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489
11:59:19 opendevreview Alexey Stupnikov proposed openstack/nova master: Fix logging in MemEncryption-related checks https://review.opendev.org/c/openstack/nova/+/873388
#openstack-nova - 2023-02-12
04:57:13 opendevreview Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594
20:28:52 opendevreview Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594
#openstack-nova - 2023-02-13
01:49:33 opendevreview Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594
02:33:26 opendevreview Nobuhiro MIKI proposed openstack/nova master: libvirt: Add 'COMPUTE_ADDRESS_SPACE_*' traits support https://review.opendev.org/c/openstack/nova/+/873221
04:12:10 opendevreview Ghanshyam proposed openstack/nova master: Add service role in nova policy https://review.opendev.org/c/openstack/nova/+/864594
07:58:14 ihti[m] Hi, question: Does anyone know why we are using warnings module instead of LOG.warning() e.g for deprecations in oslo_policy. It breaks logs parsing for us, as the warnings emitted deviates from the standard logging formatting.
08:09:58 opendevreview Amit Uniyal proposed openstack/nova master: Remove "see nova-manage.log" string from console https://review.opendev.org/c/openstack/nova/+/873496
08:38:12 bauzas morning Nova
08:53:16 gibi bauzas: o/
08:54:41 bauzas currently writing some PTL nomination email, but will be back in 20 mins
09:07:45 gibi ack
09:15:42 opendevreview Jorge San Emeterio proposed openstack/nova master: WIP: Look for cpu controller on cgroups v2 https://review.opendev.org/c/openstack/nova/+/873127
10:33:07 opendevreview Amit Uniyal proposed openstack/nova master: Remove "see nova-manage.log" string from console https://review.opendev.org/c/openstack/nova/+/873496
10:38:17 opendevreview Jorge San Emeterio proposed openstack/nova master: WIP: Look for cpu controller on cgroups v2 https://review.opendev.org/c/openstack/nova/+/873127
10:58:14 bauzas gibi: can't imagine https://review.opendev.org/c/openstack/tempest/+/873300 is still not merged
11:05:09 kashyap There is an emoji for that in Slack!
11:06:05 kashyap 🤯
11:06:05 kashyap 🤯
11:17:46 bauzas gibi: https://zuul.openstack.org/status#873300 will fuckingly fuckering fucker fail again
11:24:28 bauzas gibi: we may need to skip the ceph-multistore job by today if we wanna just merge some stuff in nova
11:41:17 sean-k-mooney bauzas: push a patch to make it non voting and add those test to the exclude regex and you can ping me to help merge it if you like
11:43:12 opendevreview Jorge San Emeterio proposed openstack/nova master: WIP: Look for cpu controller on cgroups v2 https://review.opendev.org/c/openstack/nova/+/873127
12:34:21 opendevreview Edward Hope-Morley proposed openstack/nova stable/yoga: ignore deleted server groups in validation https://review.opendev.org/c/openstack/nova/+/867989
13:11:34 opendevreview Konrad Gube proposed openstack/nova master: Use Cinder's os-extend_volume_completion volume action. https://review.opendev.org/c/openstack/nova/+/873560
13:34:32 gibi bauzas: we can make ceph non voting for now, but it does not solve the issue that it seems to be impossible to merge anything in tempest now
13:34:55 bauzas gibi: let's see what's the latest result then
13:35:35 bauzas yet again a f.f.f. failure https://review.opendev.org/c/openstack/tempest/+/873300
13:36:27 gibi bauzas: it seems you commented recheck at 12:23 and zuul voted at 12:37 so I guess you rechecked before the job finished
13:42:19 gibi I queued up again
13:53:05 opendevreview Amit Uniyal proposed openstack/nova master: nova-manage volume_attachment issues https://review.opendev.org/c/openstack/nova/+/873561
14:00:24 opendevreview Amit Uniyal proposed openstack/nova master: nova-manage volume_attachment issues https://review.opendev.org/c/openstack/nova/+/873561
14:11:05 dansmith bauzas: I wonder if we're getting to the point where we should make the ceph job n-v
14:11:13 dansmith I really hate to do that, but this is kinda nuts
14:11:27 bauzas dansmith: yeah, I dunno what to say
14:11:50 dansmith there are other such fixes in the queue and we've been unable to land them over the weekend as well
14:12:02 bauzas yup
14:13:42 bauzas so, how could we merge the changes that were accepted ? that's only my concern :)
14:14:42 dansmith the nova changes you mean? making the ceph job n-v at least gives you a chance to pass the other jobs
14:14:48 dansmith it'll still be a struggle, but..
14:43:23 gibi dansmith, bauzas: can we ask infra to promot the tempest fix by skiping the gate? Or is it too dangerous?
14:44:08 dansmith we could, but I think given that it's not landing due to all the other broken, it's not really the emergent situation that normally demands that sort of thing
14:44:34 dansmith I think we reserve that for situations where we've got two things wedged against each other and one has to merge without tests
14:47:40 bauzas https://zuul.openstack.org/status#873300 looks like tempest-full-py3 can try twice
14:49:27 bauzas gibi: dansmith: so, given the time we have before FF, what should we be doing ?
14:49:33 bauzas punting the FF ?
14:49:45 bauzas or awaiting for the change to be merged ?
14:49:59 dansmith the 100% fail is only in the ceph job, so you can just make that job n-v for the moment and watch it carefully
14:50:00 bauzas or non-voting the job ?
14:50:05 dansmith but you'll still suffer all the other issues
14:50:16 dansmith how much is still needing to land for FF?
14:50:55 bauzas https://etherpad.opendev.org/p/nova-antelope-blueprint-status at least two series
14:51:47 bauzas + https://review.opendev.org/c/openstack/nova/+/872413
14:53:30 gibi I'm OK to make the ceph job non voting to land the pending nova series
14:54:04 bauzas ok, then I'll create the change
14:54:17 bauzas and I'll also create a revert one
14:54:40 dansmith just make the revert one depends-on the fix and we can go ahead and have it pre-approved
14:54:45 gibi yeah and you can make a depends on in the revert to point to the tempest fix
14:54:57 bauzas cool

Earlier   Later