| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-06-18 | |||
| 08:10:44 | sean-k-mooney | https://bugzilla.redhat.com/show_bug.cgi?id=1651994 | |
| 08:11:01 | sean-k-mooney | oh thats private well that is what it was tracking | |
| 08:11:43 | sean-k-mooney | anyway 16.0 which is based on train was released on 8.1 16.1 will be on 8.2 | |
| 08:12:31 | sean-k-mooney | so customer are now in the situation where the default model that nova will select is deprecaed and there is no way via the api to change this on existing instance other than rebuild | |
| 08:13:08 | sean-k-mooney | we wont be able to back port this new command or whatever it will be but it might help in the future | |
| 08:39:31 | openstackgerrit | Wenping Song proposed openstack/nova master: delete sub resource provider when delete resource provider https://review.opendev.org/719163 | |
| 09:54:40 | openstackgerrit | Wenping Song proposed openstack/nova master: delete sub resource provider when delete resource provider https://review.opendev.org/719163 | |
| 10:38:26 | arne_wiebalck | One of the main points for user adoption probably is that whenever sth needs to change, this should be feasible without having the Ironic admin to re-clean nodes or the Nova admin to add new flavors (with 5000 nodes in Ironic we have around 150 resource classes and hence 150 flavors). This is why we came up with the idea of a "kickstart" driver where Ironic uses a kickstart/preseed file to do the | |
| 10:38:26 | arne_wiebalck | the whole discussion. | |
| 10:38:26 | arne_wiebalck | to deploy. However, other users (like the Ceph team) want to partition/RAID their physical instances in a specific way as needed for their service. There is no convenient way to express this is via Nova/Ironic at the moment. So, they reinstall these instances after initial deployment through Nova/Ironic once more with kickstart. This double installation is what we'd like to get rid of, and this triggered | |
| 10:38:26 | arne_wiebalck | Let me add a little background info: physical instances at CERN are all done via Nova and Ironic. The main user is OpenStack itself and we rely on Ironic's software RAID and standard cloud images | |
| 10:38:26 | arne_wiebalck | TheJulia: dansmith sean-k-mooney Sorry, I missed the discussion yesterday. | |
| 10:38:27 | arne_wiebalck | deployment (and skip the usual image deployment) as it would provide the user with the same flexibility as there is now. But since we would like to keep Nova in the mix (for various reasons), we may need some tooling help on the Nova side. | |
| 12:36:20 | aarents | Hi nova, | |
| 12:38:02 | aarents | sean-k-mooney: dansmith Let me know if it needs further amend on https://review.opendev.org/#/c/736169/ https://review.opendev.org/#/c/734776/ thanks! | |
| 12:39:07 | openstackgerrit | Alexandre Arents proposed openstack/nova master: libvirt: ensure disk_over_commit is not negative https://review.opendev.org/719008 | |
| 12:39:29 | sean-k-mooney | ^ that should be done by the config option | |
| 12:39:33 | sean-k-mooney | we set min 0 i think | |
| 12:40:41 | sean-k-mooney | oh that is not the ratio | |
| 12:41:09 | sean-k-mooney | for that to be negitive the size on disk would have to be larger then the virtual size | |
| 12:41:39 | aarents | sean-k-mooney: yes | |
| 12:41:57 | sean-k-mooney | does it being negitive break something | |
| 12:43:06 | sean-k-mooney | ah i see | |
| 12:43:18 | aarents | It just mislead calcuation of available_disk_least on which rely disk_filter | |
| 12:47:43 | sean-k-mooney | ya clamping the value shoudl be fine. | |
| 13:12:54 | rmart04 | Hey all, I'm wondering if anyone could help me. Not strictly dev specific, but I'm having trouble with NUMA information being passed through to my virtual machines. /sys/bus/pci/devices*/numa_node always = -1. I'm running Rocky on C7, with numa_nodes=2 and cpu_sockets=2. | |
| 13:13:08 | rmart04 | and cpu pinning policy set to dedicated | |
| 13:14:17 | sean-k-mooney | is the -1 on the host or in the guest | |
| 13:14:23 | sean-k-mooney | if its in the guest that is expected | |
| 13:14:26 | rmart04 | In the guest | |
| 13:14:52 | sean-k-mooney | so we do not currently create a pcie root complex per numa node | |
| 13:15:13 | sean-k-mooney | since there is only one pci root all devices are childerne of that root | |
| 13:15:42 | sean-k-mooney | wehn we do passthough we dont affiniteis the pci device to the virutal numa node of the guest | |
| 13:15:59 | sean-k-mooney | so it is reported as -1 meaning no numa affinity in the guest | |
| 13:16:55 | sean-k-mooney | to change that we would likely have to use the q35 machine type and create a pcie root complete per numa node then add the passthough deivice to the correct pci root | |
| 13:17:45 | sean-k-mooney | that has other implciation mainliy that on move operation either we have to allow the toploty to change or we have to limit the host we can select to maintain the current toplogy | |
| 13:18:14 | sean-k-mooney | if we allow the toplogy to change the virtual pci address of the devices in the guest would also change | |
| 13:19:16 | sean-k-mooney | rmart04: but yes if you have a multi numa node guest this can result in cross numa traffic worse case twice because you dont know the numa affinity of the device | |
| 13:20:18 | openstackgerrit | Merged openstack/nova master: libvirt: Mark e1000e VIF as supported https://review.opendev.org/734777 | |
| 13:22:32 | rmart04 | OK, appreciate all the info SeanKMooney. I guess another way around this is to split the host into two guests, one on each numa node with their associated pci-passthrough devices (GPUs). Currently I appear to be blocked on this by my older kernel. 3.10. I bump into an issue allocating memory from the second NUMA node for the second machine. I believe this is fixed in 4.14. | |
| 13:23:06 | rmart04 | qq, you mention the q35 machine type, what type do we use by default? | |
| 13:25:51 | sean-k-mooney | rmart04: if the guest has a numa toploty we do not allow its memory to come form a remote numa node by design | |
| 13:26:01 | sean-k-mooney | rmart04: we use pc | |
| 13:26:23 | sean-k-mooney | or pc-i440fx | |
| 13:26:27 | sean-k-mooney | something like that | |
| 13:26:43 | sean-k-mooney | rmart04: what version of openstack are you using | |
| 13:27:02 | rmart04 | Rocky (Stein upgrade this weekend) | |
| 13:27:26 | sean-k-mooney | do you have gpus on all host numa nodes | |
| 13:27:30 | sean-k-mooney | or just numa 0 | |
| 13:27:42 | rmart04 | Yep, 2 sockets, 8 GPUs | |
| 13:27:50 | rmart04 | 4 each | |
| 13:28:16 | sean-k-mooney | ok what iw was going to say is you might need to use nuam_policy=preferred in the alias | |
| 13:28:21 | sean-k-mooney | if you did not have them split | |
| 13:28:54 | sean-k-mooney | if you do then yes 2 vms with 1 numa each and the default legacy polciy which enforce numa affintiy between cpu/memory and the pci device is what you will want | |
| 13:29:42 | sean-k-mooney | you can create a dual numa guest but the limitation is you will know know what numa node in the guest maps to the actull location of the device on the host | |
| 13:29:59 | sean-k-mooney | rmart04: are your vms using 1 gpu earch or multiple | |
| 13:30:16 | sean-k-mooney | it wont affect the answer just wondering | |
| 13:30:24 | rmart04 | Initial approach was 1VM 8 GPUs, second approach is 2VM's 4 each | |
| 13:31:00 | sean-k-mooney | cool if you can horizontally scale then yes 2 vm with a singel numa node each shoudl give better performance | |
| 13:31:25 | sean-k-mooney | since there will be no corss numa trafic fo the vm cpu memory and gpus | |
| 13:32:24 | rmart04 | Thats the plan, but previously I tried this and got a cannot allocate memory issue, which seemed to be related to no dma32 on node1 in /proc/zoneinfo. Which I believe may be due to the older kernel | |
| 13:33:27 | sean-k-mooney | rmart04: oh you hit that | |
| 13:33:56 | sean-k-mooney | so that is not really a kernel issue so much as a kernl/bios/firmware issue that we worked around with a kvm change | |
| 13:34:17 | sean-k-mooney | rmart04: really there should have been a dma32 region allcoated per numa node | |
| 13:34:34 | sean-k-mooney | the kvm fix was not to require numa affinity for the dma32 region | |
| 13:34:39 | rmart04 | Oh right, interesting. Could you point me at the info for the kvm change? | |
| 13:34:50 | rmart04 | ah Ok, is that strict=false or similar | |
| 13:35:08 | sean-k-mooney | kind of but that would have done it for all the vms memory | |
| 13:35:15 | sean-k-mooney | that was the alternitive workaround | |
| 13:35:34 | rmart04 | Please tell me its fixed in Stein? :D | |
| 13:36:43 | sean-k-mooney | https://lkml.org/lkml/2018/7/24/843 | |
| 13:36:51 | sean-k-mooney | this is not an openstack bug | |
| 13:36:58 | sean-k-mooney | so we did not modify nova | |
| 13:37:09 | sean-k-mooney | what distro are you using | |
| 13:37:32 | rmart04 | ah OK, Yes this is what I was looking at, I thought it was a Kernel patch | |
| 13:37:37 | rmart04 | Centos7 | |
| 13:37:45 | rmart04 | 3.10 kernel | |
| 13:37:55 | sean-k-mooney | it is for the kvm kernel module | |
| 13:38:08 | sean-k-mooney | there might have been another patch too | |
| 13:38:47 | sean-k-mooney | ok i know we backported this in rhel 7 | |
| 13:38:52 | sean-k-mooney | may in 7.6 | |
| 13:39:04 | sean-k-mooney | so hopefully you have that in the lates centos 7 too | |
| 13:39:16 | sean-k-mooney | let me see if i have the bz for it in my history | |
| 13:39:25 | rmart04 | ah that would be amazing | |
| 13:43:30 | sean-k-mooney | so this is the nova patch we decied not to go with https://review.opendev.org/#/c/684375/ partly because we could not test it | |
| 13:43:53 | sean-k-mooney | rmart04: the commit meassage has the links to the relevent bugs and converations | |
| 13:45:39 | openstack | bugzilla.redhat.com bug 1010885 in libvirt "kvm_init_vcpu failed: Cannot allocate memory in NUMA" [Medium,Closed: errata] - Assigned to mkletzan | |
| 13:45:39 | sean-k-mooney | hum it look like https://bugzilla.redhat.com/show_bug.cgi?id=1010885#c2 might also be a workaround but i dont think it is | |
| 13:52:41 | rmart04 | Remove cpuset from cgroup controllers? | |
| 13:53:11 | rmart04 | Is that what also makes the pinning work? | |
| 13:53:13 | sean-k-mooney | rmart04: yes but i dont know if that fully disables pinning | |
| 13:53:58 | sean-k-mooney | so the issue is that the wya libvirt appliees the cgrpus it also confines the allcoations of kernel memory | |
| 13:54:37 | sean-k-mooney | one of the fixes that was only a partial fix was to move that later so that the dma region could be allocate before the cpus are pinned | |
| 13:54:55 | sean-k-mooney | that was done in https://libvirt.org/git/?p=libvirt.git;a=commit;h=7e72ac7878 | |
| 13:55:24 | sean-k-mooney | but that was backin 2014 so it obviouslyu was not a full fix or it was broken angain later | |
| 13:56:01 | rmart04 | OK :/ | |
| 13:59:57 | sean-k-mooney | rmart04: this was the final kernel fix i belive https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=ee6268ba3a68 | |
| 14:00:01 | sean-k-mooney | that was in 4.19 | |
| 14:00:15 | rmart04 | ah ok, my bad I said 4.14 earliar | |
| 14:01:00 | rmart04 | How easy is it to find out whether it was backported in C7? | |