Earlier  
Posted Nick Remark
#openstack-nova - 2020-06-18
08:39:31 openstackgerrit Wenping Song proposed openstack/nova master: delete sub resource provider when delete resource provider https://review.opendev.org/719163
09:54:40 openstackgerrit Wenping Song proposed openstack/nova master: delete sub resource provider when delete resource provider https://review.opendev.org/719163
10:38:26 arne_wiebalck TheJulia: dansmith sean-k-mooney Sorry, I missed the discussion yesterday.
10:38:26 arne_wiebalck Let me add a little background info: physical instances at CERN are all done via Nova and Ironic. The main user is OpenStack itself and we rely on Ironic's software RAID and standard cloud images
10:38:26 arne_wiebalck to deploy. However, other users (like the Ceph team) want to partition/RAID their physical instances in a specific way as needed for their service. There is no convenient way to express this is via Nova/Ironic at the moment. So, they reinstall these instances after initial deployment through Nova/Ironic once more with kickstart. This double installation is what we'd like to get rid of, and this triggered
10:38:26 arne_wiebalck the whole discussion.
10:38:26 arne_wiebalck One of the main points for user adoption probably is that whenever sth needs to change, this should be feasible without having the Ironic admin to re-clean nodes or the Nova admin to add new flavors (with 5000 nodes in Ironic we have around 150 resource classes and hence 150 flavors). This is why we came up with the idea of a "kickstart" driver where Ironic uses a kickstart/preseed file to do the
10:38:27 arne_wiebalck deployment (and skip the usual image deployment) as it would provide the user with the same flexibility as there is now. But since we would like to keep Nova in the mix (for various reasons), we may need some tooling help on the Nova side.
12:36:20 aarents Hi nova,
12:38:02 aarents sean-k-mooney: dansmith Let me know if it needs further amend on https://review.opendev.org/#/c/736169/ https://review.opendev.org/#/c/734776/ thanks!
12:39:07 openstackgerrit Alexandre Arents proposed openstack/nova master: libvirt: ensure disk_over_commit is not negative https://review.opendev.org/719008
12:39:29 sean-k-mooney ^ that should be done by the config option
12:39:33 sean-k-mooney we set min 0 i think
12:40:41 sean-k-mooney oh that is not the ratio
12:41:09 sean-k-mooney for that to be negitive the size on disk would have to be larger then the virtual size
12:41:39 aarents sean-k-mooney: yes
12:41:57 sean-k-mooney does it being negitive break something
12:43:06 sean-k-mooney ah i see
12:43:18 aarents It just mislead calcuation of available_disk_least on which rely disk_filter
12:47:43 sean-k-mooney ya clamping the value shoudl be fine.
13:12:54 rmart04 Hey all, I'm wondering if anyone could help me. Not strictly dev specific, but I'm having trouble with NUMA information being passed through to my virtual machines. /sys/bus/pci/devices*/numa_node always = -1. I'm running Rocky on C7, with numa_nodes=2 and cpu_sockets=2.
13:13:08 rmart04 and cpu pinning policy set to dedicated
13:14:17 sean-k-mooney is the -1 on the host or in the guest
13:14:23 sean-k-mooney if its in the guest that is expected
13:14:26 rmart04 In the guest
13:14:52 sean-k-mooney so we do not currently create a pcie root complex per numa node
13:15:13 sean-k-mooney since there is only one pci root all devices are childerne of that root
13:15:42 sean-k-mooney wehn we do passthough we dont affiniteis the pci device to the virutal numa node of the guest
13:15:59 sean-k-mooney so it is reported as -1 meaning no numa affinity in the guest
13:16:55 sean-k-mooney to change that we would likely have to use the q35 machine type and create a pcie root complete per numa node then add the passthough deivice to the correct pci root
13:17:45 sean-k-mooney that has other implciation mainliy that on move operation either we have to allow the toploty to change or we have to limit the host we can select to maintain the current toplogy
13:18:14 sean-k-mooney if we allow the toplogy to change the virtual pci address of the devices in the guest would also change
13:19:16 sean-k-mooney rmart04: but yes if you have a multi numa node guest this can result in cross numa traffic worse case twice because you dont know the numa affinity of the device
13:20:18 openstackgerrit Merged openstack/nova master: libvirt: Mark e1000e VIF as supported https://review.opendev.org/734777
13:22:32 rmart04 OK, appreciate all the info SeanKMooney. I guess another way around this is to split the host into two guests, one on each numa node with their associated pci-passthrough devices (GPUs). Currently I appear to be blocked on this by my older kernel. 3.10. I bump into an issue allocating memory from the second NUMA node for the second machine. I believe this is fixed in 4.14.
13:23:06 rmart04 qq, you mention the q35 machine type, what type do we use by default?
13:25:51 sean-k-mooney rmart04: if the guest has a numa toploty we do not allow its memory to come form a remote numa node by design
13:26:01 sean-k-mooney rmart04: we use pc
13:26:23 sean-k-mooney or pc-i440fx
13:26:27 sean-k-mooney something like that
13:26:43 sean-k-mooney rmart04: what version of openstack are you using
13:27:02 rmart04 Rocky (Stein upgrade this weekend)
13:27:26 sean-k-mooney do you have gpus on all host numa nodes
13:27:30 sean-k-mooney or just numa 0
13:27:42 rmart04 Yep, 2 sockets, 8 GPUs
13:27:50 rmart04 4 each
13:28:16 sean-k-mooney ok what iw was going to say is you might need to use nuam_policy=preferred in the alias
13:28:21 sean-k-mooney if you did not have them split
13:28:54 sean-k-mooney if you do then yes 2 vms with 1 numa each and the default legacy polciy which enforce numa affintiy between cpu/memory and the pci device is what you will want
13:29:42 sean-k-mooney you can create a dual numa guest but the limitation is you will know know what numa node in the guest maps to the actull location of the device on the host
13:29:59 sean-k-mooney rmart04: are your vms using 1 gpu earch or multiple
13:30:16 sean-k-mooney it wont affect the answer just wondering
13:30:24 rmart04 Initial approach was 1VM 8 GPUs, second approach is 2VM's 4 each
13:31:00 sean-k-mooney cool if you can horizontally scale then yes 2 vm with a singel numa node each shoudl give better performance
13:31:25 sean-k-mooney since there will be no corss numa trafic fo the vm cpu memory and gpus
13:32:24 rmart04 Thats the plan, but previously I tried this and got a cannot allocate memory issue, which seemed to be related to no dma32 on node1 in /proc/zoneinfo. Which I believe may be due to the older kernel
13:33:27 sean-k-mooney rmart04: oh you hit that
13:33:56 sean-k-mooney so that is not really a kernel issue so much as a kernl/bios/firmware issue that we worked around with a kvm change
13:34:17 sean-k-mooney rmart04: really there should have been a dma32 region allcoated per numa node
13:34:34 sean-k-mooney the kvm fix was not to require numa affinity for the dma32 region
13:34:39 rmart04 Oh right, interesting. Could you point me at the info for the kvm change?
13:34:50 rmart04 ah Ok, is that strict=false or similar
13:35:08 sean-k-mooney kind of but that would have done it for all the vms memory
13:35:15 sean-k-mooney that was the alternitive workaround
13:35:34 rmart04 Please tell me its fixed in Stein? :D
13:36:43 sean-k-mooney https://lkml.org/lkml/2018/7/24/843
13:36:51 sean-k-mooney this is not an openstack bug
13:36:58 sean-k-mooney so we did not modify nova
13:37:09 sean-k-mooney what distro are you using
13:37:32 rmart04 ah OK, Yes this is what I was looking at, I thought it was a Kernel patch
13:37:37 rmart04 Centos7
13:37:45 rmart04 3.10 kernel
13:37:55 sean-k-mooney it is for the kvm kernel module
13:38:08 sean-k-mooney there might have been another patch too
13:38:47 sean-k-mooney ok i know we backported this in rhel 7
13:38:52 sean-k-mooney may in 7.6
13:39:04 sean-k-mooney so hopefully you have that in the lates centos 7 too
13:39:16 sean-k-mooney let me see if i have the bz for it in my history
13:39:25 rmart04 ah that would be amazing
13:43:30 sean-k-mooney so this is the nova patch we decied not to go with https://review.opendev.org/#/c/684375/ partly because we could not test it
13:43:53 sean-k-mooney rmart04: the commit meassage has the links to the relevent bugs and converations
13:45:39 sean-k-mooney hum it look like https://bugzilla.redhat.com/show_bug.cgi?id=1010885#c2 might also be a workaround but i dont think it is
13:45:39 openstack bugzilla.redhat.com bug 1010885 in libvirt "kvm_init_vcpu failed: Cannot allocate memory in NUMA" [Medium,Closed: errata] - Assigned to mkletzan
13:52:41 rmart04 Remove cpuset from cgroup controllers?
13:53:11 rmart04 Is that what also makes the pinning work?
13:53:13 sean-k-mooney rmart04: yes but i dont know if that fully disables pinning
13:53:58 sean-k-mooney so the issue is that the wya libvirt appliees the cgrpus it also confines the allcoations of kernel memory
13:54:37 sean-k-mooney one of the fixes that was only a partial fix was to move that later so that the dma region could be allocate before the cpus are pinned
13:54:55 sean-k-mooney that was done in https://libvirt.org/git/?p=libvirt.git;a=commit;h=7e72ac7878
13:55:24 sean-k-mooney but that was backin 2014 so it obviouslyu was not a full fix or it was broken angain later
13:56:01 rmart04 OK :/
13:59:57 sean-k-mooney rmart04: this was the final kernel fix i belive https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=ee6268ba3a68
14:00:01 sean-k-mooney that was in 4.19
14:00:15 rmart04 ah ok, my bad I said 4.14 earliar
14:01:00 rmart04 How easy is it to find out whether it was backported in C7?
14:03:43 sean-k-mooney its in the rhel kernel-3.10.0-957.26.1.el7 pacakge and needs libivrt ibvirt-4.5.0-13.el7 or higher
14:06:26 rmart04 OK thats amazing thank you. I think i'll see how far we are away from those package versions and expidite moving to them
14:46:06 mordred sean-k-mooney: you remember every conversation we've had about things from the past, right?
14:46:29 sean-k-mooney mordred: ocourse i was right and you were ....
14:46:42 sean-k-mooney mordred: is there a converstation in partcalar?

Earlier   Later