Earlier  
Posted Nick Remark
#openstack-nova - 2018-10-23
11:16:24 pvc_ okay i'll set it back
11:16:25 sean-k-mooney did you follow instruction that said you should disable it from nvidia?
11:16:33 pvc_ but my compute node already have the nvidia driver
11:16:52 pvc_ because earlier i cant install the nvidia-driver
11:17:00 pvc_ after disabled the vfio_pci, i successfully installed it
11:17:04 pvc_ i will get it back
11:18:15 sean-k-mooney so looking at https://images.nvidia.com/content/grid/pdf/GRID-vGPU-User-Guide.pdf in section 3.1 it does not say anything about disableing vfio so i would guess that is the issue
11:18:46 pvc_ i will get it back sean-k-mooney wait
11:19:16 pvc_ so install the driver then enable it. I enable it first after installing the driver
11:22:37 pvc_ sean-k-mooney can i add all the available vgpu on the nova.conf?
11:23:06 pvc_ then create a flavor that have 1 VGPU, 2 VGPU, 3 VGPU?
11:25:23 sean-k-mooney pvc_: you can only have 1 vgpu per vm currently and you can only enable 1 vgpu mdev type per phyical host
11:26:40 sean-k-mooney part of the limitation comes form libvirt and part form nvidia. multiple mdevs can be attached to the same instance but its a rather new addtion to the kernel and things have not really mautred yet in libvirt/qemu
11:26:43 pvc_ so i cannot use the 24gpus on my physical host?
11:27:24 jangutter sean-k-mooney: minor correction, intel_iommu=on is required, and "iommu=pt" is discouraged.
11:28:16 jangutter sean-k-mooney: iommu=pt used to be required when DPDK hadn't set vfio-pci as it's default transport yet, and people used uio.
11:29:56 sean-k-mooney jangutter: yes but i still recommend iommu=pt as the iommu is picky somethimes and i find you hit less corner cases if you limit its scope to pasthrough devices
11:30:28 pvc_ ERROR nova.compute.manager [instance: 014fcc5f-660b-40df-92ed-9f4587993fa7] Verify all devices in group 58 are bound to vfio-<bus> or pci-stub and not already in use
11:31:00 sean-k-mooney pvc_: yes so as i seaid previously i think your error is the gpus is sharingin an iommu group with anothe rdevice
11:31:18 sean-k-mooney pvc_: all devices in a iommu group must use the same kernel driver
11:31:22 jangutter sean-k-mooney: heh, iommu=pt is one of the worst-named options. The "passthrough" mapping it enables is a way to 'bypass' the IOMMU by creating a 1:1 memory map for the PCI space.
11:31:50 pvc_ how can i check that sean-k-mooney?
11:31:59 sean-k-mooney jangutter: yes and that 1:1 mapping fixes so manny things :P
11:32:01 jangutter pvc_: yeah, you have to hand off _all_ the devices in an iommu group, and some platforms can't divide between them.
11:32:21 sean-k-mooney pvc_: am you can find this in sysfs
11:32:28 pvc_ sysfs?
11:32:29 sean-k-mooney let me see if i can remember
11:32:33 pvc_ thank you
11:32:53 pvc_ jangutter i cannot use the 12vgpus of my 1 tesla?
11:32:58 pvc_ only 1 vgpus?
11:33:43 jangutter pvc_: it depends, the chipset may allow you to pass the entire card at once, but not a portion of it.
11:34:01 sean-k-mooney pvc_: you can but if your tesla share an iommu group with a nic then they both need to be bound to vfio-pci or pci-stub
11:34:18 jangutter pvc_: the kernel documentation (low level warning) is at: https://www.kernel.org/doc/Documentation/vfio.txt
11:34:38 pvc_ thankyou jangutter, how can i do that sean-k-mooney?
11:34:53 pvc_ sorry this is my first time using vgpu, im using just the pci-passthroigh
11:35:05 sean-k-mooney first we need to see what is in iommu group 58 in /sys/class/iommu/
11:35:17 jangutter pvc_: you can do something like: ls -l /sys/bus/pci/devices/0000:06:0d.0/iommu_group/devices
11:35:29 jangutter pvc_: where the pci address is obviously yours.
11:35:45 jangutter pvc_: that's a list of PCI devices in one group
11:36:19 jangutter pvc: they _all_ have to be passed through together, or it will fail.
11:36:37 pvc_ lrwxrwxrwx. 1 root root 0 Oct 23 11:36 0000:06:00.0 -> ../../../../devices/pci0000:00/0000:00:02.0/0000:06:00.0
11:37:27 jangutter Also check /sys/kernel/iommu_groups/58 ?
11:37:38 openstackgerrit Elod Illes proposed openstack/nova master: Transform scheduler.select_destinations notification https://review.openstack.org/508506
11:38:01 pvc_ 979d010e-17a3-4ac9-987a-565f9ba4b4a6
11:38:07 pvc_ [root@overcloud-novacompute-0 iommu]# ls /sys/kernel/iommu_groups/58/devices/ 979d010e-17a3-4ac9-987a-565f9ba4b4a6
11:39:07 jangutter pvc_: interesting... I haven't seen a UUID there yet. can you ls -l it?
11:39:27 sean-k-mooney jangutter: my guess is the uuid is a mdev uuid
11:39:43 pvc_ http://paste.openstack.org/show/732803/
11:41:30 pvc_ mdev bus types plus driver http://paste.openstack.org/show/732805/
11:42:58 claudiub heyo. since it's spec review day, could you take a look again at the live-resize one? https://review.openstack.org/#/c/141219/
11:43:05 jangutter pvc_, sean-k-mooney: should the PCI passthrough libxml element set "managed=true"?
11:43:30 jangutter sean-k-mooney: managed=true generally means that it will auto-bind vfio-pci to the device before attempting passthrough?
11:43:42 bauzas jangutter, sean-k-mooney: context ?
11:43:45 sean-k-mooney jangutter: sorry where was the libvirt xml i missed that
11:43:58 sean-k-mooney bauzas: trying to help pvc_ with there vgpu issue
11:44:15 bauzas and?
11:44:16 sean-k-mooney bauzas: pvc_ is seeing a weird iommu error
11:44:22 jangutter hang on, looking up the libxml doc.
11:44:45 bauzas we don't manage the iommu group
11:45:05 jangutter (rofl) s/libxml/libvirt/
11:45:32 sean-k-mooney bauzas: yes that is managed by the uefi and kernel
11:46:12 bauzas oh wait
11:46:13 jangutter sean-k-mooney: when doing https://libvirt.org/formatdomain.html#elementsHostDevSubsys <--- there's a 'managed=yes' element in the xml. if that's set it will do the vfio-pci binding for you.
11:46:24 bauzas is pvc_ doing PXI
11:46:26 sean-k-mooney bauzas: pvc_ is getting Verify all devices in group 58 are bound to vfio-<bus> or pci-stub and not already in use
11:46:36 bauzas pci passthrough?
11:47:02 pvc_ im using vgpu bauzas
11:47:03 sean-k-mooney bauzas: no pvc_ is trying to do vgpu passthrouhg not pci
11:47:29 sean-k-mooney bauzas: this is there flavor http://paste.openstack.org/show/732799/
11:47:58 bauzas sec, GPU passthrough?
11:47:59 sean-k-mooney and pvc_ has set the gpu type in the config to nvidia-160
11:48:08 bauzas I'm confused
11:48:25 jangutter pvc_: what does "readlink /sys/bus/pci/devices/0000:06:00.0" say?
11:48:32 bauzas virtual GPU or GPU passthrough?
11:48:45 sean-k-mooney bauzas: pvc_ has a tesla gp100 and is trying to ues the mdev based virtual gpu
11:49:36 bauzas then don't do vfio bus
11:49:44 bauzas or pci stub
11:50:06 bauzas just use the nvidia kernel module
11:50:10 jangutter aaah, the penny drops.
11:50:11 sean-k-mooney pvc_: can you provide the libvirt xml that nova generated so we can see what its doing
11:50:37 pvc_ i add an option of options vfio-pci ids=10de:15f8 should i disable this?
11:50:44 pvc_ then new error is occured
11:50:49 bauzas sean-k-mooney : I think pvc_ is mixing two different things
11:50:55 pvc_ 2018-10-23 11:49:46.754 7 WARNING nova.virt.libvirt.driver [req-a3570604-eed3-4f8f-a244-202ac2b92b7d - - - - -] Error from libvirt while getting description of instance-00000001: [Error Code 42] Domain not found: no domain with matching uuid '2ac6c395-5f92-4e9e-a52a-cf90b9d551c5' (instance-00000001): libvirtError: Domain not found: no domain with matching uuid '2ac6c395-5f92-4e9e-a52a-cf90b9d551c5' (instance-00000001)
11:50:57 openstackgerrit Jim Rollenhagen proposed openstack/nova-specs master: Use conductor groups to partition nova-compute services for Ironic https://review.openstack.org/609709
11:51:34 sean-k-mooney pvc_: you should not be disableing or foceing vfio-pci
11:52:11 sean-k-mooney you should allow it to be loaded if need but oterwise do not set any options for the vfio kernel module at all
11:53:32 pvc_ http://paste.openstack.org/show/732806/
11:54:06 bauzas yup this
11:54:08 sean-k-mooney pvc_: the phyical gpu needs to be bound to the nvidia grid driver and the mdevs will be created via the vfio framwork in the kenel but you should not force the pgpu to use vfio-pci
11:54:22 pvc_ i need to remove that?
11:54:24 bauzas yup it's only for GPU passthrough
11:54:27 pvc_ okay wait
11:54:30 bauzas hence my confusion
11:54:30 pvc_ i'll remove
11:54:56 sean-k-mooney bauzas: im guessing the driver in use should be nvidia_vgpu_vfio or nvida correct
11:55:10 sean-k-mooney praobly nvidia_vgpu_vfio
11:55:15 pvc_ i reboot again the hypervisor wait
11:56:52 bauzas sean-k-mooney, I don't remember the module name but yeah something like that
11:57:30 pvc_ bauzas is it possible to use the 12vgpus of my gpu?

Earlier   Later