Earlier  
Posted Nick Remark
#openstack-nova - 2018-10-23
12:04:54 pvc_ i add this on module
12:05:02 pvc_ options nvidia_vgpu_vfio ids=10de:15f8
12:05:39 sean-k-mooney can yo bind it by hand instad of via the moduel file to test it
12:05:57 pvc_ how can i bind it? i'm sorry im not done it before
12:05:59 bauzas that's super weird
12:06:33 pvc_ http://paste.openstack.org/show/732809/
12:06:33 pvc_ this is my conf
12:07:16 bauzas pvc_ did you remove the nouveau driver ?
12:07:32 sean-k-mooney echo 06:00.0 | sudo tee /sys/bus/pci/drivers/nvidia/unbind
12:07:50 sean-k-mooney echo 06:00.0 | sudo tee /sys/bus/pci/drivers/nvidia_vgpu_vfio/bind
12:07:53 pvc_ [root@overcloud-novacompute-0 nova]# lsmod | grep nou [root@overcloud-novacompute-0 nova]#
12:07:56 pvc_ yes bauzas
12:08:48 pvc_ tee: /sys/bus/pci/drivers/nvidia/unbind: No such device
12:09:00 pvc_ I install this driver
12:09:09 pvc_ NVIDIA-Linux-x86_64-390.72-vgpu-kvm.run
12:09:15 bauzas I need to drop, planned outage here
12:10:01 sean-k-mooney pvc_: from http://paste.openstack.org/show/732807/ that should have been the driver in use
12:10:10 pvc_ i use this 0000:06:00.0
12:10:31 pvc_ tee: /sys/bus/pci/drivers/nvidia_vgpu_vfio/bind: No such file or directory
12:10:43 pvc_ it is already unbind
12:11:04 pvc_ no nvidia_vgpu_vfio on drivers
12:11:08 pvc_ just nvidia
12:14:16 sean-k-mooney pvc_: this is the latest verion of the nvdia vgpu user guide https://docs.nvidia.com/grid/5.0/pdf/grid-vgpu-user-guide.pdf i think you need to back through it and section 4.2 specifically
12:14:33 pvc_ sean-k-mooney im using a ubuntu image with img_hide_hypervisor_id='true'
12:15:13 sean-k-mooney pvc_: i asked thi earliar but is the compute node a phyical server or a vm
12:15:34 pvc_ compute node is a physical server
12:15:42 pvc_ on docs it said the grid driver
12:15:48 pvc_ but this is the driver i installed
12:16:01 pvc_ installed NVIDIA-Linux-x86_64-390.72-vgpu-kvm.run
12:16:10 pvc_ i have this grid driver NVIDIA-Linux-x86_64-390.75-grid.run
12:17:03 sean-k-mooney pvc_: oh ok i think i understand the issue then
12:17:12 sean-k-mooney you install the guest driver on the host
12:17:25 pvc_ to enable this
12:17:25 pvc_ yes sean-k-mooney
12:17:31 pvc_ /sys/class/mdev_bus/*/mdev_supported_types
12:18:04 pvc_ i install the driver on my compute node ( baremetal server ) im using a tripleo-deployment
12:22:26 pvc_ sean-k-mooney what will i do then?
12:22:26 aperevalov hello, do nova or neutron has functional test for direct (SR-IOV) port (something like tempest test)?
12:23:37 sean-k-mooney aperevalov: i dont belive so in upstream tempest. neutron may have fullstack test but ingerneral our sriov testing is limited
12:24:58 sean-k-mooney pvc_: this seams to be a driver issue not a nova one. there is little more advice i can give other then they to follow the nvida docs form start to finish exactly
12:25:01 aperevalov sean-k-mooney: I assume, if such functional test exists it uses real HW with SR-IOV, but not simulators or emulators.
12:25:24 pvc_ so you think i installed the grid one?
12:25:31 pvc_ my baremetal is centos 7
12:25:34 sean-k-mooney aperevalov: so the fact we can use simulator/emulators for sriov is the main reason we have such limited testing
12:25:58 sean-k-mooney pvc_: yes you need to install the grid one
12:26:13 sean-k-mooney at least i think so.
12:26:26 pvc_ okay wait
12:27:09 sean-k-mooney aperevalov: i recently looked at using the netdevsim kernel module for sriov testing but it only emulates the kernel netdevs it does not emulate the pci devices or virtual function so we cant use it for testing
12:28:40 aperevalov sean-k-mooney: I checked netdevsim on the latest kernel too, there are no device to put it into docker container. Docker container it's because I'm doing such research for kuryr-kubernetes.
12:29:49 aperevalov sean-k-mooney: yes netdevsim is based on its own bus, but not pci. I also found attemp to submit to QEMU's pci emulation the SR-IOV support.
12:30:36 sean-k-mooney aperevalov: yes so there is a netdev but no pci device on the virtual pci device. so you can use ip link and allocate vfs via sysfs but they dont show up on the virtual pci bus
12:31:20 sean-k-mooney aperevalov: yes i have see that in the past but that wont help us with testing as we would need the hosting vms provided by the could providers to emulate pci device that support sriov
12:31:30 aperevalov sean-k-mooney: but it was postponed due to lack of existing working qemu drivers, initial author suggested copied e1000, but it wasn't in working condition.
12:32:21 sean-k-mooney aperevalov: so first qemu would have to be extended then libvirt and then nova. once that is done we would need cloud providers to proved devices with virt sriov capable nics then we could start using it in the upstream gate
12:32:41 sean-k-mooney *provided vms with ...
12:33:29 sean-k-mooney aperevalov: effectivly we rely on thirdpart cis to test sriov. either intels or melonox's ci
12:33:54 sean-k-mooney intels ci was intented to have signifcatily more sriov testing then it currently has but that never happened
12:34:31 aperevalov sean-k-mooney, I see it's a long way, and seems netdevsim (just kernel module) looks like easiest (if kernel community will be agreed to bind it with pci bus).
12:35:50 sean-k-mooney aperevalov: most people dont know that sr-iov is a specificaiton from the PCI-SIG i dont think the current netdevsim moudle is technically a confroming sriov implentaiton without the pci emulation
12:36:19 pvc_ sean-k-mooney ls: cannot access /sys/class/mdev_bus/*/mdev_supported_types: No such file or directory
12:36:28 sean-k-mooney aperevalov: that said its a netdev simulator not a sriov simulator so they skipped the bits they did not need for there own testing
12:36:32 pvc_ i reboot again to reflect the new driver
12:38:35 aperevalov sean-k-mooney: ok, I'll talk with authors of netdevsim. Thank you for information. BTW is intel or mellanox ci is publicly available. Or that work is going through their teems involved into openstack community?
12:40:40 aperevalov sean-k-mooney: I think moshele knows it.
12:41:18 pvc_ hi sean-k-mooney
12:41:20 moshele aperevalov: I know what?
12:41:30 pvc_ after installign the grid driver the mdev is gone
12:41:33 pvc_ ls: cannot access /sys/class/mdev_bus/*/mdev_supported_types: No such file or directory
12:41:33 sean-k-mooney so i used to work at intel and one of my roles there was product owner of the intel nfv ci. it used to fall to the upstream teams to identify which features needed ci testing and either add it or request that the team maintaining it add it to there backlog
12:41:59 pvc_ sean-k-mooney any ideas? ls: cannot access /sys/class/mdev_bus/*/mdev_supported_types: No such file or directory
12:42:13 sean-k-mooney im not sure how the intel nfv ci currently works but if you reach out to the new maintianer im sure they will respond
12:42:13 pvc_ after installing the grid driver only not the vgpu
12:45:30 moshele aperevalov: The Mellanox CI is public but not part of the openstack community. and I think intel CI is the same, because it depend on nic vendor
12:45:33 sean-k-mooney pvc_: sorry not really. as i said this seams to be an nvidia driver issue. i dod not have acess to the hardware to test it my self so beyond reading there docs there is not much more advice i can give
12:46:06 pvc_ but sean-k-mooney when i install the kvm-vgpu i can list the mdev
12:47:03 aperevalov moshele: does it use in gerrit integration tests by zuul?
12:48:32 moshele aperevalov: we run tempest scenarios and configure the tempest to use vnic_type=direct
12:49:35 moshele aperevalov: see https://github.com/openstack/tempest/blob/master/tempest/config.py#L628
12:50:10 sean-k-mooney moshele: oh when was that option added?
12:50:15 openstackgerrit Stephen Finucane proposed openstack/nova stable/rocky: fixtures: Track volume attachments within CinderFixtureNewAttachFlow https://review.openstack.org/612485
12:50:16 openstackgerrit Stephen Finucane proposed openstack/nova stable/rocky: conductor: Recreate volume attachments during a reschedule https://review.openstack.org/612487
12:50:16 openstackgerrit Stephen Finucane proposed openstack/nova stable/rocky: Add regression test for bug#1784353 https://review.openstack.org/612486
12:50:16 sean-k-mooney that is useful to know about
12:50:25 moshele aperevalov: long time ago
12:50:39 moshele sean-k-mooney: long time ago
12:51:02 pvc_ vfio_iommu_type1.allow_unsafe_interrupts=1
12:51:07 pvc_ sean-k-mooney vfio_iommu_type1.allow_unsafe_interrupts=1?
12:51:19 sean-k-mooney moshele: cool aperevalov the intel ci does somthing similar
12:52:07 sean-k-mooney aperevalov: the intel ci uses the standard senario test but addes extraflaovr extraspecs for cpu pinning hugepages numa toplogy exctra
12:52:44 aperevalov sean-k-mooney: If I trully understood, kuryr-kubernetes (tempest test) can also be running there?
12:52:47 sean-k-mooney pvc_: are you getting a message in dmesg? vfio_iommu_type1.allow_unsafe_interrupts=1 is specificaly for working around old buggy hardware
12:53:28 sean-k-mooney aperevalov: the intel nfv ci does not load the kuryr-kuberntese tempest module or deploy tempetst
12:53:40 sean-k-mooney at least it didnt when i was invovled with it
12:53:52 pvc_ Oct 23 12:53:40 overcloud-novacompute-0 journal: 2018-10-23 12:53:40.708+0000: 3128: warning : virDomainAuditHostdev:424 : Unexpected hostdev type while en
12:53:56 pvc_ Oct 23 12:53:40 overcloud-novacompute-0 journal: libvirt: QEMU Driver error : Requested operation is not valid: domain is not running
12:53:58 sean-k-mooney * or deply kuryr-kubernetes
12:54:59 sean-k-mooney pvc_: i dont really have time to contiue debugging sorry. i need to update some review and catch up on spec review today
12:56:41 pvc_ 2018-10-23 12:41:10.331+0000: 3216: error : virPCIDeviceNew:1787 : Device 0003:01:05.1 not found: could not access /sys/bus/pci/devices/0003:01:05.1/config
12:57:57 pvc_ virPidFileAcquirePath:422 : Failed to acquire pid file '/var/run/libvirtd.pid': Resource temporarily unavailable
13:02:22 pvc_ sean-k-mooney there is an issue on my nova_libvirt

Earlier   Later