Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-21
10:44:15 sean-k-mooney i dont think it actully works we cannot enable it for nics that use sriov
10:44:26 sean-k-mooney i.e. nics that present the vf directly to the guest
10:44:37 sean-k-mooney but it might eb possibel for macvtap or vdpa
10:45:03 sean-k-mooney dvo-plv_: in terms of pings the active member of the nova core team
10:46:13 sean-k-mooney dvo-plv_: so gibi, bauzas, melwitt, gmann and dansmith are you best bet in addtion to me. stephenfin is also around somethimes but they mainly work on non nova related thigns day to day
10:47:51 dvo-plv_ okay, thanks
10:49:32 sean-k-mooney so looping back to multi queue
10:50:06 sean-k-mooney implementing this in the way you wanted is not going to be easy or really desireable form a nova point of view
10:51:19 sean-k-mooney there are a few parts to this problem
10:51:56 sean-k-mooney first we need to detach and recored the number of queues avaialbel in the vf
10:52:28 sean-k-mooney second we need to be able to schudle based on that (either updatign the pci filter or recording this in placement)
10:53:16 sean-k-mooney third we need a way to request a device with a min number of queuse multiup queuse out side the falvoar/image (likely on the nutron port)
10:54:09 sean-k-mooney finally we need to take the resouce request form teh port and include that in our scueduling reeust and wonce we find a host/device that meets that need we need to ensure that the qemu device is cofnigured correctly
10:54:37 dvo-plv_ regarding first question, we investigated it and the best option on our opinio is to parse other config. we configure queues like that -a 0000:65:00.0,representor=[4-6],portqueues=[4:2,5:2,6:2]. Yes it will work only for our nic
10:54:57 sean-k-mooney which config
10:55:01 sean-k-mooney the pci config space
10:55:15 dvo-plv_ ovs. other_config
10:55:15 sean-k-mooney you can useuslly get this form sysfs i tought
10:55:31 sean-k-mooney we cant do that
10:55:42 sean-k-mooney the ovs port wont exist at that point
10:55:45 dvo-plv_ no we can not, this is untrivial task for our dpdk driver
10:56:05 sean-k-mooney oh right so honestly you cant start this work
10:56:24 sean-k-mooney until the basic work of supporting the dpdk represtors is done
10:56:28 dvo-plv_ ovs port not, but vf yes. We would like to parse this config and fill it to the device_spec
10:56:54 sean-k-mooney the ovs port will be created and added by os-vif
10:57:11 sean-k-mooney only after we ahve selected a vf
10:57:46 sean-k-mooney so i think we need https://review.opendev.org/c/openstack/nova-specs/+/859290 to be done before we can talk about multiqueue
10:59:17 dvo-plv_ I see, we thought we can start to find solution for all comments to the blueprint in the parallel at the moment
10:59:47 sean-k-mooney well we could but its going to be diffuctly to compelte jsut one of the 3-4 specs you have propsoed this cycle
10:59:51 sean-k-mooney maybe 2
11:00:00 sean-k-mooney its very unlikely that all of them will land
11:00:41 dvo-plv_ You mentined that multiqueue functional will be possible for vdpa. So maybe it will be better to move from virtio-forwarder to the vdpa nvic type for future purposes
11:01:17 sean-k-mooney well there is work in dpdk to supprot vdpa
11:01:31 sean-k-mooney i was expecting to have a vdpa-user type at some point for that
11:02:47 sean-k-mooney https://doc.dpdk.org/guides/vdpadevs/features_overview.html
11:02:52 sean-k-mooney i have not looked into it much
11:03:25 sean-k-mooney i dont think that is supported by ovs-dpdk currenlty but i have not really been following it closely
11:04:08 sean-k-mooney dvo-plv_: so for the basic enablement we are going to be trackign napatec VF which we will add to ovs as dpdk prots corret
11:04:34 sean-k-mooney and then those will be exposed to the guest as vhost-user ports
11:05:01 sean-k-mooney so for multi queue we would need to read the number of queues on the vf ideally
11:05:07 dvo-plv_ This multiqueue functional with queue mq and vector is in the our ovs fork at the moment
11:05:08 sean-k-mooney because we need that info for schduling
11:05:27 sean-k-mooney ok so thats kind of a problem
11:05:52 sean-k-mooney we do not really allow enablment of forked functionality in nova
11:06:11 dvo-plv_ Yes, I remember that it requires for placement to handle scheduler with queues number
11:06:57 sean-k-mooney if we can do it generically we enabel it so i was hoppng we coudl do somehting liek read /sys/bus/pci/device/<address>/num_queus or somehting like that
11:07:09 sean-k-mooney ideally vai libvirt nodedev api
11:07:15 sean-k-mooney not reading sys directly
11:09:18 sean-k-mooney so you can get the queue like this https://paste.opendev.org/show/bVSM5IDtJTRwhcIcuTFs/
11:09:27 sean-k-mooney that is a pf
11:09:33 sean-k-mooney but i belive the same is true for VFs
11:11:44 sean-k-mooney i done see the queues in libvirt https://paste.opendev.org/show/bJHNfZR8JNpxJVvLggCI/
11:12:18 sean-k-mooney so the first step woudl really be to add the ablityu to get the queus form libvirt to libvirt
11:14:39 sean-k-mooney dvo-plv_: do you need the VFs to be bound to vfio-pci
11:15:05 sean-k-mooney i assume use so i geuss this infor will not be aviable via the vf since it wont have a netdev
11:16:14 dvo-plv_ we probe vfio-pci driver modprobe vfio-pci enable_sriov=1
11:16:25 dvo-plv_ ane then allocate vf echo "$NUMVFS" > /sys/bus/pci/devices/0000:$BUS:00.0/sriov_numvfs
11:17:14 sean-k-mooney yep thats pretty standard for dpdk although the enable_sriov bit is relitvly recent
11:17:18 dvo-plv_ we don not have netdev devices for that. so this is hard to get queue at linux layer
11:17:26 sean-k-mooney ya
11:17:46 sean-k-mooney so the probelm is the device spec is not intended for configuration
11:17:52 sean-k-mooney it was orgianly just for filtering
11:18:02 sean-k-mooney we have since added some metadta to it
11:18:14 sean-k-mooney im not sure hwo peopel would fell about adding the number of queues
11:18:35 dvo-plv_ yes, so this is why we firstly decided that user can fill device_spec with additional rapameter for filtering
11:18:58 dvo-plv_ queue_number
11:19:10 sean-k-mooney dvo-plv_: right so that approch has been rejected in the past
11:19:17 dvo-plv_ yes
11:19:26 sean-k-mooney i mean before you propsoed it
11:19:55 sean-k-mooney there have been attempts to do this in teh past and it was rejected
11:20:07 sean-k-mooney that said we now have enough things like this that we might be ok with it
11:20:26 sean-k-mooney we now have things liek remote_managed and resouce class
11:20:51 dvo-plv_ I see. now you would like to see that queue parameter gets automatically like metadata and placemnet filter nodes according to the required queue. But you would not like to increase resource provider
11:20:58 sean-k-mooney so addign queue_pairs=<count> might be ok
11:21:58 sean-k-mooney am no we can model this in placment but it would need use to have one RP per vf
11:22:05 sean-k-mooney which is not soemthgn we wanted to do if we coudl avoid it
11:22:37 sean-k-mooney we would need to adress the placment scaling bug first
11:23:24 sean-k-mooney dvo-plv_: https://review.opendev.org/c/openstack/nova/+/855885
11:23:33 sean-k-mooney so we could not track this in placment initally
11:23:51 sean-k-mooney we would have to track this in nova and use the pci filter to filter based on the queus
11:24:20 sean-k-mooney eventully it could be done in placement but we also need to start trackign neutron consumabel pci devices in placemnet before that
11:25:20 sean-k-mooney the only workable solution i see in the next 6-12 months is to do this in nova
11:26:00 sean-k-mooney if we require all VFs in the same pool to have the same queue count
11:26:31 sean-k-mooney then we can add the queue_pair count to the extra_info on the pci_device in the nova db
11:26:46 sean-k-mooney and the pci_passhtough filter can use that
11:28:10 dvo-plv_ lets assume that we have already dealt with the automatic queue getting. Lets use traits config. if this node has vf with 2 and 3 queues. Placemnet will add traits 2_queus and 3_queues to the scheduler to fitler node by queue. and than, when node is choosen, nova and choose appropariate vf by queue number
11:28:28 sean-k-mooney no
11:28:35 sean-k-mooney thsi is not a correct use of traits
11:29:38 sean-k-mooney traits cannot be used for Quantitative aspect of a resuce i.e the number of queuse or frequency of a cpu
11:30:36 sean-k-mooney HW_NIC_MULTIQUEUE is an accpaable trait which we already have https://github.com/openstack/os-traits/blob/master/os_traits/hw/nic/__init__.py#L18
11:30:45 sean-k-mooney but 2_queus is not
11:31:29 sean-k-mooney Quantitative aspects must eb tracked as inventories in a resouce provider
11:32:13 sean-k-mooney dvo-plv_: we cannot currently track neutron consumable pci devices in placment by the way
11:32:24 sean-k-mooney vmaccel: has said they want to work on that this cycle
11:33:02 sean-k-mooney so for bobcat it would be risky to asume that work woudl be compelte in time for multi queu to be implemetned
11:33:16 dvo-plv_ does it will be part of this spec ? https://specs.openstack.org/openstack/nova-specs/specs/2023.1/implemented/pci-device-tracking-in-placement.html
11:33:41 sean-k-mooney dvo-plv_: no that spec epxlictly dose not support any pci device that can be use via neutron
11:37:16 dvo-plv_ so, the main problem taht we can not filter specific node fro the pool with required vf queue number, right ?
11:37:53 sean-k-mooney we can solve that todya with the pci_passhtough_filter
11:38:16 sean-k-mooney as i said there are 3-4 peice that need to be done

Earlier   Later