Earlier  
Posted Nick Remark
#openstack-nova - 2022-05-09
12:22:33 gibi What to do if both ``resource_class`` and ``vendor_id`` and ``product_id`` are provided in the alias?
12:23:10 sean-k-mooney good question. i guess we have too options
12:23:15 sean-k-mooney first is consider that an error
12:23:33 sean-k-mooney second is use the resouce_class for placment queries
12:23:50 sean-k-mooney and the the rest for the pci/numa filter
12:23:53 gibi the only use case I can think of is that deployer use a generic RC but then later want to refine the alias via product id
12:24:20 sean-k-mooney my secret plan is to eventually remove the need for the alias
12:24:41 sean-k-mooney so over time it woudl be nice if we coudl move to jsut having the resouce class in the alias
12:24:42 gibi do you want to go with flavor extra_spec based resource?
12:25:00 sean-k-mooney yes and no
12:25:11 sean-k-mooney i current hate that we allow grouping in the extra_specs
12:25:36 sean-k-mooney so part of me wants to keep the alisa as we can insulate operators form that
12:26:14 sean-k-mooney but i do kind of like the idea of have just resouce:CUSTOM_<whatever>=1
12:26:51 gibi OK, I get the goal that we want to move deployers to RC based alias in the future and if the deployer still want product id based filtering then the deployer can use traits for that
12:26:55 sean-k-mooney so i think having the RC take prescidence for the placment query and allowign vendor_id and product_id makes sense
12:27:40 gibi OK
12:28:06 sean-k-mooney so really what i would like is for operators to take devces with the custom resouce class and use the RC name as the "alias"
12:28:26 gibi so we keep the PCIFilter to keep filtering for vendor / product
12:28:33 gibi at least for now
12:28:46 gibi the RC name as alias make sense
12:28:57 sean-k-mooney ya for now althogh in princial i think you could trun it off if we did this right
12:29:42 sean-k-mooney my idea is
12:29:45 gibi yes if we let the product id filtering case go and no SRIOV in the deployment then we can turn off the PCIFilter
12:30:02 gibi neutron based SRIOV
12:30:16 sean-k-mooney device_list = some device -> RC gpu_gold
12:30:36 sean-k-mooney and then you could jsut ask for RC gpu_goal in the alias
12:31:01 sean-k-mooney right now the reason i dont wnat to go directly to resouce:gpu_gold=1
12:31:19 sean-k-mooney is that intoduced problems wiht vgpu and generic mdev usage
12:31:31 sean-k-mooney it shoudl be resovlable
12:31:59 sean-k-mooney i.e. if we see that the resouce does not match any of the generic_mdev types listed in teh config we know tis a pci passthough request
12:32:18 sean-k-mooney but i tought that would complicate the spec more then need initally
12:33:11 gibi ahh yeah
12:33:22 gibi thanks for the background
12:33:35 gibi next
12:34:04 gibi I feel that both you and me want to keep the dependent device handling supported. But as stephenfin said, it is a lot of complexity
12:34:14 gibi so just double checking it that you still think this is needed
12:34:27 gibi as per https://review.opendev.org/c/openstack/nova-specs/+/791047/4/specs/zed/approved/pci-device-tracking-in-placement.rst#320
12:34:40 sean-k-mooney honestly i would like to be able to remove it. but im concerned by the upgrade impact
12:35:11 sean-k-mooney i do know we have customer that want to dynically choose if they consime a device a a PF or VF when they boot the workload
12:35:23 sean-k-mooney but its not very reliable today
12:36:00 sean-k-mooney as in its easy for vms to consume 1 vf on all the devices
12:36:12 gibi I also think that there is many deployment out there that is was used without knowing it. I mean if somebody whitelisted a PF that had VFs then the VFs become scheduleable automatically
12:36:23 sean-k-mooney basically meaning you can not allocate PF even though your could have allocate the vfs differntly
12:37:03 sean-k-mooney right so i think what stephenfin had in mind was if you whitelist the PF and it has VFs we would only expose the VFs
12:37:22 sean-k-mooney where as today unless you use the product_id to filter
12:37:30 sean-k-mooney we expos both the VFs and PFs
12:37:55 sean-k-mooney if we maintian the current behavior we obvilly need ot dynamically adjst the reserved value
12:38:04 gibi yes, that is the complexity
12:38:07 sean-k-mooney to emulated the unclaimable state
12:38:08 gibi but it is solveable
12:38:31 gibi I think I will keep this open for bauzas or other reviews to chime in
12:38:59 sean-k-mooney sure
12:39:13 sean-k-mooney we decided to reduce flexiblity for cpu pinning
12:39:15 sean-k-mooney with isolate
12:39:25 sean-k-mooney and we know that not everyone was happy with that
12:40:01 sean-k-mooney we can elect to do the same here but we need to be deliberiate about it and comunicate it well if we want to force this change
12:40:09 ralonsoh sean-k-mooney, sorry, I was having lunch
12:40:13 ralonsoh what do you need?
12:40:22 sean-k-mooney if we can live with the complexity then we proably shoudl keep it
12:40:28 sean-k-mooney ralonsoh: i found it
12:40:36 gibi OK, I will plan with the complexity
12:40:37 ralonsoh ah perfect
12:40:39 sean-k-mooney ralonsoh: https://review.opendev.org/q/topic:bp%252Fenable-sriov-nic-features
12:40:47 sean-k-mooney ralonsoh: your code for tracking pci device in placment
12:41:12 sean-k-mooney ralonsoh: we are just discussing the spec to enable it. gibi will be taking on that feature
12:41:20 ralonsoh perfect
12:41:39 gibi ralonsoh: https://review.opendev.org/c/openstack/nova-specs/+/791047/4/specs/zed/approved/pci-device-tracking-in-placement.rst this is the spec if you are interested :)
12:41:47 ralonsoh sure
12:42:01 gibi sean-k-mooney: so I have one more open question
12:42:08 sean-k-mooney gibi: go for it
12:42:09 gibi upgrade
12:42:19 gibi obviously rolling upgrade is a pain
12:42:35 sean-k-mooney ah well you burn down the datacenter and build a new one in the ashes
12:42:45 sean-k-mooney obvisouly the least painful approch
12:42:46 gibi in the PCPU case we did a fallback query
12:43:01 gibi to allow schedling to not-yet upgraded computes
12:43:05 sean-k-mooney yep
12:43:13 sean-k-mooney we could make the prefilter configurable
12:43:14 gibi I'm not sure how we did the allocation in that case
12:43:19 gibi but
12:43:51 gibi in the PCI case if we select a host based on the fallback query then on that host the scheduler will not allocate PCI devices in placement
12:44:12 sean-k-mooney right so i would not use a fallback
12:44:21 sean-k-mooney by default we shoudl reprot the inventores to placemnt
12:44:28 sean-k-mooney and ahve a prefileter
12:44:43 sean-k-mooney the prefilter would add the pci device request to the query
12:44:52 sean-k-mooney and we disable it by default in zed
12:44:59 sean-k-mooney then enable it by default in AA
12:45:18 sean-k-mooney so you would rolling upgrade to Zed
12:45:29 sean-k-mooney then enable the prefilter once all host are upgraded
12:45:59 sean-k-mooney you likely would have to then do a heal-allocation like command to update the allcotiosn of existign instances
12:46:07 gibi ahh i see
12:46:34 sean-k-mooney we could also have a nova-status check
12:46:36 gibi so we report devices but we don't allocate yet
12:46:44 sean-k-mooney yep
12:46:52 gibi then when every compute is ready we do a heal and then start allocating
12:47:06 sean-k-mooney yes using the claim in the pci_devices table
12:47:26 gibi yepp we keep using the claim and the pci_device table
12:47:29 gibi anyhow
12:47:36 gibi as that tracks exact VF PCI addresses
12:47:41 gibi Placement wont

Earlier   Later