| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-05-09 | |||
| 12:26:51 | gibi | OK, I get the goal that we want to move deployers to RC based alias in the future and if the deployer still want product id based filtering then the deployer can use traits for that | |
| 12:26:55 | sean-k-mooney | so i think having the RC take prescidence for the placment query and allowign vendor_id and product_id makes sense | |
| 12:27:40 | gibi | OK | |
| 12:28:06 | sean-k-mooney | so really what i would like is for operators to take devces with the custom resouce class and use the RC name as the "alias" | |
| 12:28:26 | gibi | so we keep the PCIFilter to keep filtering for vendor / product | |
| 12:28:33 | gibi | at least for now | |
| 12:28:46 | gibi | the RC name as alias make sense | |
| 12:28:57 | sean-k-mooney | ya for now althogh in princial i think you could trun it off if we did this right | |
| 12:29:42 | sean-k-mooney | my idea is | |
| 12:29:45 | gibi | yes if we let the product id filtering case go and no SRIOV in the deployment then we can turn off the PCIFilter | |
| 12:30:02 | gibi | neutron based SRIOV | |
| 12:30:16 | sean-k-mooney | device_list = some device -> RC gpu_gold | |
| 12:30:36 | sean-k-mooney | and then you could jsut ask for RC gpu_goal in the alias | |
| 12:31:01 | sean-k-mooney | right now the reason i dont wnat to go directly to resouce:gpu_gold=1 | |
| 12:31:19 | sean-k-mooney | is that intoduced problems wiht vgpu and generic mdev usage | |
| 12:31:31 | sean-k-mooney | it shoudl be resovlable | |
| 12:31:59 | sean-k-mooney | i.e. if we see that the resouce does not match any of the generic_mdev types listed in teh config we know tis a pci passthough request | |
| 12:32:18 | sean-k-mooney | but i tought that would complicate the spec more then need initally | |
| 12:33:11 | gibi | ahh yeah | |
| 12:33:22 | gibi | thanks for the background | |
| 12:33:35 | gibi | next | |
| 12:34:04 | gibi | I feel that both you and me want to keep the dependent device handling supported. But as stephenfin said, it is a lot of complexity | |
| 12:34:14 | gibi | so just double checking it that you still think this is needed | |
| 12:34:27 | gibi | as per https://review.opendev.org/c/openstack/nova-specs/+/791047/4/specs/zed/approved/pci-device-tracking-in-placement.rst#320 | |
| 12:34:40 | sean-k-mooney | honestly i would like to be able to remove it. but im concerned by the upgrade impact | |
| 12:35:11 | sean-k-mooney | i do know we have customer that want to dynically choose if they consime a device a a PF or VF when they boot the workload | |
| 12:35:23 | sean-k-mooney | but its not very reliable today | |
| 12:36:00 | sean-k-mooney | as in its easy for vms to consume 1 vf on all the devices | |
| 12:36:12 | gibi | I also think that there is many deployment out there that is was used without knowing it. I mean if somebody whitelisted a PF that had VFs then the VFs become scheduleable automatically | |
| 12:36:23 | sean-k-mooney | basically meaning you can not allocate PF even though your could have allocate the vfs differntly | |
| 12:37:03 | sean-k-mooney | right so i think what stephenfin had in mind was if you whitelist the PF and it has VFs we would only expose the VFs | |
| 12:37:22 | sean-k-mooney | where as today unless you use the product_id to filter | |
| 12:37:30 | sean-k-mooney | we expos both the VFs and PFs | |
| 12:37:55 | sean-k-mooney | if we maintian the current behavior we obvilly need ot dynamically adjst the reserved value | |
| 12:38:04 | gibi | yes, that is the complexity | |
| 12:38:07 | sean-k-mooney | to emulated the unclaimable state | |
| 12:38:08 | gibi | but it is solveable | |
| 12:38:31 | gibi | I think I will keep this open for bauzas or other reviews to chime in | |
| 12:38:59 | sean-k-mooney | sure | |
| 12:39:13 | sean-k-mooney | we decided to reduce flexiblity for cpu pinning | |
| 12:39:15 | sean-k-mooney | with isolate | |
| 12:39:25 | sean-k-mooney | and we know that not everyone was happy with that | |
| 12:40:01 | sean-k-mooney | we can elect to do the same here but we need to be deliberiate about it and comunicate it well if we want to force this change | |
| 12:40:09 | ralonsoh | sean-k-mooney, sorry, I was having lunch | |
| 12:40:13 | ralonsoh | what do you need? | |
| 12:40:22 | sean-k-mooney | if we can live with the complexity then we proably shoudl keep it | |
| 12:40:28 | sean-k-mooney | ralonsoh: i found it | |
| 12:40:36 | gibi | OK, I will plan with the complexity | |
| 12:40:37 | ralonsoh | ah perfect | |
| 12:40:39 | sean-k-mooney | ralonsoh: https://review.opendev.org/q/topic:bp%252Fenable-sriov-nic-features | |
| 12:40:47 | sean-k-mooney | ralonsoh: your code for tracking pci device in placment | |
| 12:41:12 | sean-k-mooney | ralonsoh: we are just discussing the spec to enable it. gibi will be taking on that feature | |
| 12:41:20 | ralonsoh | perfect | |
| 12:41:39 | gibi | ralonsoh: https://review.opendev.org/c/openstack/nova-specs/+/791047/4/specs/zed/approved/pci-device-tracking-in-placement.rst this is the spec if you are interested :) | |
| 12:41:47 | ralonsoh | sure | |
| 12:42:01 | gibi | sean-k-mooney: so I have one more open question | |
| 12:42:08 | sean-k-mooney | gibi: go for it | |
| 12:42:09 | gibi | upgrade | |
| 12:42:19 | gibi | obviously rolling upgrade is a pain | |
| 12:42:35 | sean-k-mooney | ah well you burn down the datacenter and build a new one in the ashes | |
| 12:42:45 | sean-k-mooney | obvisouly the least painful approch | |
| 12:42:46 | gibi | in the PCPU case we did a fallback query | |
| 12:43:01 | gibi | to allow schedling to not-yet upgraded computes | |
| 12:43:05 | sean-k-mooney | yep | |
| 12:43:13 | sean-k-mooney | we could make the prefilter configurable | |
| 12:43:14 | gibi | I'm not sure how we did the allocation in that case | |
| 12:43:19 | gibi | but | |
| 12:43:51 | gibi | in the PCI case if we select a host based on the fallback query then on that host the scheduler will not allocate PCI devices in placement | |
| 12:44:12 | sean-k-mooney | right so i would not use a fallback | |
| 12:44:21 | sean-k-mooney | by default we shoudl reprot the inventores to placemnt | |
| 12:44:28 | sean-k-mooney | and ahve a prefileter | |
| 12:44:43 | sean-k-mooney | the prefilter would add the pci device request to the query | |
| 12:44:52 | sean-k-mooney | and we disable it by default in zed | |
| 12:44:59 | sean-k-mooney | then enable it by default in AA | |
| 12:45:18 | sean-k-mooney | so you would rolling upgrade to Zed | |
| 12:45:29 | sean-k-mooney | then enable the prefilter once all host are upgraded | |
| 12:45:59 | sean-k-mooney | you likely would have to then do a heal-allocation like command to update the allcotiosn of existign instances | |
| 12:46:07 | gibi | ahh i see | |
| 12:46:34 | sean-k-mooney | we could also have a nova-status check | |
| 12:46:36 | gibi | so we report devices but we don't allocate yet | |
| 12:46:44 | sean-k-mooney | yep | |
| 12:46:52 | gibi | then when every compute is ready we do a heal and then start allocating | |
| 12:47:06 | sean-k-mooney | yes using the claim in the pci_devices table | |
| 12:47:26 | gibi | yepp we keep using the claim and the pci_device table | |
| 12:47:29 | gibi | anyhow | |
| 12:47:36 | gibi | as that tracks exact VF PCI addresses | |
| 12:47:41 | gibi | Placement wont | |
| 12:47:48 | sean-k-mooney | yep so form the contolers | |
| 12:48:00 | sean-k-mooney | we will have all the info in the pci_devices table to heal the allocations | |
| 12:48:17 | sean-k-mooney | since we also have the parent adresses | |
| 12:48:24 | sean-k-mooney | we can constuct the RP names | |
| 12:49:07 | gibi | yepp | |
| 12:49:39 | sean-k-mooney | we could consider | |
| 12:49:51 | sean-k-mooney | if we can activate teh filter based on min compute service version | |
| 12:50:21 | sean-k-mooney | if we were to do that we would likely need the compute agent to heal the allcoations automaticaly | |
| 12:50:38 | sean-k-mooney | perhaps on starup or in the upsadate_avaiable_resouces periodic task | |
| 12:51:29 | sean-k-mooney | im not sure if we want that level of compleixty but we already do reshapes in init_host | |
| 12:52:11 | sean-k-mooney | do you think that is too much "magic"/complexity | |
| 12:52:38 | sean-k-mooney | it would make the operator experince much nicer as it woudl just start working once everythin was upgraded | |
| 12:52:42 | gibi | this reshap will be just RP creation, we won't move things | |