Earlier  
Posted Nick Remark
#openstack-nova - 2021-06-23
13:26:36 gibi you as a user get the same vgpu in your vm
13:26:37 bauzas that's my point
13:27:17 sean-k-mooney that point is incompatbale with "use different RC to tack ownwership" they are identical form a user point of view so shoudl have the same resouce class
13:27:17 bauzas if a device profile is a set of request groups, OK
13:27:18 gibi so if two service represents non interchangeable resource then the difference in capability needs to be modelled in placement (as a trait)
13:27:54 gibi if the two service represents the interchangeable resource then we use the same RC and we forbid to deploy a mixed cloud
13:28:11 gibi as a mixed could would be undecideable
13:28:22 bauzas sean-k-mooney: in this case (the mixed one), this is absolutely OK to get allocation candidates from a device profile that AREN'T cyborg managed hosts
13:28:27 sean-k-mooney gibi: well you could use aggreate for that
13:28:36 gibi sean-k-mooney: yeah we can allow that
13:28:40 gibi maybe
13:29:23 sean-k-mooney bauzas: i dont agree
13:29:52 sean-k-mooney if that is the case then its absolutely ok to get allcoations form a pci passthough alias form cyborg
13:29:59 sean-k-mooney or to get mdevs from cyborg
13:30:07 sean-k-mooney when using generic mdevs
13:30:39 bauzas agreed, with VGPU standard request, you could end up getting a cyborg managed host
13:30:42 sean-k-mooney we can support that but if we do it has to be symetirc
13:30:45 bauzas and I'm OK with this
13:31:06 bauzas but you said it's more than a standard RC
13:31:24 sean-k-mooney lets take that case
13:31:32 sean-k-mooney you have resouce:vgpu1
13:31:34 bauzas but as I said, I'm OK with getting results using an explicit request group that maps a device profile
13:31:42 sean-k-mooney its alocated form a cyborg RP
13:31:49 sean-k-mooney how do you boot that vm
13:32:23 sean-k-mooney bauzas: the device profile is in its own requet group by the way
13:32:24 bauzas the conductor which gets the candidates would see this is cyborg target and accordingly would ask the agent to bind the arq
13:32:32 sean-k-mooney bauzas: what arq
13:32:41 sean-k-mooney we just had resouces:vgpu
13:32:55 sean-k-mooney there is not arq or device type referenced in the flavor
13:33:02 sean-k-mooney *device-profile
13:33:17 sean-k-mooney and without a device-profile you cannot create an arq
13:34:01 bauzas so, again, that means we CAN'T be scheduler agnostic
13:34:23 bauzas in this case, you're asking for orange
13:34:29 bauzas and not appel
13:34:31 bauzas apple
13:34:45 sean-k-mooney but we agree that the device the guest sees is idential
13:34:53 sean-k-mooney so they are apples
13:35:13 gibi nova implementation needs a distintion but for the user the result is the same
13:35:14 sean-k-mooney they just have different requiremetn to consume them
13:35:24 sean-k-mooney gibi: exactly
13:35:30 gibi that is bad :/
13:35:40 sean-k-mooney this is an internal impleemntaiton detail of how the cloud was deployed
13:35:52 gibi we have two way to get the to the same goal :/
13:36:05 sean-k-mooney technially 3
13:36:11 gibi :D
13:36:28 sean-k-mooney amd and nviai now supprot vgpu via palin sriov
13:36:38 sean-k-mooney so you can jsut use pci passthough
13:36:39 bauzas what I'm sad is that I thought cyborg would accept the mixed scenario
13:36:51 bauzas but they explicitely ask to split the cloud
13:37:01 bauzas due to implementation details
13:37:12 sean-k-mooney its not cyborg fault
13:37:23 sean-k-mooney nova is not being expresive enough in its query to placment
13:37:24 bauzas as they're unable to post-create the resources
13:37:33 bauzas no
13:37:41 sean-k-mooney no we coud but we do not have the info required
13:37:43 bauzas you're asking placement to give you apples
13:38:04 bauzas it's not nova's fault that you ask for apples coming specially from Tesco
13:38:07 sean-k-mooney for example we could list all the device-profile that request the vgpu resouce and select one
13:38:18 sean-k-mooney and use that to create teh arq
13:38:35 sean-k-mooney ah but that is not what we asked for
13:38:42 sean-k-mooney we assed for resouce:vgpu=1
13:38:47 sean-k-mooney with no other qulifyer
13:39:01 bauzas ok, back to the original thoughts (that's already 2 hours we're tangling the problem)
13:39:07 bauzas we progressed
13:39:07 sean-k-mooney the current propsal uses ownwer traits and a prefilter to make that owrk
13:39:32 bauzas we know we wanna express 'placement, give me apples that come from tesco and not from carrefour"
13:39:53 sean-k-mooney transparent ot the user or admin by adding a required trati for nova or forbidnint trait for cyborg in the resouce:vgpu case
13:40:06 bauzas sean-k-mooney: in this case, I'd prefer the forbidden trait way
13:40:22 sean-k-mooney we can do that but it makes nova special
13:40:23 bauzas "placement, give me apples from Tesco" in the cyborg case
13:40:41 bauzas "placement, give me apples not coming from Tesco as I know you got some of them"
13:40:57 bauzas sean-k-mooney: only if cyborg is configured on this cloud
13:41:05 bauzas this is the "as I know"
13:41:34 sean-k-mooney that works if the class is only shared across 2 service but not 3+
13:41:36 bauzas nothing would change for operators who don't give a single penny to cyborg
13:41:54 sean-k-mooney there is another way i guess
13:42:05 sean-k-mooney we could use placment aggreate to track ownership
13:42:07 bauzas the other way is aggregates
13:42:29 bauzas "placement, don't give me apples coming from those supermarkets"
13:42:32 sean-k-mooney each proejct woudl add all there rps to an aggate
13:42:49 bauzas sean-k-mooney: I insist
13:42:52 sean-k-mooney and then nova would in the nova case add a member_of requirement to the vgpu request group
13:43:21 bauzas sean-k-mooney: whatever the solution is, the prefilter which would have the excluding logic (for the nova case) should only do this if cyborg manages hosts
13:43:26 opendevreview Lee Yarwood proposed openstack/nova master: Add check job for FIPS https://review.opendev.org/c/openstack/nova/+/790519
13:43:45 sean-k-mooney bauzas: that is trival to do
13:43:57 sean-k-mooney bauzas: we just have a config option for it
13:44:06 sean-k-mooney like pcpus
13:44:06 bauzas sean-k-mooney: I don't want upgrade impacts if the operator doesn't use cyborg yet
13:44:44 bauzas sean-k-mooney: can we express "give me hosts that are not from this aggregate" ?
13:44:44 sean-k-mooney there wont be one if we do this correctly regarless of ownwer traits ro differnt resouce classes
13:44:54 sean-k-mooney bauzas: yes
13:45:09 bauzas ok, so there are no operations for nova-managed hosts
13:45:12 sean-k-mooney but i do not think we shoudl do a negitive member_or for cyborg
13:45:20 bauzas sean-k-mooney: the other way
13:45:33 bauzas sean-k-mooney: cyborg would manage an aggregate of managed hosts
13:45:53 sean-k-mooney for nova the prefilter when enabled would do member_op=nova-aggate or member_of=cyborg-aggreate
13:45:56 bauzas the prefilter would ask for hosts that are members of this agg in the cyborg case
13:46:03 sean-k-mooney when prefilter is off it wont do anything
13:46:25 bauzas but in the nova case, the prefilter would ask hosts that AREN'T hosts of this agg
13:46:40 bauzas sean-k-mooney: again, I don't wanna manage a fleet of nova hosts

Earlier   Later