Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-25
13:35:49 efried alex_xu Like-a-this? ^
13:35:50 jianghuaw bauzas, Thanks for have reviewed my spec for may rounds and given may useful comments.
13:35:55 jaypipes bauzas: done last week: https://review.openstack.org/#/q/topic:bp/nested-resource-providers
13:36:09 bauzas jaypipes: cool, that's then my top prio
13:36:23 jaypipes efried: answered on spec.
13:36:27 efried thx
13:36:49 jianghuaw bauzas, jaypipes: I do need your help on how to define that options to restrict only one vGPU type is exposed.
13:37:25 jaypipes jianghuaw: sorry, not following you... is this a patch you've proposed?
13:37:54 mriedem bauzas: yes, you'll see i was already triaging it
13:38:14 jianghuaw jaypipes, I mean this patch:https://review.openstack.org/#/c/450122
13:39:06 jianghuaw sahid suggested to make the option to be general.
13:39:17 bauzas mriedem: yup, I saw hence my ping
13:39:30 bauzas mriedem: it was an implicit 'good call, but what should we do next" ?
13:39:58 mriedem idk, i hoped that the rhops guys would know since you guys run with clustered rabbit
13:39:59 bauzas because looks like it's something that worked in the past but we never officially supported it
13:40:03 mriedem *rhosp
13:40:33 bauzas owalsh: around ?
13:41:44 bauzas jaypipes: the only point that is still concerning me about https://review.openstack.org/#/c/450122 is how we draft the whitelist
13:42:18 bauzas jaypipes: as it can be different for each virt driver
13:42:18 jaypipes bauzas: I had specifically asked jianghuaw to make the config option *not* the pci_passthrough_whitelist
13:42:26 jaypipes bauzas: how so?
13:42:34 mriedem jianghuaw: bauzas: jaypipes: ew, yeah, was just going to ask if this is a new pci whitelist but for gpus
13:42:51 jaypipes mriedem: yes, I asked for that.
13:43:07 bauzas jaypipes: well, jianghuaw made a very explicit pci-tied config option
13:43:13 bauzas I probably missed your point then
13:43:33 bauzas but I'm fine with just a ListOpt containing strings that would match device IDs
13:43:54 bauzas it would be up to the driver to find the right GPU that matches the string
13:44:16 jaypipes bauzas: are you referring to the fact that the proposed enabled_vgpu_types CONF option has vendor_id and product_id keys?
13:44:27 bauzas like enabled_vgpu_types = ['nvidia-11', 'nvidia-10']
13:44:38 bauzas yeah, that is too specific
13:45:08 bauzas jianghuaw: ^
13:45:22 edleafe Scheduler subteam meeting in 15 minutes in #openstack-meeting-alt
13:45:54 jianghuaw how about use the vGPU type's name instead of the id? by considering the type name is common from both XenServer and kvm? enabled_vgpu_types = ['GRID K160Q'...]
13:46:44 jianghuaw As I explained in the comment, both virts can get the type name - "GRID K160"
13:47:09 jianghuaw And that's also the name recorded in the user guide.
13:48:16 sdague mriedem: with you and I both with fingers in the qemu 2.10 patch, you want a 4th core to look into it - https://review.openstack.org/#/c/505673/ - or you want to just make sure my update isn't crazy and put it in
13:48:26 bauzas jianghuaw: and what if I have two exact same cards ?
13:48:49 sahid mriedem: did you notice my request, continuing development on /pci and in parallel working on porting virt features on RP
13:49:01 mriedem sdague: i looked at it last week, which is how we went down the paused rabbit hole
13:49:09 mriedem sdague: but we have plenty of cores, what's one more to review
13:49:10 sahid by this way we could still provide proction-ready features and have at some point a deprecating phease
13:49:22 sdague mriedem: right, that's fixed now
13:49:49 mriedem sahid: this? https://review.openstack.org/#/c/485522/
13:50:18 mriedem sahid: sorry no i'm not sure what you're referring to
13:50:34 sahid mriedem: no worries let me give you a pointer
13:50:48 sahid mriedem: http://lists.openstack.org/pipermail/openstack-dev/2017-September/122591.html
13:50:56 jianghuaw bauzas, hmmm. yes, that means we can only expose one type per compute node.
13:51:13 jianghuaw I think that's the reason I used a pci-tied config option;
13:52:23 jianghuaw we can specify the vGPU enabled a a specific PGPU keyed by pci address.
13:52:32 openstackgerrit Merged openstack/nova master: Add some inline code docs tracing the cold migrate flow https://review.openstack.org/496861
13:52:47 mriedem jaypipes: "The decision of whether to allow an approach that adds more to the existing /pci module is ultimately Matt's."
13:52:53 mriedem jaypipes: i didn't even see that bus coming
13:53:22 jaypipes mriedem: lol. well, sorry, but you are the PTL :)
13:54:02 bauzas jianghuaw: sec, reading your last comments then
13:56:51 bauzas jianghuaw: so, in case of libvirt, we have the capabilities XML that shows up different device IDs for two same cards (or types if you prefer)
13:56:52 efried jianghuaw Is it possible you have multiple identical GPUs (or mulitple pGPUs on the same card) where you would want to whitelist some but not others?
13:57:00 efried :)
13:57:02 sahid mriedem: jaypipes, my point is to continue make our users happy, have a deprecating phase to fix any issues for the work ongoing with RP and at some point, switch
13:57:26 bauzas jianghuaw: because each of them is corresponding to a mediated device
13:57:50 jaypipes sahid: well, we're never going to make users happy :)
13:57:56 bauzas jianghuaw: and that's the deployer's responsibility to configure libvirt so that the GPUs are shown as mediated devices
13:58:03 openstackgerrit Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275
13:58:09 jaypipes sahid: but anyway, like I said, I'm fine with you adding to the existing code.
13:58:21 jaypipes sahid: I just will be focusing on the n-r-p stuff, that's all
13:58:40 bauzas it's open-source development, anyone can just propose
13:58:58 sahid jaypipes: yes i understand that which it make sense and i will be happy to help and review some part if i can
13:59:03 bauzas I just felt we discussed on what could be achievable for Queens in parallel of the nested RPs implementation
14:00:02 bauzas the main focus for Queens was to provide the libvirt feature to returning RPs that were augmented by the VGPU resources
14:00:38 bauzas jianghuaw: can you clarify how Xen would see those GPU types ?
14:00:54 edleafe Scheduler subteam meeting running now in #openstack-meeting-alt
14:00:55 sahid bauzas: so you want to push on top of something under heavy developpement a feature which we are going to be used by large industries ? i think we should be reasonable
14:01:28 bauzas sahid: I'm just saying it's orthogonal :)
14:01:33 sahid RP is doind a lof of things, we need a deprecating phase, where some users can migrate and so we can fix the issues
14:01:35 jianghuaw bauzas, ok. XenServer will detect the pGPU types and make the same type of pGPUs into a single group.
14:01:46 bauzas vGPU tracking is done by providing a new set of resource classes
14:01:53 bauzas and traits
14:02:11 bauzas while nested resource providers is focusing on providing a tree of resource providers with dependencies
14:02:19 bauzas those don't overlap
14:02:24 jianghuaw And expose the vGPU types supported by the pGPU group.
14:02:47 jaypipes sahid: for the record, we *have* been doing that deprecation/migration period with resource providers. for example, we had a deprecation/migration period for tracking Ironic nodes as atomic resources.
14:02:49 bauzas snap, scheduler meeting
14:03:03 jianghuaw when requesting a vGPU, xenserver will schedule a PGPU and create a vGPU on it.
14:03:14 jaypipes sahid: so it's not true that we're just imposing resource providers modeling without providing a migration path for existing resource classes.
14:03:41 jianghuaw bauzas, so we can't define whitelist per pGPU by per pGPU group.
14:03:51 bauzas sahid: jaypipes: probably a good call for discussing that in the scheduler meeting
14:03:54 sahid jaypipes: oh yes right, but, what about NUMA, CPU Pinning, Huge Pages and all the virt specific features?
14:04:08 sahid don't we have to migr incrementally?
14:04:18 jaypipes sahid: yes, absolutely.
14:04:32 jaypipes sahid: none of those resources are planned to migrate in Queens, BTW.
14:04:41 sahid jaypipes: yes and what is the ETA for sometging production-ready?
14:04:48 sahid ok cool that is my point
14:04:50 bauzas well, with the fact that CPU pinning isn't targeted to be pushed to Placement, right?
14:04:50 owalsh bauzas: hey
14:04:53 jianghuaw bauzas, that's why we make the pGPU group as the vGPU's resource provider.
14:05:16 bauzas jianghuaw: so, say I have two same cards, I'd only see one pGPU group right?
14:05:27 jianghuaw true.
14:05:28 sahid jaypipes: it will take at least 2 or 3 releases and during that time we can't make any development?
14:05:39 bauzas jianghuaw: okay, then that's a bit different from libvirt
14:06:01 sahid i just suggest we work in parallel
14:06:29 jianghuaw bauzas, could you give an example for 'that shows up different device IDs for two same cards.'

Earlier   Later