Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-25
13:30:37 jianghuaw mriedem, bauzas, thanks. In that case, I believe we will proceed to implement vGPU with resource provider. right?
13:30:49 jianghuaw bauzas, cool.
13:31:15 bauzas jianghuaw: for the moment, I see vGPU resources as just traits and resource classes attached to a specific RP
13:31:54 bauzas if one implements NUMA reporting with nested RPs, that will just mean that vGPU resource classes will be attached to the child RP that is providing the NUMA cell
13:32:04 bauzas if that's a PCI device
13:32:22 jianghuaw yeah, fair enough.
13:32:52 bauzas jianghuaw: I actually raised that point in a comment in your change
13:33:11 bauzas jianghuaw: but I was waiting for jaypipes acking or not that
13:33:18 jaypipes bauzas: ack
13:33:31 bauzas cool then
13:33:34 jianghuaw :-)
13:33:41 bauzas then, no need to wait for NUMA-isms
13:33:46 jianghuaw that's great.
13:34:06 bauzas jianghuaw: I actually owe you a new review of your spec
13:34:32 bauzas and I owe jaypipes a serious review of his nested RP series
13:34:42 bauzas jaypipes: you said you were about to rebase, right?
13:35:02 bauzas mriedem: interesting bug https://bugs.launchpad.net/nova/+bug/1717915
13:35:03 openstack Launchpad bug 1717915 in oslo.messaging "nova services and transport_url, cannot connect to vhost if specified" [Undecided,New]
13:35:42 openstackgerrit Eric Fried proposed openstack/nova master: Live Migration sequence diagram https://review.openstack.org/506370
13:35:49 efried alex_xu Like-a-this? ^
13:35:50 jianghuaw bauzas, Thanks for have reviewed my spec for may rounds and given may useful comments.
13:35:55 jaypipes bauzas: done last week: https://review.openstack.org/#/q/topic:bp/nested-resource-providers
13:36:09 bauzas jaypipes: cool, that's then my top prio
13:36:23 jaypipes efried: answered on spec.
13:36:27 efried thx
13:36:49 jianghuaw bauzas, jaypipes: I do need your help on how to define that options to restrict only one vGPU type is exposed.
13:37:25 jaypipes jianghuaw: sorry, not following you... is this a patch you've proposed?
13:37:54 mriedem bauzas: yes, you'll see i was already triaging it
13:38:14 jianghuaw jaypipes, I mean this patch:https://review.openstack.org/#/c/450122
13:39:06 jianghuaw sahid suggested to make the option to be general.
13:39:17 bauzas mriedem: yup, I saw hence my ping
13:39:30 bauzas mriedem: it was an implicit 'good call, but what should we do next" ?
13:39:58 mriedem idk, i hoped that the rhops guys would know since you guys run with clustered rabbit
13:39:59 bauzas because looks like it's something that worked in the past but we never officially supported it
13:40:03 mriedem *rhosp
13:40:33 bauzas owalsh: around ?
13:41:44 bauzas jaypipes: the only point that is still concerning me about https://review.openstack.org/#/c/450122 is how we draft the whitelist
13:42:18 bauzas jaypipes: as it can be different for each virt driver
13:42:18 jaypipes bauzas: I had specifically asked jianghuaw to make the config option *not* the pci_passthrough_whitelist
13:42:26 jaypipes bauzas: how so?
13:42:34 mriedem jianghuaw: bauzas: jaypipes: ew, yeah, was just going to ask if this is a new pci whitelist but for gpus
13:42:51 jaypipes mriedem: yes, I asked for that.
13:43:07 bauzas jaypipes: well, jianghuaw made a very explicit pci-tied config option
13:43:13 bauzas I probably missed your point then
13:43:33 bauzas but I'm fine with just a ListOpt containing strings that would match device IDs
13:43:54 bauzas it would be up to the driver to find the right GPU that matches the string
13:44:16 jaypipes bauzas: are you referring to the fact that the proposed enabled_vgpu_types CONF option has vendor_id and product_id keys?
13:44:27 bauzas like enabled_vgpu_types = ['nvidia-11', 'nvidia-10']
13:44:38 bauzas yeah, that is too specific
13:45:08 bauzas jianghuaw: ^
13:45:22 edleafe Scheduler subteam meeting in 15 minutes in #openstack-meeting-alt
13:45:54 jianghuaw how about use the vGPU type's name instead of the id? by considering the type name is common from both XenServer and kvm? enabled_vgpu_types = ['GRID K160Q'...]
13:46:44 jianghuaw As I explained in the comment, both virts can get the type name - "GRID K160"
13:47:09 jianghuaw And that's also the name recorded in the user guide.
13:48:16 sdague mriedem: with you and I both with fingers in the qemu 2.10 patch, you want a 4th core to look into it - https://review.openstack.org/#/c/505673/ - or you want to just make sure my update isn't crazy and put it in
13:48:26 bauzas jianghuaw: and what if I have two exact same cards ?
13:48:49 sahid mriedem: did you notice my request, continuing development on /pci and in parallel working on porting virt features on RP
13:49:01 mriedem sdague: i looked at it last week, which is how we went down the paused rabbit hole
13:49:09 mriedem sdague: but we have plenty of cores, what's one more to review
13:49:10 sahid by this way we could still provide proction-ready features and have at some point a deprecating phease
13:49:22 sdague mriedem: right, that's fixed now
13:49:49 mriedem sahid: this? https://review.openstack.org/#/c/485522/
13:50:18 mriedem sahid: sorry no i'm not sure what you're referring to
13:50:34 sahid mriedem: no worries let me give you a pointer
13:50:48 sahid mriedem: http://lists.openstack.org/pipermail/openstack-dev/2017-September/122591.html
13:50:56 jianghuaw bauzas, hmmm. yes, that means we can only expose one type per compute node.
13:51:13 jianghuaw I think that's the reason I used a pci-tied config option;
13:52:23 jianghuaw we can specify the vGPU enabled a a specific PGPU keyed by pci address.
13:52:32 openstackgerrit Merged openstack/nova master: Add some inline code docs tracing the cold migrate flow https://review.openstack.org/496861
13:52:47 mriedem jaypipes: "The decision of whether to allow an approach that adds more to the existing /pci module is ultimately Matt's."
13:52:53 mriedem jaypipes: i didn't even see that bus coming
13:53:22 jaypipes mriedem: lol. well, sorry, but you are the PTL :)
13:54:02 bauzas jianghuaw: sec, reading your last comments then
13:56:51 bauzas jianghuaw: so, in case of libvirt, we have the capabilities XML that shows up different device IDs for two same cards (or types if you prefer)
13:56:52 efried jianghuaw Is it possible you have multiple identical GPUs (or mulitple pGPUs on the same card) where you would want to whitelist some but not others?
13:57:00 efried :)
13:57:02 sahid mriedem: jaypipes, my point is to continue make our users happy, have a deprecating phase to fix any issues for the work ongoing with RP and at some point, switch
13:57:26 bauzas jianghuaw: because each of them is corresponding to a mediated device
13:57:50 jaypipes sahid: well, we're never going to make users happy :)
13:57:56 bauzas jianghuaw: and that's the deployer's responsibility to configure libvirt so that the GPUs are shown as mediated devices
13:58:03 openstackgerrit Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275
13:58:09 jaypipes sahid: but anyway, like I said, I'm fine with you adding to the existing code.
13:58:21 jaypipes sahid: I just will be focusing on the n-r-p stuff, that's all
13:58:40 bauzas it's open-source development, anyone can just propose
13:58:58 sahid jaypipes: yes i understand that which it make sense and i will be happy to help and review some part if i can
13:59:03 bauzas I just felt we discussed on what could be achievable for Queens in parallel of the nested RPs implementation
14:00:02 bauzas the main focus for Queens was to provide the libvirt feature to returning RPs that were augmented by the VGPU resources
14:00:38 bauzas jianghuaw: can you clarify how Xen would see those GPU types ?
14:00:54 edleafe Scheduler subteam meeting running now in #openstack-meeting-alt
14:00:55 sahid bauzas: so you want to push on top of something under heavy developpement a feature which we are going to be used by large industries ? i think we should be reasonable
14:01:28 bauzas sahid: I'm just saying it's orthogonal :)
14:01:33 sahid RP is doind a lof of things, we need a deprecating phase, where some users can migrate and so we can fix the issues
14:01:35 jianghuaw bauzas, ok. XenServer will detect the pGPU types and make the same type of pGPUs into a single group.
14:01:46 bauzas vGPU tracking is done by providing a new set of resource classes
14:01:53 bauzas and traits
14:02:11 bauzas while nested resource providers is focusing on providing a tree of resource providers with dependencies
14:02:19 bauzas those don't overlap
14:02:24 jianghuaw And expose the vGPU types supported by the pGPU group.
14:02:47 jaypipes sahid: for the record, we *have* been doing that deprecation/migration period with resource providers. for example, we had a deprecation/migration period for tracking Ironic nodes as atomic resources.
14:02:49 bauzas snap, scheduler meeting

Earlier   Later