| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-25 | |||
| 13:29:29 | mriedem | jianghuaw: we plan on implemented nested resource providers in queens | |
| 13:29:39 | mriedem | but i don't think we plan on focusing on numa use cases | |
| 13:29:49 | mriedem | *implementing | |
| 13:29:49 | bauzas | jianghuaw: I don't think we should worry on NUMA resources now | |
| 13:30:29 | bauzas | jianghuaw: because once compute nodes will report NUMA cells using nested RPs, we'll get both vGPU support and NUMA handling tied together | |
| 13:30:37 | jianghuaw | mriedem, bauzas, thanks. In that case, I believe we will proceed to implement vGPU with resource provider. right? | |
| 13:30:49 | jianghuaw | bauzas, cool. | |
| 13:31:15 | bauzas | jianghuaw: for the moment, I see vGPU resources as just traits and resource classes attached to a specific RP | |
| 13:31:54 | bauzas | if one implements NUMA reporting with nested RPs, that will just mean that vGPU resource classes will be attached to the child RP that is providing the NUMA cell | |
| 13:32:04 | bauzas | if that's a PCI device | |
| 13:32:22 | jianghuaw | yeah, fair enough. | |
| 13:32:52 | bauzas | jianghuaw: I actually raised that point in a comment in your change | |
| 13:33:11 | bauzas | jianghuaw: but I was waiting for jaypipes acking or not that | |
| 13:33:18 | jaypipes | bauzas: ack | |
| 13:33:31 | bauzas | cool then | |
| 13:33:34 | jianghuaw | :-) | |
| 13:33:41 | bauzas | then, no need to wait for NUMA-isms | |
| 13:33:46 | jianghuaw | that's great. | |
| 13:34:06 | bauzas | jianghuaw: I actually owe you a new review of your spec | |
| 13:34:32 | bauzas | and I owe jaypipes a serious review of his nested RP series | |
| 13:34:42 | bauzas | jaypipes: you said you were about to rebase, right? | |
| 13:35:02 | bauzas | mriedem: interesting bug https://bugs.launchpad.net/nova/+bug/1717915 | |
| 13:35:03 | openstack | Launchpad bug 1717915 in oslo.messaging "nova services and transport_url, cannot connect to vhost if specified" [Undecided,New] | |
| 13:35:42 | openstackgerrit | Eric Fried proposed openstack/nova master: Live Migration sequence diagram https://review.openstack.org/506370 | |
| 13:35:49 | efried | alex_xu Like-a-this? ^ | |
| 13:35:50 | jianghuaw | bauzas, Thanks for have reviewed my spec for may rounds and given may useful comments. | |
| 13:35:55 | jaypipes | bauzas: done last week: https://review.openstack.org/#/q/topic:bp/nested-resource-providers | |
| 13:36:09 | bauzas | jaypipes: cool, that's then my top prio | |
| 13:36:23 | jaypipes | efried: answered on spec. | |
| 13:36:27 | efried | thx | |
| 13:36:49 | jianghuaw | bauzas, jaypipes: I do need your help on how to define that options to restrict only one vGPU type is exposed. | |
| 13:37:25 | jaypipes | jianghuaw: sorry, not following you... is this a patch you've proposed? | |
| 13:37:54 | mriedem | bauzas: yes, you'll see i was already triaging it | |
| 13:38:14 | jianghuaw | jaypipes, I mean this patch:https://review.openstack.org/#/c/450122 | |
| 13:39:06 | jianghuaw | sahid suggested to make the option to be general. | |
| 13:39:17 | bauzas | mriedem: yup, I saw hence my ping | |
| 13:39:30 | bauzas | mriedem: it was an implicit 'good call, but what should we do next" ? | |
| 13:39:58 | mriedem | idk, i hoped that the rhops guys would know since you guys run with clustered rabbit | |
| 13:39:59 | bauzas | because looks like it's something that worked in the past but we never officially supported it | |
| 13:40:03 | mriedem | *rhosp | |
| 13:40:33 | bauzas | owalsh: around ? | |
| 13:41:44 | bauzas | jaypipes: the only point that is still concerning me about https://review.openstack.org/#/c/450122 is how we draft the whitelist | |
| 13:42:18 | jaypipes | bauzas: I had specifically asked jianghuaw to make the config option *not* the pci_passthrough_whitelist | |
| 13:42:18 | bauzas | jaypipes: as it can be different for each virt driver | |
| 13:42:26 | jaypipes | bauzas: how so? | |
| 13:42:34 | mriedem | jianghuaw: bauzas: jaypipes: ew, yeah, was just going to ask if this is a new pci whitelist but for gpus | |
| 13:42:51 | jaypipes | mriedem: yes, I asked for that. | |
| 13:43:07 | bauzas | jaypipes: well, jianghuaw made a very explicit pci-tied config option | |
| 13:43:13 | bauzas | I probably missed your point then | |
| 13:43:33 | bauzas | but I'm fine with just a ListOpt containing strings that would match device IDs | |
| 13:43:54 | bauzas | it would be up to the driver to find the right GPU that matches the string | |
| 13:44:16 | jaypipes | bauzas: are you referring to the fact that the proposed enabled_vgpu_types CONF option has vendor_id and product_id keys? | |
| 13:44:27 | bauzas | like enabled_vgpu_types = ['nvidia-11', 'nvidia-10'] | |
| 13:44:38 | bauzas | yeah, that is too specific | |
| 13:45:08 | bauzas | jianghuaw: ^ | |
| 13:45:22 | edleafe | Scheduler subteam meeting in 15 minutes in #openstack-meeting-alt | |
| 13:45:54 | jianghuaw | how about use the vGPU type's name instead of the id? by considering the type name is common from both XenServer and kvm? enabled_vgpu_types = ['GRID K160Q'...] | |
| 13:46:44 | jianghuaw | As I explained in the comment, both virts can get the type name - "GRID K160" | |
| 13:47:09 | jianghuaw | And that's also the name recorded in the user guide. | |
| 13:48:16 | sdague | mriedem: with you and I both with fingers in the qemu 2.10 patch, you want a 4th core to look into it - https://review.openstack.org/#/c/505673/ - or you want to just make sure my update isn't crazy and put it in | |
| 13:48:26 | bauzas | jianghuaw: and what if I have two exact same cards ? | |
| 13:48:49 | sahid | mriedem: did you notice my request, continuing development on /pci and in parallel working on porting virt features on RP | |
| 13:49:01 | mriedem | sdague: i looked at it last week, which is how we went down the paused rabbit hole | |
| 13:49:09 | mriedem | sdague: but we have plenty of cores, what's one more to review | |
| 13:49:10 | sahid | by this way we could still provide proction-ready features and have at some point a deprecating phease | |
| 13:49:22 | sdague | mriedem: right, that's fixed now | |
| 13:49:49 | mriedem | sahid: this? https://review.openstack.org/#/c/485522/ | |
| 13:50:18 | mriedem | sahid: sorry no i'm not sure what you're referring to | |
| 13:50:34 | sahid | mriedem: no worries let me give you a pointer | |
| 13:50:48 | sahid | mriedem: http://lists.openstack.org/pipermail/openstack-dev/2017-September/122591.html | |
| 13:50:56 | jianghuaw | bauzas, hmmm. yes, that means we can only expose one type per compute node. | |
| 13:51:13 | jianghuaw | I think that's the reason I used a pci-tied config option; | |
| 13:52:23 | jianghuaw | we can specify the vGPU enabled a a specific PGPU keyed by pci address. | |
| 13:52:32 | openstackgerrit | Merged openstack/nova master: Add some inline code docs tracing the cold migrate flow https://review.openstack.org/496861 | |
| 13:52:47 | mriedem | jaypipes: "The decision of whether to allow an approach that adds more to the existing /pci module is ultimately Matt's." | |
| 13:52:53 | mriedem | jaypipes: i didn't even see that bus coming | |
| 13:53:22 | jaypipes | mriedem: lol. well, sorry, but you are the PTL :) | |
| 13:54:02 | bauzas | jianghuaw: sec, reading your last comments then | |
| 13:56:51 | bauzas | jianghuaw: so, in case of libvirt, we have the capabilities XML that shows up different device IDs for two same cards (or types if you prefer) | |
| 13:56:52 | efried | jianghuaw Is it possible you have multiple identical GPUs (or mulitple pGPUs on the same card) where you would want to whitelist some but not others? | |
| 13:57:00 | efried | :) | |
| 13:57:02 | sahid | mriedem: jaypipes, my point is to continue make our users happy, have a deprecating phase to fix any issues for the work ongoing with RP and at some point, switch | |
| 13:57:26 | bauzas | jianghuaw: because each of them is corresponding to a mediated device | |
| 13:57:50 | jaypipes | sahid: well, we're never going to make users happy :) | |
| 13:57:56 | bauzas | jianghuaw: and that's the deployer's responsibility to configure libvirt so that the GPUs are shown as mediated devices | |
| 13:58:03 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275 | |
| 13:58:09 | jaypipes | sahid: but anyway, like I said, I'm fine with you adding to the existing code. | |
| 13:58:21 | jaypipes | sahid: I just will be focusing on the n-r-p stuff, that's all | |
| 13:58:40 | bauzas | it's open-source development, anyone can just propose | |
| 13:58:58 | sahid | jaypipes: yes i understand that which it make sense and i will be happy to help and review some part if i can | |
| 13:59:03 | bauzas | I just felt we discussed on what could be achievable for Queens in parallel of the nested RPs implementation | |
| 14:00:02 | bauzas | the main focus for Queens was to provide the libvirt feature to returning RPs that were augmented by the VGPU resources | |
| 14:00:38 | bauzas | jianghuaw: can you clarify how Xen would see those GPU types ? | |
| 14:00:54 | edleafe | Scheduler subteam meeting running now in #openstack-meeting-alt | |
| 14:00:55 | sahid | bauzas: so you want to push on top of something under heavy developpement a feature which we are going to be used by large industries ? i think we should be reasonable | |
| 14:01:28 | bauzas | sahid: I'm just saying it's orthogonal :) | |
| 14:01:33 | sahid | RP is doind a lof of things, we need a deprecating phase, where some users can migrate and so we can fix the issues | |
| 14:01:35 | jianghuaw | bauzas, ok. XenServer will detect the pGPU types and make the same type of pGPUs into a single group. | |
| 14:01:46 | bauzas | vGPU tracking is done by providing a new set of resource classes | |
| 14:01:53 | bauzas | and traits | |