| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-31 | |||
| 18:32:57 | efried | But where did we find out that one of the 8 GPUs was called Y? | |
| 18:33:36 | dansmith | well, none of that actually works today really, so.. nowhere :) | |
| 18:33:37 | dansmith | however, | |
| 18:33:44 | dansmith | nova has tables for pci devices reported from the compute nodes, with actual addresses and stuff | |
| 18:33:46 | dansmith | if that's what you mean | |
| 18:35:06 | efried | dansmith But those tables aren't plugged into placement/RP yet, right? | |
| 18:35:37 | dansmith | they won't be really | |
| 18:35:50 | efried | Certainly not in their current form, yeah. | |
| 18:36:02 | dansmith | there is some work to do to plug in a pci device as a provider, under the compute provider, which provides VF_THINGYs or whatever | |
| 18:36:10 | dansmith | that's the nested resource providers stuff yeah | |
| 18:36:53 | openstackgerrit | Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950 | |
| 18:37:24 | efried | dansmith Yeah, okay. I'm trying to synthesize ^ with stuff jaypipes has said plus functionality needed by e.g. HyperV (claudiub) and PowerVM (moi) and others (various DOA blueprints) for discussion at the PTG. | |
| 18:37:43 | dansmith | yeah, good effing luck with that :P | |
| 18:38:48 | efried | So it seems like what came out of the above is: We don't have a framework in place that knows how to identify individual resources within a RP. | |
| 18:38:52 | efried | And | |
| 18:38:58 | efried | We need that. | |
| 18:39:23 | efried | If we're going to talk about device passthrough of any sort (GPU, SR-IOV, etc.) | |
| 18:39:37 | dansmith | yes, I still have minor issues with the way you're phrasing it, | |
| 18:39:39 | dansmith | but only because RP is really a placement thing and placement won't ever know *which* one | |
| 18:39:48 | dansmith | but from nova's perspective it'll be closer to what you're saying | |
| 18:39:58 | dansmith | well, | |
| 18:40:00 | efried | dansmith Please do help me phrase things correctly. | |
| 18:40:07 | dansmith | pci passthrough of specific SRIOV devices works today, | |
| 18:40:14 | efried | via placement? | |
| 18:40:40 | dansmith | it won't ever be *via* placement, but we'll be using placement to do our accounting of *how many* of those things we have available | |
| 18:40:49 | efried | So part of what I'm driving here is total replacement of the existing PCI passthrough setup. | |
| 18:40:57 | dansmith | it won't be materially different from today, except faster, cleaner, and less racy | |
| 18:41:15 | dansmith | no, because the existing pci passthrough stuff knows *which* and placement does not and will not | |
| 18:41:37 | dansmith | in tha' future, | |
| 18:41:50 | dansmith | we'll pick a host with enough GPU_THINGs, reserve one of the eight it provides, | |
| 18:42:04 | dansmith | and then later we get to the compute mostly guaranteed that one will be available, where we'll pick the actual one | |
| 18:42:15 | dansmith | but placement will only ever know that we're using one of the eight | |
| 18:42:16 | efried | Who's that second 'we'? | |
| 18:42:24 | efried | The virt driver? | |
| 18:42:38 | efried | The resource tracker? | |
| 18:42:40 | dansmith | the we before "pick the actual one" ? | |
| 18:42:44 | efried | yeah | |
| 18:43:04 | dansmith | well, today it's super confusing I think, as it's kindof "the scheduler filter" I think | |
| 18:43:07 | dansmith | for certain things | |
| 18:43:11 | dansmith | but regardless, | |
| 18:43:37 | dansmith | that we will be something on the nova side, likely closer to the compute node than not | |
| 18:43:38 | dansmith | clear as mud? :) | |
| 18:43:39 | dansmith | jaypipes may have a better idea of how that will actually shake out long term, | |
| 18:43:48 | dansmith | but regardless, placement will only ever be counting things, not choosing things | |
| 18:44:07 | efried | dansmith Yeah, this is a shared jaypipes dream. | |
| 18:44:19 | efried | Lest you should think it was just efried's pipe dream | |
| 18:45:49 | efried | Okay, so there's gonna be a generic device passthrough management module of some sort that sits in nova and coordinates among placement, scheduler, and virt driver to do whitelisting, claims, allocations | |
| 18:46:04 | dansmith | well, there is that today :) | |
| 18:46:24 | efried | Just for PCI, right? | |
| 18:46:29 | dansmith | (sans the telling-placement-about-it) thing | |
| 18:46:55 | dansmith | ...just for pci.... aren't we talking about pci here? | |
| 18:47:23 | efried | Well, that's one of Jay's bugbears: it should work the same for any device, not just PCI. | |
| 18:47:38 | dansmith | um | |
| 18:47:43 | dansmith | I'm confused then | |
| 18:47:51 | efried | Which intersects with one of my bugbears, which is: not every device has a PCI address. | |
| 18:48:13 | dansmith | GPU devices need not be managed as raw PCI devices for this I think | |
| 18:48:35 | dansmith | maybe that's what you mean | |
| 18:48:55 | dansmith | so I think eventually you should be able to say I want VGPU=1, traits=nvidia,gen5 | |
| 18:49:06 | dansmith | and placement will pick you a host that has one of those available | |
| 18:49:36 | dansmith | when you get to the host, that compute node will tell the virt driver (which I think will be generalizing gpus, at least in the case of libvirt), to give it a gpu, which it'll assign to your guest | |
| 18:49:41 | artom | efried, is not every device having a PCI address a problem? My understanding is that the compute reports its resources (devices) to placement | |
| 18:49:50 | artom | How it figures out what resources it has is up to it | |
| 18:50:03 | dansmith | artom: it doesn't report which resources it has, it reports how many of each it has | |
| 18:50:32 | artom | dansmith, ah, right, thanks | |
| 18:50:45 | efried | Right right. So there's the gap. Right now, artom, not having a PCI address *is* a pretty serious problem. | |
| 18:50:51 | openstackgerrit | Lucian Petrut proposed openstack/nova master: HyperV: Perform proper cleanup after failed instance spawns https://review.openstack.org/499690 | |
| 18:50:56 | artom | So for PCI devices it can use addresses, for other stuff, or PCI devices that don't have addresses for whatever reason, it can use a different mechanism | |
| 18:51:15 | efried | Because the existing PCI device management subsystem expects every device it gets from the virt driver to have a PCI address in domain:bus:slot.func format. | |
| 18:51:22 | artom | efried, ah, it's a problem because we can't manage them with the current way of doing things? | |
| 18:51:37 | dansmith | for devices we can generalize (like a GPU or NIC) we can talk about what we want in terms of quantity and traits | |
| 18:51:56 | dansmith | we can kindof ask for that in the flavor today, | |
| 18:52:14 | dansmith | but we're not exposing things like GPUs from things like libvirt yet, | |
| 18:52:18 | dansmith | which needs to happen | |
| 18:53:18 | efried | "ask for that in the flavor today" - are you talking about pci_passthrough:alias or something else? | |
| 18:53:23 | dansmith | no, | |
| 18:53:34 | dansmith | in the flavor, you can do things like resources:VGPU=1 | |
| 18:53:47 | dansmith | and placement will not consider hosts that don't expose a VGPU resource type that has at least one available | |
| 18:53:54 | dansmith | nothing does right now, but you can ask for it | |
| 18:55:00 | efried | heh, okay. But we don't have any accomodation for resources:PCI_DEVICE=1,traits=whatever ? | |
| 18:55:19 | dansmith | well, that doesn't make sense | |
| 18:55:26 | dansmith | you wouldn't ask for one pci device, any pci device | |
| 18:55:42 | dansmith | like "give me a nic, or a sata controller, or a serial UART, I'll take anything!" | |
| 18:55:58 | efried | No, I agree that's too broad. | |
| 18:56:09 | artom | Maybe they're working on the pci kernel subsystem ;) | |
| 18:56:24 | dansmith | but you might ask for a NIC or a NIC_PF or a NIC_VF I thnk | |
| 18:56:41 | efried | And I *think* I see that we can accomodate broad groupings | |
| 18:56:58 | efried | But I'm concerned that we can't accomodate narrow groupings. | |
| 18:57:11 | efried | In particular, what if I want a specific device? | |
| 18:57:13 | dansmith | jay might argue that we should ask for a NIC=1, traits=definiteily_a_pf | |
| 18:57:30 | artom | efried, define 'specific device'? | |
| 18:57:44 | dansmith | efried: like "give me device 00:01:5f on host $foo" ? | |
| 18:57:44 | efried | "device with UUID X" | |
| 18:57:49 | artom | Because the 'non-cloud way of doing things' argument will come up really fast with that one ;) | |
| 18:57:51 | dansmith | because that's not a thing we want to do | |
| 18:57:53 | dansmith | yeah | |
| 18:57:56 | efried | heh | |
| 18:57:56 | dansmith | reeeeeal fast :) | |
| 18:58:13 | efried | Well, okay, let me put it a different way. | |
| 18:58:41 | efried | I want my aliasing to be able to create a group of specific devices. | |
| 18:58:56 | openstackgerrit | Merged openstack/nova master: Fix _delete_inventory log message in report client https://review.openstack.org/498833 | |