Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-31
18:40:14 efried via placement?
18:40:40 dansmith it won't ever be *via* placement, but we'll be using placement to do our accounting of *how many* of those things we have available
18:40:49 efried So part of what I'm driving here is total replacement of the existing PCI passthrough setup.
18:40:57 dansmith it won't be materially different from today, except faster, cleaner, and less racy
18:41:15 dansmith no, because the existing pci passthrough stuff knows *which* and placement does not and will not
18:41:37 dansmith in tha' future,
18:41:50 dansmith we'll pick a host with enough GPU_THINGs, reserve one of the eight it provides,
18:42:04 dansmith and then later we get to the compute mostly guaranteed that one will be available, where we'll pick the actual one
18:42:15 dansmith but placement will only ever know that we're using one of the eight
18:42:16 efried Who's that second 'we'?
18:42:24 efried The virt driver?
18:42:38 efried The resource tracker?
18:42:40 dansmith the we before "pick the actual one" ?
18:42:44 efried yeah
18:43:04 dansmith well, today it's super confusing I think, as it's kindof "the scheduler filter" I think
18:43:07 dansmith for certain things
18:43:11 dansmith but regardless,
18:43:37 dansmith that we will be something on the nova side, likely closer to the compute node than not
18:43:38 dansmith clear as mud? :)
18:43:39 dansmith jaypipes may have a better idea of how that will actually shake out long term,
18:43:48 dansmith but regardless, placement will only ever be counting things, not choosing things
18:44:07 efried dansmith Yeah, this is a shared jaypipes dream.
18:44:19 efried Lest you should think it was just efried's pipe dream
18:45:49 efried Okay, so there's gonna be a generic device passthrough management module of some sort that sits in nova and coordinates among placement, scheduler, and virt driver to do whitelisting, claims, allocations
18:46:04 dansmith well, there is that today :)
18:46:24 efried Just for PCI, right?
18:46:29 dansmith (sans the telling-placement-about-it) thing
18:46:55 dansmith ...just for pci.... aren't we talking about pci here?
18:47:23 efried Well, that's one of Jay's bugbears: it should work the same for any device, not just PCI.
18:47:38 dansmith um
18:47:43 dansmith I'm confused then
18:47:51 efried Which intersects with one of my bugbears, which is: not every device has a PCI address.
18:48:13 dansmith GPU devices need not be managed as raw PCI devices for this I think
18:48:35 dansmith maybe that's what you mean
18:48:55 dansmith so I think eventually you should be able to say I want VGPU=1, traits=nvidia,gen5
18:49:06 dansmith and placement will pick you a host that has one of those available
18:49:36 dansmith when you get to the host, that compute node will tell the virt driver (which I think will be generalizing gpus, at least in the case of libvirt), to give it a gpu, which it'll assign to your guest
18:49:41 artom efried, is not every device having a PCI address a problem? My understanding is that the compute reports its resources (devices) to placement
18:49:50 artom How it figures out what resources it has is up to it
18:50:03 dansmith artom: it doesn't report which resources it has, it reports how many of each it has
18:50:32 artom dansmith, ah, right, thanks
18:50:45 efried Right right. So there's the gap. Right now, artom, not having a PCI address *is* a pretty serious problem.
18:50:51 openstackgerrit Lucian Petrut proposed openstack/nova master: HyperV: Perform proper cleanup after failed instance spawns https://review.openstack.org/499690
18:50:56 artom So for PCI devices it can use addresses, for other stuff, or PCI devices that don't have addresses for whatever reason, it can use a different mechanism
18:51:15 efried Because the existing PCI device management subsystem expects every device it gets from the virt driver to have a PCI address in domain:bus:slot.func format.
18:51:22 artom efried, ah, it's a problem because we can't manage them with the current way of doing things?
18:51:37 dansmith for devices we can generalize (like a GPU or NIC) we can talk about what we want in terms of quantity and traits
18:51:56 dansmith we can kindof ask for that in the flavor today,
18:52:14 dansmith but we're not exposing things like GPUs from things like libvirt yet,
18:52:18 dansmith which needs to happen
18:53:18 efried "ask for that in the flavor today" - are you talking about pci_passthrough:alias or something else?
18:53:23 dansmith no,
18:53:34 dansmith in the flavor, you can do things like resources:VGPU=1
18:53:47 dansmith and placement will not consider hosts that don't expose a VGPU resource type that has at least one available
18:53:54 dansmith nothing does right now, but you can ask for it
18:55:00 efried heh, okay. But we don't have any accomodation for resources:PCI_DEVICE=1,traits=whatever ?
18:55:19 dansmith well, that doesn't make sense
18:55:26 dansmith you wouldn't ask for one pci device, any pci device
18:55:42 dansmith like "give me a nic, or a sata controller, or a serial UART, I'll take anything!"
18:55:58 efried No, I agree that's too broad.
18:56:09 artom Maybe they're working on the pci kernel subsystem ;)
18:56:24 dansmith but you might ask for a NIC or a NIC_PF or a NIC_VF I thnk
18:56:41 efried And I *think* I see that we can accomodate broad groupings
18:56:58 efried But I'm concerned that we can't accomodate narrow groupings.
18:57:11 efried In particular, what if I want a specific device?
18:57:13 dansmith jay might argue that we should ask for a NIC=1, traits=definiteily_a_pf
18:57:30 artom efried, define 'specific device'?
18:57:44 efried "device with UUID X"
18:57:44 dansmith efried: like "give me device 00:01:5f on host $foo" ?
18:57:49 artom Because the 'non-cloud way of doing things' argument will come up really fast with that one ;)
18:57:51 dansmith because that's not a thing we want to do
18:57:53 dansmith yeah
18:57:56 dansmith reeeeeal fast :)
18:57:56 efried heh
18:58:13 efried Well, okay, let me put it a different way.
18:58:41 efried I want my aliasing to be able to create a group of specific devices.
18:58:56 openstackgerrit Merged openstack/nova master: Fix _delete_inventory log message in report client https://review.openstack.org/498833
18:58:58 dansmith you mean like we can do with pci today
18:59:04 efried EXCEPT
18:59:07 dansmith where you choose a vendor and model, and we gather up all those things
18:59:11 efried today we can only do it by product/vendor ID.
18:59:23 efried And only a single product/vendor ID pair.
18:59:31 dansmith right, because that's the industry standard
18:59:47 efried nono, this isn't a weirdo platform thing anymore.
18:59:57 efried I have two NICs, exact same vendor/product IDs
19:00:02 efried But they're attached to different networks.
19:00:08 efried I don't want them in the same alias group
19:00:15 dansmith right, so you ask for a nic attached to network foo,
19:00:22 dansmith not "give me the fifth nic from the left"
19:00:37 dansmith NIC=1, traits=network-public
19:00:38 dansmith or something
19:00:51 efried Right, got it.
19:01:10 artom So that would put potentially mutable deployment details into the nova.conf on the compute?
19:01:26 artom That seems... unwise
19:01:31 dansmith artom: well, that's kinda part of the problem with the pci stuff today,
19:01:40 dansmith but the nic attachment thing is managed dynamically today I think
19:01:52 dansmith like we have some way of correlating actual nics to what physical network they're attached to
19:02:01 efried For SR-IOV we do
19:02:05 dansmith right
19:02:21 efried pci_passthrough_devices lists a physical_network on a PF

Earlier   Later