| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-31 | |||
| 18:52:18 | dansmith | which needs to happen | |
| 18:53:18 | efried | "ask for that in the flavor today" - are you talking about pci_passthrough:alias or something else? | |
| 18:53:23 | dansmith | no, | |
| 18:53:34 | dansmith | in the flavor, you can do things like resources:VGPU=1 | |
| 18:53:47 | dansmith | and placement will not consider hosts that don't expose a VGPU resource type that has at least one available | |
| 18:53:54 | dansmith | nothing does right now, but you can ask for it | |
| 18:55:00 | efried | heh, okay. But we don't have any accomodation for resources:PCI_DEVICE=1,traits=whatever ? | |
| 18:55:19 | dansmith | well, that doesn't make sense | |
| 18:55:26 | dansmith | you wouldn't ask for one pci device, any pci device | |
| 18:55:42 | dansmith | like "give me a nic, or a sata controller, or a serial UART, I'll take anything!" | |
| 18:55:58 | efried | No, I agree that's too broad. | |
| 18:56:09 | artom | Maybe they're working on the pci kernel subsystem ;) | |
| 18:56:24 | dansmith | but you might ask for a NIC or a NIC_PF or a NIC_VF I thnk | |
| 18:56:41 | efried | And I *think* I see that we can accomodate broad groupings | |
| 18:56:58 | efried | But I'm concerned that we can't accomodate narrow groupings. | |
| 18:57:11 | efried | In particular, what if I want a specific device? | |
| 18:57:13 | dansmith | jay might argue that we should ask for a NIC=1, traits=definiteily_a_pf | |
| 18:57:30 | artom | efried, define 'specific device'? | |
| 18:57:44 | dansmith | efried: like "give me device 00:01:5f on host $foo" ? | |
| 18:57:44 | efried | "device with UUID X" | |
| 18:57:49 | artom | Because the 'non-cloud way of doing things' argument will come up really fast with that one ;) | |
| 18:57:51 | dansmith | because that's not a thing we want to do | |
| 18:57:53 | dansmith | yeah | |
| 18:57:56 | efried | heh | |
| 18:57:56 | dansmith | reeeeeal fast :) | |
| 18:58:13 | efried | Well, okay, let me put it a different way. | |
| 18:58:41 | efried | I want my aliasing to be able to create a group of specific devices. | |
| 18:58:56 | openstackgerrit | Merged openstack/nova master: Fix _delete_inventory log message in report client https://review.openstack.org/498833 | |
| 18:58:58 | dansmith | you mean like we can do with pci today | |
| 18:59:04 | efried | EXCEPT | |
| 18:59:07 | dansmith | where you choose a vendor and model, and we gather up all those things | |
| 18:59:11 | efried | today we can only do it by product/vendor ID. | |
| 18:59:23 | efried | And only a single product/vendor ID pair. | |
| 18:59:31 | dansmith | right, because that's the industry standard | |
| 18:59:47 | efried | nono, this isn't a weirdo platform thing anymore. | |
| 18:59:57 | efried | I have two NICs, exact same vendor/product IDs | |
| 19:00:02 | efried | But they're attached to different networks. | |
| 19:00:08 | efried | I don't want them in the same alias group | |
| 19:00:15 | dansmith | right, so you ask for a nic attached to network foo, | |
| 19:00:22 | dansmith | not "give me the fifth nic from the left" | |
| 19:00:37 | dansmith | NIC=1, traits=network-public | |
| 19:00:38 | dansmith | or something | |
| 19:00:51 | efried | Right, got it. | |
| 19:01:10 | artom | So that would put potentially mutable deployment details into the nova.conf on the compute? | |
| 19:01:26 | artom | That seems... unwise | |
| 19:01:31 | dansmith | artom: well, that's kinda part of the problem with the pci stuff today, | |
| 19:01:40 | dansmith | but the nic attachment thing is managed dynamically today I think | |
| 19:01:52 | dansmith | like we have some way of correlating actual nics to what physical network they're attached to | |
| 19:02:01 | efried | For SR-IOV we do | |
| 19:02:05 | dansmith | right | |
| 19:02:21 | efried | pci_passthrough_devices lists a physical_network on a PF | |
| 19:02:28 | dansmith | yeah | |
| 19:02:38 | efried | But that's a single, specific attr | |
| 19:03:10 | dansmith | artom: so yeah if someone goes and moves the yellow cable to the blue port in the datacenter, then bad things happen, but that's kinda the risk you run with giving people hardware passthrough It hink | |
| 19:03:29 | efried | So with generic device management, you could have arbitrary traits, one of which is physical_network | |
| 19:03:44 | artom | Yeah, and now that I think about it more, there's no way to avoid it really | |
| 19:03:51 | artom | No sane way, at any rate | |
| 19:04:06 | dansmith | efried: yeah like maybe you'd want has-tls-offload on your nic as well as physnet-private37 | |
| 19:04:11 | artom | Unless you want to start having admin REST APIs where you associate device addresses with what network they're connected to | |
| 19:04:41 | efried | So qualitative attributes are handle via traits; and in the example of physical_network, I sorta have to have the key and value both stuffed into the value? | |
| 19:05:02 | efried | Like, CUSTOM_NETWORK_MYNET would be a trait? | |
| 19:05:03 | dansmith | efried: well, you have to arrange things as traits, not k=v | |
| 19:05:06 | dansmith | yes | |
| 19:05:18 | efried | And I'd have to have a separate one for every network. | |
| 19:05:29 | efried | Okay; so how are qualitative attrs handled? | |
| 19:05:42 | efried | Like, "give me a VGPU >= 10GHz" | |
| 19:05:52 | dansmith | we don't do that | |
| 19:05:55 | artom | Actually, in that situation, wouldn't the physical network become a resource, not a trait? | |
| 19:06:01 | efried | sorry ^ quantitative | |
| 19:06:10 | dansmith | efried: we don't do that :) | |
| 19:06:28 | efried | Are we gonna? | |
| 19:06:50 | dansmith | every time someone brings it up, jay has an aneurism | |
| 19:06:58 | efried | Really? Cause he's the one who brought it up. | |
| 19:06:59 | dansmith | it's too much complexity I think | |
| 19:07:02 | dansmith | we have to model something | |
| 19:07:37 | dansmith | so let's say you had some 5GHz and some 10GHz gpus | |
| 19:07:56 | efried | Here's a jaypipes quote: "With the advent of the resource provider modeling, which includes an appropriate modeling of quantitative and qualitative things split out into resources (with inventory and allocation) and traits, we have a system that can (and IMHO should) be the basis for any future improvements to device management in Nova." | |
| 19:07:58 | dansmith | you could have traits of gpu-5ghz on the 5ghz ones and gpu-5ghz and gpu-10ghz on the 10s | |
| 19:08:14 | dansmith | efried: yep, totes | |
| 19:08:26 | efried | dansmith So what did he mean by "quantitative"? | |
| 19:08:34 | dansmith | efried: counts of resources | |
| 19:08:39 | efried | oh. | |
| 19:08:40 | dansmith | efried: counts of actual things | |
| 19:08:41 | efried | Bugger. | |
| 19:08:56 | dansmith | so my model above | |
| 19:09:09 | dansmith | if you care about the fast one, you ask for a gpu-10ghz and you would weed out the 5s | |
| 19:09:25 | efried | But if you want to set a minimum, you'll never get the 10s | |
| 19:09:37 | dansmith | if you want at least 5, you would ask for that, and might get a 10 if that's how it falls | |
| 19:09:45 | dansmith | I don't get that last statement | |
| 19:09:50 | efried | Huh? How would that work? | |
| 19:10:00 | dansmith | so, | |
| 19:10:19 | efried | Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz" | |
| 19:10:19 | dansmith | the 10g gpus has at-least-5g and at-leadt-10g traits attached | |
| 19:10:23 | efried | It doesn't even know they're related | |
| 19:10:25 | dansmith | right, | |
| 19:10:30 | dansmith | which is why you put both traits on the 10s | |
| 19:10:33 | efried | ohjeez | |
| 19:10:43 | dansmith | and arrange your grammar in terms of traits | |
| 19:11:05 | efried | okay, I reread your original statement, which I misread the first time. | |
| 19:11:47 | artom | We need nested traits ;) | |