| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-31 | |||
| 18:58:41 | efried | I want my aliasing to be able to create a group of specific devices. | |
| 18:58:56 | openstackgerrit | Merged openstack/nova master: Fix _delete_inventory log message in report client https://review.openstack.org/498833 | |
| 18:58:58 | dansmith | you mean like we can do with pci today | |
| 18:59:04 | efried | EXCEPT | |
| 18:59:07 | dansmith | where you choose a vendor and model, and we gather up all those things | |
| 18:59:11 | efried | today we can only do it by product/vendor ID. | |
| 18:59:23 | efried | And only a single product/vendor ID pair. | |
| 18:59:31 | dansmith | right, because that's the industry standard | |
| 18:59:47 | efried | nono, this isn't a weirdo platform thing anymore. | |
| 18:59:57 | efried | I have two NICs, exact same vendor/product IDs | |
| 19:00:02 | efried | But they're attached to different networks. | |
| 19:00:08 | efried | I don't want them in the same alias group | |
| 19:00:15 | dansmith | right, so you ask for a nic attached to network foo, | |
| 19:00:22 | dansmith | not "give me the fifth nic from the left" | |
| 19:00:37 | dansmith | NIC=1, traits=network-public | |
| 19:00:38 | dansmith | or something | |
| 19:00:51 | efried | Right, got it. | |
| 19:01:10 | artom | So that would put potentially mutable deployment details into the nova.conf on the compute? | |
| 19:01:26 | artom | That seems... unwise | |
| 19:01:31 | dansmith | artom: well, that's kinda part of the problem with the pci stuff today, | |
| 19:01:40 | dansmith | but the nic attachment thing is managed dynamically today I think | |
| 19:01:52 | dansmith | like we have some way of correlating actual nics to what physical network they're attached to | |
| 19:02:01 | efried | For SR-IOV we do | |
| 19:02:05 | dansmith | right | |
| 19:02:21 | efried | pci_passthrough_devices lists a physical_network on a PF | |
| 19:02:28 | dansmith | yeah | |
| 19:02:38 | efried | But that's a single, specific attr | |
| 19:03:10 | dansmith | artom: so yeah if someone goes and moves the yellow cable to the blue port in the datacenter, then bad things happen, but that's kinda the risk you run with giving people hardware passthrough It hink | |
| 19:03:29 | efried | So with generic device management, you could have arbitrary traits, one of which is physical_network | |
| 19:03:44 | artom | Yeah, and now that I think about it more, there's no way to avoid it really | |
| 19:03:51 | artom | No sane way, at any rate | |
| 19:04:06 | dansmith | efried: yeah like maybe you'd want has-tls-offload on your nic as well as physnet-private37 | |
| 19:04:11 | artom | Unless you want to start having admin REST APIs where you associate device addresses with what network they're connected to | |
| 19:04:41 | efried | So qualitative attributes are handle via traits; and in the example of physical_network, I sorta have to have the key and value both stuffed into the value? | |
| 19:05:02 | efried | Like, CUSTOM_NETWORK_MYNET would be a trait? | |
| 19:05:03 | dansmith | efried: well, you have to arrange things as traits, not k=v | |
| 19:05:06 | dansmith | yes | |
| 19:05:18 | efried | And I'd have to have a separate one for every network. | |
| 19:05:29 | efried | Okay; so how are qualitative attrs handled? | |
| 19:05:42 | efried | Like, "give me a VGPU >= 10GHz" | |
| 19:05:52 | dansmith | we don't do that | |
| 19:05:55 | artom | Actually, in that situation, wouldn't the physical network become a resource, not a trait? | |
| 19:06:01 | efried | sorry ^ quantitative | |
| 19:06:10 | dansmith | efried: we don't do that :) | |
| 19:06:28 | efried | Are we gonna? | |
| 19:06:50 | dansmith | every time someone brings it up, jay has an aneurism | |
| 19:06:58 | efried | Really? Cause he's the one who brought it up. | |
| 19:06:59 | dansmith | it's too much complexity I think | |
| 19:07:02 | dansmith | we have to model something | |
| 19:07:37 | dansmith | so let's say you had some 5GHz and some 10GHz gpus | |
| 19:07:56 | efried | Here's a jaypipes quote: "With the advent of the resource provider modeling, which includes an appropriate modeling of quantitative and qualitative things split out into resources (with inventory and allocation) and traits, we have a system that can (and IMHO should) be the basis for any future improvements to device management in Nova." | |
| 19:07:58 | dansmith | you could have traits of gpu-5ghz on the 5ghz ones and gpu-5ghz and gpu-10ghz on the 10s | |
| 19:08:14 | dansmith | efried: yep, totes | |
| 19:08:26 | efried | dansmith So what did he mean by "quantitative"? | |
| 19:08:34 | dansmith | efried: counts of resources | |
| 19:08:39 | efried | oh. | |
| 19:08:40 | dansmith | efried: counts of actual things | |
| 19:08:41 | efried | Bugger. | |
| 19:08:56 | dansmith | so my model above | |
| 19:09:09 | dansmith | if you care about the fast one, you ask for a gpu-10ghz and you would weed out the 5s | |
| 19:09:25 | efried | But if you want to set a minimum, you'll never get the 10s | |
| 19:09:37 | dansmith | if you want at least 5, you would ask for that, and might get a 10 if that's how it falls | |
| 19:09:45 | dansmith | I don't get that last statement | |
| 19:09:50 | efried | Huh? How would that work? | |
| 19:10:00 | dansmith | so, | |
| 19:10:19 | dansmith | the 10g gpus has at-least-5g and at-leadt-10g traits attached | |
| 19:10:19 | efried | Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz" | |
| 19:10:23 | efried | It doesn't even know they're related | |
| 19:10:25 | dansmith | right, | |
| 19:10:30 | dansmith | which is why you put both traits on the 10s | |
| 19:10:33 | efried | ohjeez | |
| 19:10:43 | dansmith | and arrange your grammar in terms of traits | |
| 19:11:05 | efried | okay, I reread your original statement, which I misread the first time. | |
| 19:11:47 | artom | We need nested traits ;) | |
| 19:11:51 | dansmith | the reason for doing this is that it's easy to model to everything | |
| 19:11:57 | efried | "easy" | |
| 19:12:05 | dansmith | compare this which is numerical to processor flags | |
| 19:12:15 | dansmith | I can't say "give me a cpu with at least sse3" | |
| 19:12:36 | efried | I don't know what that means | |
| 19:12:37 | dansmith | because there were later procs without those things because they weren't the high-end one, but were faster, and had some other features that earlier ones didn't | |
| 19:12:53 | dansmith | i.e. processor flags are not numerical, linear, and neat sets | |
| 19:13:06 | efried | okay, right, I get that - those are qualitative measures. That's different. | |
| 19:13:15 | efried | A proc either has that feature or it doesn't. | |
| 19:13:42 | efried | I just don't see this model-quantity-as-quality thing scaling. | |
| 19:14:11 | efried | What if I've got to deal with microversions or something. | |
| 19:14:26 | dansmith | um, what? | |
| 19:15:10 | efried | I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,... | |
| 19:15:54 | dansmith | efried: so you want to add semver understandings to a request processor for things? | |
| 19:16:04 | dansmith | is 1.4.3 > 1.3.7? | |
| 19:16:22 | efried | yes it is; and yes, but not limited to semver. | |
| 19:16:29 | dansmith | no it's not | |
| 19:16:47 | efried | 1.4.3 isn't > 1.3.7 ? | |
| 19:16:47 | dansmith | because I need a particular bug fix that was backported to 1.3.7 and 1.4.5 | |
| 19:16:55 | dansmith | 1.4.3 will not give me the bug fix, but you would choose it | |
| 19:17:24 | dansmith | so then I need to say >= 1.3.7 || (! <1.4.6) | |
| 19:17:26 | dansmith | or something | |
| 19:17:32 | dansmith | and that gets crazy | |
| 19:17:47 | dansmith | which is why boiling this down to abstract traits is useful | |
| 19:18:08 | dansmith | it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic | |
| 19:18:08 | artom | What real resource would even have something like thing? | |