| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-31 | |||
| 19:02:21 | efried | pci_passthrough_devices lists a physical_network on a PF | |
| 19:02:28 | dansmith | yeah | |
| 19:02:38 | efried | But that's a single, specific attr | |
| 19:03:10 | dansmith | artom: so yeah if someone goes and moves the yellow cable to the blue port in the datacenter, then bad things happen, but that's kinda the risk you run with giving people hardware passthrough It hink | |
| 19:03:29 | efried | So with generic device management, you could have arbitrary traits, one of which is physical_network | |
| 19:03:44 | artom | Yeah, and now that I think about it more, there's no way to avoid it really | |
| 19:03:51 | artom | No sane way, at any rate | |
| 19:04:06 | dansmith | efried: yeah like maybe you'd want has-tls-offload on your nic as well as physnet-private37 | |
| 19:04:11 | artom | Unless you want to start having admin REST APIs where you associate device addresses with what network they're connected to | |
| 19:04:41 | efried | So qualitative attributes are handle via traits; and in the example of physical_network, I sorta have to have the key and value both stuffed into the value? | |
| 19:05:02 | efried | Like, CUSTOM_NETWORK_MYNET would be a trait? | |
| 19:05:03 | dansmith | efried: well, you have to arrange things as traits, not k=v | |
| 19:05:06 | dansmith | yes | |
| 19:05:18 | efried | And I'd have to have a separate one for every network. | |
| 19:05:29 | efried | Okay; so how are qualitative attrs handled? | |
| 19:05:42 | efried | Like, "give me a VGPU >= 10GHz" | |
| 19:05:52 | dansmith | we don't do that | |
| 19:05:55 | artom | Actually, in that situation, wouldn't the physical network become a resource, not a trait? | |
| 19:06:01 | efried | sorry ^ quantitative | |
| 19:06:10 | dansmith | efried: we don't do that :) | |
| 19:06:28 | efried | Are we gonna? | |
| 19:06:50 | dansmith | every time someone brings it up, jay has an aneurism | |
| 19:06:58 | efried | Really? Cause he's the one who brought it up. | |
| 19:06:59 | dansmith | it's too much complexity I think | |
| 19:07:02 | dansmith | we have to model something | |
| 19:07:37 | dansmith | so let's say you had some 5GHz and some 10GHz gpus | |
| 19:07:56 | efried | Here's a jaypipes quote: "With the advent of the resource provider modeling, which includes an appropriate modeling of quantitative and qualitative things split out into resources (with inventory and allocation) and traits, we have a system that can (and IMHO should) be the basis for any future improvements to device management in Nova." | |
| 19:07:58 | dansmith | you could have traits of gpu-5ghz on the 5ghz ones and gpu-5ghz and gpu-10ghz on the 10s | |
| 19:08:14 | dansmith | efried: yep, totes | |
| 19:08:26 | efried | dansmith So what did he mean by "quantitative"? | |
| 19:08:34 | dansmith | efried: counts of resources | |
| 19:08:39 | efried | oh. | |
| 19:08:40 | dansmith | efried: counts of actual things | |
| 19:08:41 | efried | Bugger. | |
| 19:08:56 | dansmith | so my model above | |
| 19:09:09 | dansmith | if you care about the fast one, you ask for a gpu-10ghz and you would weed out the 5s | |
| 19:09:25 | efried | But if you want to set a minimum, you'll never get the 10s | |
| 19:09:37 | dansmith | if you want at least 5, you would ask for that, and might get a 10 if that's how it falls | |
| 19:09:45 | dansmith | I don't get that last statement | |
| 19:09:50 | efried | Huh? How would that work? | |
| 19:10:00 | dansmith | so, | |
| 19:10:19 | dansmith | the 10g gpus has at-least-5g and at-leadt-10g traits attached | |
| 19:10:19 | efried | Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz" | |
| 19:10:23 | efried | It doesn't even know they're related | |
| 19:10:25 | dansmith | right, | |
| 19:10:30 | dansmith | which is why you put both traits on the 10s | |
| 19:10:33 | efried | ohjeez | |
| 19:10:43 | dansmith | and arrange your grammar in terms of traits | |
| 19:11:05 | efried | okay, I reread your original statement, which I misread the first time. | |
| 19:11:47 | artom | We need nested traits ;) | |
| 19:11:51 | dansmith | the reason for doing this is that it's easy to model to everything | |
| 19:11:57 | efried | "easy" | |
| 19:12:05 | dansmith | compare this which is numerical to processor flags | |
| 19:12:15 | dansmith | I can't say "give me a cpu with at least sse3" | |
| 19:12:36 | efried | I don't know what that means | |
| 19:12:37 | dansmith | because there were later procs without those things because they weren't the high-end one, but were faster, and had some other features that earlier ones didn't | |
| 19:12:53 | dansmith | i.e. processor flags are not numerical, linear, and neat sets | |
| 19:13:06 | efried | okay, right, I get that - those are qualitative measures. That's different. | |
| 19:13:15 | efried | A proc either has that feature or it doesn't. | |
| 19:13:42 | efried | I just don't see this model-quantity-as-quality thing scaling. | |
| 19:14:11 | efried | What if I've got to deal with microversions or something. | |
| 19:14:26 | dansmith | um, what? | |
| 19:15:10 | efried | I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,... | |
| 19:15:54 | dansmith | efried: so you want to add semver understandings to a request processor for things? | |
| 19:16:04 | dansmith | is 1.4.3 > 1.3.7? | |
| 19:16:22 | efried | yes it is; and yes, but not limited to semver. | |
| 19:16:29 | dansmith | no it's not | |
| 19:16:47 | efried | 1.4.3 isn't > 1.3.7 ? | |
| 19:16:47 | dansmith | because I need a particular bug fix that was backported to 1.3.7 and 1.4.5 | |
| 19:16:55 | dansmith | 1.4.3 will not give me the bug fix, but you would choose it | |
| 19:17:24 | dansmith | so then I need to say >= 1.3.7 || (! <1.4.6) | |
| 19:17:26 | dansmith | or something | |
| 19:17:32 | dansmith | and that gets crazy | |
| 19:17:47 | dansmith | which is why boiling this down to abstract traits is useful | |
| 19:18:08 | dansmith | it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic | |
| 19:18:08 | artom | What real resource would even have something like thing? | |
| 19:18:13 | dansmith | artom: a hypervisor | |
| 19:18:18 | artom | I understand clock speeds of various kinds of processors, but... | |
| 19:18:24 | dansmith | artom: I assume that's where he was going | |
| 19:18:36 | dansmith | and it's a common thing | |
| 19:18:49 | dansmith | people need to be on at least hyperv 2012.4 in order to be able to use some fragile windows driver | |
| 19:19:17 | dansmith | so instead of them needing to know versions, you can expose has-multiqueue-nics or something | |
| 19:19:27 | dansmith | and then if you have an older release with that fix backported, then that's all that matters | |
| 19:19:27 | efried | Yeah; I was trying to come up with an example of something that could have a lot of different quantitative traits where enumerating them becomes very unwieldy. | |
| 19:19:35 | mriedem | bauzas: didn't you say that before 2.29 and the force flag on the evacuate API, that if you specified a host we'd bypass the scheduler, but now we only do it if you specify a host AND pass force=True? | |
| 19:19:40 | dansmith | instead of needing to know the version semantics of the thing you're asking for | |
| 19:20:03 | edleafe | dansmith: didn't we discuss trait comparisons like this a while ago, and concluded that anything beyond simple boolean "has trait" would be a unmanageable tangle of special comparison cases? | |
| 19:20:15 | dansmith | edleafe: yes, precisely | |
| 19:20:45 | efried | Spitballin here, but what I think I would like to see is a way for me to pass some part of the request spec down to the virt driver, who's responsible for understanding it in whatever manner it wants. | |
| 19:21:03 | mriedem | bauzas: nvm, that's correct | |
| 19:21:03 | dansmith | efried: and that is precisely what we do not want you to be able to do | |
| 19:21:09 | artom | That would *definitely* not scale | |
| 19:21:10 | dansmith | we have way too much of that shit today | |
| 19:21:14 | dansmith | yeeeeah. | |
| 19:21:18 | artom | Every boot request would go to every compute? | |
| 19:21:22 | mriedem | if you pass 2.1 and host, force ends up being None in the API and we pass the host to conductor, which then bypasses the scheduler | |
| 19:21:34 | efried | artom ohh, I see. | |
| 19:21:49 | edleafe | Hey! Why don't we allow *extensions*? That'll solve all these problems! | |
| 19:22:03 | efried | Way it's set up today, you can do the whole resource allocation business based on what the compute told you last time it checked in. | |
| 19:22:40 | efried | ...from a centralized spot (placement/scheduler) | |