| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-31 | |||
| 19:09:50 | efried | Huh? How would that work? | |
| 19:10:00 | dansmith | so, | |
| 19:10:19 | dansmith | the 10g gpus has at-least-5g and at-leadt-10g traits attached | |
| 19:10:19 | efried | Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz" | |
| 19:10:23 | efried | It doesn't even know they're related | |
| 19:10:25 | dansmith | right, | |
| 19:10:30 | dansmith | which is why you put both traits on the 10s | |
| 19:10:33 | efried | ohjeez | |
| 19:10:43 | dansmith | and arrange your grammar in terms of traits | |
| 19:11:05 | efried | okay, I reread your original statement, which I misread the first time. | |
| 19:11:47 | artom | We need nested traits ;) | |
| 19:11:51 | dansmith | the reason for doing this is that it's easy to model to everything | |
| 19:11:57 | efried | "easy" | |
| 19:12:05 | dansmith | compare this which is numerical to processor flags | |
| 19:12:15 | dansmith | I can't say "give me a cpu with at least sse3" | |
| 19:12:36 | efried | I don't know what that means | |
| 19:12:37 | dansmith | because there were later procs without those things because they weren't the high-end one, but were faster, and had some other features that earlier ones didn't | |
| 19:12:53 | dansmith | i.e. processor flags are not numerical, linear, and neat sets | |
| 19:13:06 | efried | okay, right, I get that - those are qualitative measures. That's different. | |
| 19:13:15 | efried | A proc either has that feature or it doesn't. | |
| 19:13:42 | efried | I just don't see this model-quantity-as-quality thing scaling. | |
| 19:14:11 | efried | What if I've got to deal with microversions or something. | |
| 19:14:26 | dansmith | um, what? | |
| 19:15:10 | efried | I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,... | |
| 19:15:54 | dansmith | efried: so you want to add semver understandings to a request processor for things? | |
| 19:16:04 | dansmith | is 1.4.3 > 1.3.7? | |
| 19:16:22 | efried | yes it is; and yes, but not limited to semver. | |
| 19:16:29 | dansmith | no it's not | |
| 19:16:47 | efried | 1.4.3 isn't > 1.3.7 ? | |
| 19:16:47 | dansmith | because I need a particular bug fix that was backported to 1.3.7 and 1.4.5 | |
| 19:16:55 | dansmith | 1.4.3 will not give me the bug fix, but you would choose it | |
| 19:17:24 | dansmith | so then I need to say >= 1.3.7 || (! <1.4.6) | |
| 19:17:26 | dansmith | or something | |
| 19:17:32 | dansmith | and that gets crazy | |
| 19:17:47 | dansmith | which is why boiling this down to abstract traits is useful | |
| 19:18:08 | dansmith | it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic | |
| 19:18:08 | artom | What real resource would even have something like thing? | |
| 19:18:13 | dansmith | artom: a hypervisor | |
| 19:18:18 | artom | I understand clock speeds of various kinds of processors, but... | |
| 19:18:24 | dansmith | artom: I assume that's where he was going | |
| 19:18:36 | dansmith | and it's a common thing | |
| 19:18:49 | dansmith | people need to be on at least hyperv 2012.4 in order to be able to use some fragile windows driver | |
| 19:19:17 | dansmith | so instead of them needing to know versions, you can expose has-multiqueue-nics or something | |
| 19:19:27 | dansmith | and then if you have an older release with that fix backported, then that's all that matters | |
| 19:19:27 | efried | Yeah; I was trying to come up with an example of something that could have a lot of different quantitative traits where enumerating them becomes very unwieldy. | |
| 19:19:35 | mriedem | bauzas: didn't you say that before 2.29 and the force flag on the evacuate API, that if you specified a host we'd bypass the scheduler, but now we only do it if you specify a host AND pass force=True? | |
| 19:19:40 | dansmith | instead of needing to know the version semantics of the thing you're asking for | |
| 19:20:03 | edleafe | dansmith: didn't we discuss trait comparisons like this a while ago, and concluded that anything beyond simple boolean "has trait" would be a unmanageable tangle of special comparison cases? | |
| 19:20:15 | dansmith | edleafe: yes, precisely | |
| 19:20:45 | efried | Spitballin here, but what I think I would like to see is a way for me to pass some part of the request spec down to the virt driver, who's responsible for understanding it in whatever manner it wants. | |
| 19:21:03 | mriedem | bauzas: nvm, that's correct | |
| 19:21:03 | dansmith | efried: and that is precisely what we do not want you to be able to do | |
| 19:21:09 | artom | That would *definitely* not scale | |
| 19:21:10 | dansmith | we have way too much of that shit today | |
| 19:21:14 | dansmith | yeeeeah. | |
| 19:21:18 | artom | Every boot request would go to every compute? | |
| 19:21:22 | mriedem | if you pass 2.1 and host, force ends up being None in the API and we pass the host to conductor, which then bypasses the scheduler | |
| 19:21:34 | efried | artom ohh, I see. | |
| 19:21:49 | edleafe | Hey! Why don't we allow *extensions*? That'll solve all these problems! | |
| 19:22:03 | efried | Way it's set up today, you can do the whole resource allocation business based on what the compute told you last time it checked in. | |
| 19:22:40 | efried | ...from a centralized spot (placement/scheduler) | |
| 19:22:58 | dansmith | and don't forget multi-hypervisor clouds | |
| 19:23:12 | dansmith | where much of the details aren't known until you get down to the compute, so if you couldn't pick the right one based on some generic properties, | |
| 19:23:21 | dansmith | you get there and libvirt says "wtf is a hyperv 2012.4?" | |
| 19:23:28 | dansmith | that's hugely wasteful | |
| 19:24:26 | efried | So if I need to do something funkily virt-specific, I have to do it at get_inventory time. | |
| 19:24:37 | efried | and setting-up-resource-providers time. | |
| 19:24:49 | dansmith | you have to expose it as a generic count of resources with traits | |
| 19:24:56 | efried | but constrained to ... yeah, what you said. | |
| 19:25:08 | dansmith | because that's a model that the rest of nova can understand | |
| 19:25:10 | artom | There was talk of doing virt driver capabilities as resources (or traits?), I think? | |
| 19:25:24 | dansmith | yep, exactly his hypervisor version query | |
| 19:26:31 | artom | Well, nova virt driver != hypervisor, but yeah, same gist | |
| 19:26:41 | efried | Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait? | |
| 19:26:47 | dansmith | artom: sure, some of both | |
| 19:27:11 | dansmith | efried: no, no trait inheritance plans that I know of, | |
| 19:27:37 | dansmith | but if you request that trait from your compute node, then you won't consider resources under compute nodes that don't | |
| 19:27:45 | artom | efried, I think the nested stuff was thought up for PFs/VFs | |
| 19:27:54 | efried | How do I request a trait for my compute node? | |
| 19:28:18 | efried | I can't... ask for a compute node as a resource | |
| 19:28:22 | efried | can I? | |
| 19:28:24 | dansmith | no | |
| 19:28:45 | edleafe | Flavor extra_specs | |
| 19:28:47 | dansmith | I just mean a trait for whatever the top-level resource provider is that has your custom_aix_foo thing on it | |
| 19:30:00 | efried | What, like a VCPU?? | |
| 19:30:27 | efried | Sorry, thought I was getting the picture here, but that last one threw me for a loop. | |
| 19:30:30 | dansmith | efried: it was your example above | |
| 19:30:34 | dansmith | <efried>Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait? | |
| 19:30:52 | efried | It's the hypervisor that's AIX-capable | |
| 19:30:54 | dansmith | I'm just drilling in on your comment about inheritance | |
| 19:31:02 | efried | As in, you can deploy an AIX VM on it. | |
| 19:31:04 | dansmith | right so that would be virt driver capabilities | |
| 19:31:28 | efried | Right, where/how do those get exposed? Is that still somehow in the placement universe? | |
| 19:32:10 | dansmith | so we're asking for a resource provider that has at least 2 VCPUs available, at least 1G of memory, with a trait of CUSTOM_AIX_WHYWOULDIWANNA | |
| 19:32:26 | dansmith | which would be your node with a virt driver that exposes sufficient inventory and traits | |
| 19:32:47 | artom | Wait, do traits apply to resources or resource providers? | |
| 19:33:04 | dansmith | there is no resources, only resource providers :) | |
| 19:33:16 | dansmith | a resource provider has a set of traits | |
| 19:33:38 | dansmith | a request demands or prefers traits, and asks for quantities of resources by name | |
| 19:33:46 | artom | Err | |