Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-31
19:08:14 dansmith efried: yep, totes
19:08:26 efried dansmith So what did he mean by "quantitative"?
19:08:34 dansmith efried: counts of resources
19:08:39 efried oh.
19:08:40 dansmith efried: counts of actual things
19:08:41 efried Bugger.
19:08:56 dansmith so my model above
19:09:09 dansmith if you care about the fast one, you ask for a gpu-10ghz and you would weed out the 5s
19:09:25 efried But if you want to set a minimum, you'll never get the 10s
19:09:37 dansmith if you want at least 5, you would ask for that, and might get a 10 if that's how it falls
19:09:45 dansmith I don't get that last statement
19:09:50 efried Huh? How would that work?
19:10:00 dansmith so,
19:10:19 dansmith the 10g gpus has at-least-5g and at-leadt-10g traits attached
19:10:19 efried Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz"
19:10:23 efried It doesn't even know they're related
19:10:25 dansmith right,
19:10:30 dansmith which is why you put both traits on the 10s
19:10:33 efried ohjeez
19:10:43 dansmith and arrange your grammar in terms of traits
19:11:05 efried okay, I reread your original statement, which I misread the first time.
19:11:47 artom We need nested traits ;)
19:11:51 dansmith the reason for doing this is that it's easy to model to everything
19:11:57 efried "easy"
19:12:05 dansmith compare this which is numerical to processor flags
19:12:15 dansmith I can't say "give me a cpu with at least sse3"
19:12:36 efried I don't know what that means
19:12:37 dansmith because there were later procs without those things because they weren't the high-end one, but were faster, and had some other features that earlier ones didn't
19:12:53 dansmith i.e. processor flags are not numerical, linear, and neat sets
19:13:06 efried okay, right, I get that - those are qualitative measures. That's different.
19:13:15 efried A proc either has that feature or it doesn't.
19:13:42 efried I just don't see this model-quantity-as-quality thing scaling.
19:14:11 efried What if I've got to deal with microversions or something.
19:14:26 dansmith um, what?
19:15:10 efried I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,...
19:15:54 dansmith efried: so you want to add semver understandings to a request processor for things?
19:16:04 dansmith is 1.4.3 > 1.3.7?
19:16:22 efried yes it is; and yes, but not limited to semver.
19:16:29 dansmith no it's not
19:16:47 efried 1.4.3 isn't > 1.3.7 ?
19:16:47 dansmith because I need a particular bug fix that was backported to 1.3.7 and 1.4.5
19:16:55 dansmith 1.4.3 will not give me the bug fix, but you would choose it
19:17:24 dansmith so then I need to say >= 1.3.7 || (! <1.4.6)
19:17:26 dansmith or something
19:17:32 dansmith and that gets crazy
19:17:47 dansmith which is why boiling this down to abstract traits is useful
19:18:08 dansmith it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic
19:18:08 artom What real resource would even have something like thing?
19:18:13 dansmith artom: a hypervisor
19:18:18 artom I understand clock speeds of various kinds of processors, but...
19:18:24 dansmith artom: I assume that's where he was going
19:18:36 dansmith and it's a common thing
19:18:49 dansmith people need to be on at least hyperv 2012.4 in order to be able to use some fragile windows driver
19:19:17 dansmith so instead of them needing to know versions, you can expose has-multiqueue-nics or something
19:19:27 dansmith and then if you have an older release with that fix backported, then that's all that matters
19:19:27 efried Yeah; I was trying to come up with an example of something that could have a lot of different quantitative traits where enumerating them becomes very unwieldy.
19:19:35 mriedem bauzas: didn't you say that before 2.29 and the force flag on the evacuate API, that if you specified a host we'd bypass the scheduler, but now we only do it if you specify a host AND pass force=True?
19:19:40 dansmith instead of needing to know the version semantics of the thing you're asking for
19:20:03 edleafe dansmith: didn't we discuss trait comparisons like this a while ago, and concluded that anything beyond simple boolean "has trait" would be a unmanageable tangle of special comparison cases?
19:20:15 dansmith edleafe: yes, precisely
19:20:45 efried Spitballin here, but what I think I would like to see is a way for me to pass some part of the request spec down to the virt driver, who's responsible for understanding it in whatever manner it wants.
19:21:03 mriedem bauzas: nvm, that's correct
19:21:03 dansmith efried: and that is precisely what we do not want you to be able to do
19:21:09 artom That would *definitely* not scale
19:21:10 dansmith we have way too much of that shit today
19:21:14 dansmith yeeeeah.
19:21:18 artom Every boot request would go to every compute?
19:21:22 mriedem if you pass 2.1 and host, force ends up being None in the API and we pass the host to conductor, which then bypasses the scheduler
19:21:34 efried artom ohh, I see.
19:21:49 edleafe Hey! Why don't we allow *extensions*? That'll solve all these problems!
19:22:03 efried Way it's set up today, you can do the whole resource allocation business based on what the compute told you last time it checked in.
19:22:40 efried ...from a centralized spot (placement/scheduler)
19:22:58 dansmith and don't forget multi-hypervisor clouds
19:23:12 dansmith where much of the details aren't known until you get down to the compute, so if you couldn't pick the right one based on some generic properties,
19:23:21 dansmith you get there and libvirt says "wtf is a hyperv 2012.4?"
19:23:28 dansmith that's hugely wasteful
19:24:26 efried So if I need to do something funkily virt-specific, I have to do it at get_inventory time.
19:24:37 efried and setting-up-resource-providers time.
19:24:49 dansmith you have to expose it as a generic count of resources with traits
19:24:56 efried but constrained to ... yeah, what you said.
19:25:08 dansmith because that's a model that the rest of nova can understand
19:25:10 artom There was talk of doing virt driver capabilities as resources (or traits?), I think?
19:25:24 dansmith yep, exactly his hypervisor version query
19:26:31 artom Well, nova virt driver != hypervisor, but yeah, same gist
19:26:41 efried Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait?
19:26:47 dansmith artom: sure, some of both
19:27:11 dansmith efried: no, no trait inheritance plans that I know of,
19:27:37 dansmith but if you request that trait from your compute node, then you won't consider resources under compute nodes that don't
19:27:45 artom efried, I think the nested stuff was thought up for PFs/VFs
19:27:54 efried How do I request a trait for my compute node?
19:28:18 efried I can't... ask for a compute node as a resource
19:28:22 efried can I?
19:28:24 dansmith no
19:28:45 edleafe Flavor extra_specs
19:28:47 dansmith I just mean a trait for whatever the top-level resource provider is that has your custom_aix_foo thing on it
19:30:00 efried What, like a VCPU??
19:30:27 efried Sorry, thought I was getting the picture here, but that last one threw me for a loop.
19:30:30 dansmith efried: it was your example above
19:30:34 dansmith <efried>Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait?
19:30:52 efried It's the hypervisor that's AIX-capable

Earlier   Later