Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-31
19:13:06 efried okay, right, I get that - those are qualitative measures. That's different.
19:13:15 efried A proc either has that feature or it doesn't.
19:13:42 efried I just don't see this model-quantity-as-quality thing scaling.
19:14:11 efried What if I've got to deal with microversions or something.
19:14:26 dansmith um, what?
19:15:10 efried I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,...
19:15:54 dansmith efried: so you want to add semver understandings to a request processor for things?
19:16:04 dansmith is 1.4.3 > 1.3.7?
19:16:22 efried yes it is; and yes, but not limited to semver.
19:16:29 dansmith no it's not
19:16:47 efried 1.4.3 isn't > 1.3.7 ?
19:16:47 dansmith because I need a particular bug fix that was backported to 1.3.7 and 1.4.5
19:16:55 dansmith 1.4.3 will not give me the bug fix, but you would choose it
19:17:24 dansmith so then I need to say >= 1.3.7 || (! <1.4.6)
19:17:26 dansmith or something
19:17:32 dansmith and that gets crazy
19:17:47 dansmith which is why boiling this down to abstract traits is useful
19:18:08 dansmith it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic
19:18:08 artom What real resource would even have something like thing?
19:18:13 dansmith artom: a hypervisor
19:18:18 artom I understand clock speeds of various kinds of processors, but...
19:18:24 dansmith artom: I assume that's where he was going
19:18:36 dansmith and it's a common thing
19:18:49 dansmith people need to be on at least hyperv 2012.4 in order to be able to use some fragile windows driver
19:19:17 dansmith so instead of them needing to know versions, you can expose has-multiqueue-nics or something
19:19:27 dansmith and then if you have an older release with that fix backported, then that's all that matters
19:19:27 efried Yeah; I was trying to come up with an example of something that could have a lot of different quantitative traits where enumerating them becomes very unwieldy.
19:19:35 mriedem bauzas: didn't you say that before 2.29 and the force flag on the evacuate API, that if you specified a host we'd bypass the scheduler, but now we only do it if you specify a host AND pass force=True?
19:19:40 dansmith instead of needing to know the version semantics of the thing you're asking for
19:20:03 edleafe dansmith: didn't we discuss trait comparisons like this a while ago, and concluded that anything beyond simple boolean "has trait" would be a unmanageable tangle of special comparison cases?
19:20:15 dansmith edleafe: yes, precisely
19:20:45 efried Spitballin here, but what I think I would like to see is a way for me to pass some part of the request spec down to the virt driver, who's responsible for understanding it in whatever manner it wants.
19:21:03 mriedem bauzas: nvm, that's correct
19:21:03 dansmith efried: and that is precisely what we do not want you to be able to do
19:21:09 artom That would *definitely* not scale
19:21:10 dansmith we have way too much of that shit today
19:21:14 dansmith yeeeeah.
19:21:18 artom Every boot request would go to every compute?
19:21:22 mriedem if you pass 2.1 and host, force ends up being None in the API and we pass the host to conductor, which then bypasses the scheduler
19:21:34 efried artom ohh, I see.
19:21:49 edleafe Hey! Why don't we allow *extensions*? That'll solve all these problems!
19:22:03 efried Way it's set up today, you can do the whole resource allocation business based on what the compute told you last time it checked in.
19:22:40 efried ...from a centralized spot (placement/scheduler)
19:22:58 dansmith and don't forget multi-hypervisor clouds
19:23:12 dansmith where much of the details aren't known until you get down to the compute, so if you couldn't pick the right one based on some generic properties,
19:23:21 dansmith you get there and libvirt says "wtf is a hyperv 2012.4?"
19:23:28 dansmith that's hugely wasteful
19:24:26 efried So if I need to do something funkily virt-specific, I have to do it at get_inventory time.
19:24:37 efried and setting-up-resource-providers time.
19:24:49 dansmith you have to expose it as a generic count of resources with traits
19:24:56 efried but constrained to ... yeah, what you said.
19:25:08 dansmith because that's a model that the rest of nova can understand
19:25:10 artom There was talk of doing virt driver capabilities as resources (or traits?), I think?
19:25:24 dansmith yep, exactly his hypervisor version query
19:26:31 artom Well, nova virt driver != hypervisor, but yeah, same gist
19:26:41 efried Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait?
19:26:47 dansmith artom: sure, some of both
19:27:11 dansmith efried: no, no trait inheritance plans that I know of,
19:27:37 dansmith but if you request that trait from your compute node, then you won't consider resources under compute nodes that don't
19:27:45 artom efried, I think the nested stuff was thought up for PFs/VFs
19:27:54 efried How do I request a trait for my compute node?
19:28:18 efried I can't... ask for a compute node as a resource
19:28:22 efried can I?
19:28:24 dansmith no
19:28:45 edleafe Flavor extra_specs
19:28:47 dansmith I just mean a trait for whatever the top-level resource provider is that has your custom_aix_foo thing on it
19:30:00 efried What, like a VCPU??
19:30:27 efried Sorry, thought I was getting the picture here, but that last one threw me for a loop.
19:30:30 dansmith efried: it was your example above
19:30:34 dansmith <efried>Which is why we need nested resource providers: so my hypervisor can be an outermost one and tag itself with e.g. "CUSTOM_AIX_CAPABLE", and then everything else underneath that will inherit that trait?
19:30:52 efried It's the hypervisor that's AIX-capable
19:30:54 dansmith I'm just drilling in on your comment about inheritance
19:31:02 efried As in, you can deploy an AIX VM on it.
19:31:04 dansmith right so that would be virt driver capabilities
19:31:28 efried Right, where/how do those get exposed? Is that still somehow in the placement universe?
19:32:10 dansmith so we're asking for a resource provider that has at least 2 VCPUs available, at least 1G of memory, with a trait of CUSTOM_AIX_WHYWOULDIWANNA
19:32:26 dansmith which would be your node with a virt driver that exposes sufficient inventory and traits
19:32:47 artom Wait, do traits apply to resources or resource providers?
19:33:04 dansmith there is no resources, only resource providers :)
19:33:16 dansmith a resource provider has a set of traits
19:33:38 dansmith a request demands or prefers traits, and asks for quantities of resources by name
19:33:46 artom Err
19:33:48 dansmith we find providers that have sufficient inventory and traits
19:33:49 efried okay, right, my host is the resource provider that provides VCPU and MEM_GB and DISK_GB
19:33:52 edleafe artom: inventory == resources
19:33:54 artom "quantities of resource"
19:34:00 efried and the resource provider has traits
19:34:01 artom "there is no resources"
19:34:18 dansmith artom: meaning there is no object called a resource
19:34:22 artom This isn't the Matrix, yo
19:34:33 efried artom I *think* dansmith means that *placement* doesn't know anything about specific resources.
19:34:45 dansmith artom: there are resource providers that have quantities of a given type, but that type isn't a thing that can have a trait, only the provider
19:34:53 dansmith efried: yeah
19:34:54 edleafe artom: no, but I have an inventory with resource_class == "spoon"
19:34:57 efried (Could it have been Ghostbusters, not Matrix, dansmith?)
19:35:21 dansmith I think I was thinking "there is no try, only do" or whatever
19:35:33 artom That's Star Wars
19:35:43 dansmith yeah, I dunno, I'm not a big enough nerd
19:35:46 efried Yup, and we've reached our limit of three nerd movie franchises.
19:36:02 dansmith yeah,

Earlier   Later