Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-31
19:03:10 dansmith artom: so yeah if someone goes and moves the yellow cable to the blue port in the datacenter, then bad things happen, but that's kinda the risk you run with giving people hardware passthrough It hink
19:03:29 efried So with generic device management, you could have arbitrary traits, one of which is physical_network
19:03:44 artom Yeah, and now that I think about it more, there's no way to avoid it really
19:03:51 artom No sane way, at any rate
19:04:06 dansmith efried: yeah like maybe you'd want has-tls-offload on your nic as well as physnet-private37
19:04:11 artom Unless you want to start having admin REST APIs where you associate device addresses with what network they're connected to
19:04:41 efried So qualitative attributes are handle via traits; and in the example of physical_network, I sorta have to have the key and value both stuffed into the value?
19:05:02 efried Like, CUSTOM_NETWORK_MYNET would be a trait?
19:05:03 dansmith efried: well, you have to arrange things as traits, not k=v
19:05:06 dansmith yes
19:05:18 efried And I'd have to have a separate one for every network.
19:05:29 efried Okay; so how are qualitative attrs handled?
19:05:42 efried Like, "give me a VGPU >= 10GHz"
19:05:52 dansmith we don't do that
19:05:55 artom Actually, in that situation, wouldn't the physical network become a resource, not a trait?
19:06:01 efried sorry ^ quantitative
19:06:10 dansmith efried: we don't do that :)
19:06:28 efried Are we gonna?
19:06:50 dansmith every time someone brings it up, jay has an aneurism
19:06:58 efried Really? Cause he's the one who brought it up.
19:06:59 dansmith it's too much complexity I think
19:07:02 dansmith we have to model something
19:07:37 dansmith so let's say you had some 5GHz and some 10GHz gpus
19:07:56 efried Here's a jaypipes quote: "With the advent of the resource provider modeling, which includes an appropriate modeling of quantitative and qualitative things split out into resources (with inventory and allocation) and traits, we have a system that can (and IMHO should) be the basis for any future improvements to device management in Nova."
19:07:58 dansmith you could have traits of gpu-5ghz on the 5ghz ones and gpu-5ghz and gpu-10ghz on the 10s
19:08:14 dansmith efried: yep, totes
19:08:26 efried dansmith So what did he mean by "quantitative"?
19:08:34 dansmith efried: counts of resources
19:08:39 efried oh.
19:08:40 dansmith efried: counts of actual things
19:08:41 efried Bugger.
19:08:56 dansmith so my model above
19:09:09 dansmith if you care about the fast one, you ask for a gpu-10ghz and you would weed out the 5s
19:09:25 efried But if you want to set a minimum, you'll never get the 10s
19:09:37 dansmith if you want at least 5, you would ask for that, and might get a 10 if that's how it falls
19:09:45 dansmith I don't get that last statement
19:09:50 efried Huh? How would that work?
19:10:00 dansmith so,
19:10:19 dansmith the 10g gpus has at-least-5g and at-leadt-10g traits attached
19:10:19 efried Placement doesn't know that "gpu-10ghz" is an okay thing to match to "gpu-5ghz"
19:10:23 efried It doesn't even know they're related
19:10:25 dansmith right,
19:10:30 dansmith which is why you put both traits on the 10s
19:10:33 efried ohjeez
19:10:43 dansmith and arrange your grammar in terms of traits
19:11:05 efried okay, I reread your original statement, which I misread the first time.
19:11:47 artom We need nested traits ;)
19:11:51 dansmith the reason for doing this is that it's easy to model to everything
19:11:57 efried "easy"
19:12:05 dansmith compare this which is numerical to processor flags
19:12:15 dansmith I can't say "give me a cpu with at least sse3"
19:12:36 efried I don't know what that means
19:12:37 dansmith because there were later procs without those things because they weren't the high-end one, but were faster, and had some other features that earlier ones didn't
19:12:53 dansmith i.e. processor flags are not numerical, linear, and neat sets
19:13:06 efried okay, right, I get that - those are qualitative measures. That's different.
19:13:15 efried A proc either has that feature or it doesn't.
19:13:42 efried I just don't see this model-quantity-as-quality thing scaling.
19:14:11 efried What if I've got to deal with microversions or something.
19:14:26 dansmith um, what?
19:15:10 efried I don't want to have to tag my resource provider with version-at-least-1.1,version-at-least-1.2,version-at-least-1.3,version-at-least-1.4,...
19:15:54 dansmith efried: so you want to add semver understandings to a request processor for things?
19:16:04 dansmith is 1.4.3 > 1.3.7?
19:16:22 efried yes it is; and yes, but not limited to semver.
19:16:29 dansmith no it's not
19:16:47 efried 1.4.3 isn't > 1.3.7 ?
19:16:47 dansmith because I need a particular bug fix that was backported to 1.3.7 and 1.4.5
19:16:55 dansmith 1.4.3 will not give me the bug fix, but you would choose it
19:17:24 dansmith so then I need to say >= 1.3.7 || (! <1.4.6)
19:17:26 dansmith or something
19:17:32 dansmith and that gets crazy
19:17:47 dansmith which is why boiling this down to abstract traits is useful
19:18:08 dansmith it won't let you model everything down to the tiniest detail, but it's also something that people can grasp and can be reasonably generic
19:18:08 artom What real resource would even have something like thing?
19:18:13 dansmith artom: a hypervisor
19:18:18 artom I understand clock speeds of various kinds of processors, but...
19:18:24 dansmith artom: I assume that's where he was going
19:18:36 dansmith and it's a common thing
19:18:49 dansmith people need to be on at least hyperv 2012.4 in order to be able to use some fragile windows driver
19:19:17 dansmith so instead of them needing to know versions, you can expose has-multiqueue-nics or something
19:19:27 dansmith and then if you have an older release with that fix backported, then that's all that matters
19:19:27 efried Yeah; I was trying to come up with an example of something that could have a lot of different quantitative traits where enumerating them becomes very unwieldy.
19:19:35 mriedem bauzas: didn't you say that before 2.29 and the force flag on the evacuate API, that if you specified a host we'd bypass the scheduler, but now we only do it if you specify a host AND pass force=True?
19:19:40 dansmith instead of needing to know the version semantics of the thing you're asking for
19:20:03 edleafe dansmith: didn't we discuss trait comparisons like this a while ago, and concluded that anything beyond simple boolean "has trait" would be a unmanageable tangle of special comparison cases?
19:20:15 dansmith edleafe: yes, precisely
19:20:45 efried Spitballin here, but what I think I would like to see is a way for me to pass some part of the request spec down to the virt driver, who's responsible for understanding it in whatever manner it wants.
19:21:03 mriedem bauzas: nvm, that's correct
19:21:03 dansmith efried: and that is precisely what we do not want you to be able to do
19:21:09 artom That would *definitely* not scale
19:21:10 dansmith we have way too much of that shit today
19:21:14 dansmith yeeeeah.
19:21:18 artom Every boot request would go to every compute?
19:21:22 mriedem if you pass 2.1 and host, force ends up being None in the API and we pass the host to conductor, which then bypasses the scheduler
19:21:34 efried artom ohh, I see.
19:21:49 edleafe Hey! Why don't we allow *extensions*? That'll solve all these problems!
19:22:03 efried Way it's set up today, you can do the whole resource allocation business based on what the compute told you last time it checked in.
19:22:40 efried ...from a centralized spot (placement/scheduler)
19:22:58 dansmith and don't forget multi-hypervisor clouds
19:23:12 dansmith where much of the details aren't known until you get down to the compute, so if you couldn't pick the right one based on some generic properties,
19:23:21 dansmith you get there and libvirt says "wtf is a hyperv 2012.4?"

Earlier   Later