Earlier  
Posted Nick Remark
#openstack-nova - 2018-03-27
14:53:36 dansmith efried: ah, I was going to ask if we should do that too, but I expected it would be outside this spec, no?
14:54:32 efried dansmith: Whether it is or not, that should be stated. I could go either way. But slight preference for including it. For a couple of reasons...
14:54:39 dansmith ack
14:55:23 efried dansmith: First, a single microversion introducing multiple-member_of syntax. I like that better than one microversion introducing it for one URI, another for introducing what's effectively the same feature to another URI.
14:55:26 openstackgerrit Dan Smith proposed openstack/nova-specs master: Amend the member_of spec for multiple query sets https://review.openstack.org/555413
14:55:54 efried dansmith: Second, it's actually going to make the code easier to write. Because we currently do the processing of member_of in common code for both; so we can continue to do that.
14:56:00 dansmith efried: sure
14:56:45 sean-k-mooney[m] stephenfin: that wont work als ovs has no idea if there is enough memory on the numa node for the vm also the memory is not allocated until the vhost-user frontend in qemu connects to the vhost-user backend in ovs
14:56:54 efried dansmith: Technically, the Work Items section should be updated...
14:57:15 dansmith efried: for /rps?
14:57:31 efried dansmith: On rereading, it's sufficiently vague to be acceptable as is.
14:57:36 dansmith yeah, I was going to say..
14:57:45 efried dansmith: Which is fine by me; I've always thought that section was pretty much redundant anyway.
14:58:20 efried dansmith: +1. One spec down!
14:58:29 sean-k-mooney[m] stephenfin: the other thing is that when the vhost-user frontend connects to the vhost-user backend in ovs it tries to allocate memory and pmd for the vhost-user interface from on of the numa nodes of the guest automatically
14:58:43 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Handle nested providers for allocation candidates https://review.openstack.org/556873
14:59:50 jaypipes efried: ^
14:59:58 efried jaypipes: ack
15:00:53 mriedem bhagyashris: comments inline https://review.openstack.org/#/c/511825/
15:04:00 jaypipes gibi: maybe I'm just being thick... I still don't get it. You will have multiple backends on the same compute host supporting the same physical networks that support the same vNIC types and you want to be able to choose which backend to allocate a piece of bandwidth from?
15:04:12 bauzas jaypipes: I'm not sure we need a spec for https://review.openstack.org/#/c/556873/1/specs/rocky/approved/nested-resource-providers-allocation-candidates.rst
15:04:24 bauzas jaypipes: it's just fixes we need to merge IMHO
15:05:01 mriedem bauzas: it's an api change and microversion right?
15:05:03 mriedem so spec is required yeah
15:05:05 mriedem ?
15:05:17 gibi jaypipes: exactly. Neutron today does it in the following way:
15:05:28 gibi jaypipes: Neutron has a mechnism driver config
15:05:32 bauzas mriedem: from the spec itself, looks like it's not changing the API
15:05:37 stephenfin sean-k-mooney[m]: Potentially dumb question, but in what way would lack of guest memory be related?
15:05:47 gibi jaypipes: Neutron iterates throught that list and try binding the port with the given driver
15:05:48 jaypipes mriedem: no API change, no.
15:06:05 stephenfin I figured if you could do 'ovs-appctl dpif-netdev/pmd-rxq-show' to show the affinity for a given interface and just return that
15:06:12 gibi jaypipes: the first driver that returns a positive result from that bind call will be the one Neutron use
15:06:30 gibi jaypipes: so that config option defines a preference order between backends supporting the same physnet
15:06:34 mriedem jaypipes: bauzas: it's a behavior change for the alloc candidates api though
15:06:35 stephenfin Or does that even exist before the port is attached to the interface?
15:06:45 bhagyashris mriedem: thank you for review i will look into it as i am working in IST time zone it's end of day for me :)
15:06:53 stephenfin this would be so much easier if I have a machine to experiment on. Stupid fried motherboard is killing me :(
15:07:07 jaypipes mriedem: if "behaviour change" means "it will work when there are nested providers", then yes. :)
15:07:12 bauzas mriedem: technically, alloc-candidates doesn't work yet with nested RPs
15:07:19 jaypipes stephenfin: at least it's not an efried motherboard.
15:07:25 bauzas mriedem: it's not changing the existing
15:07:35 mriedem bauzas: technically volume-backed rebuild with a new image doesn't work either,
15:07:40 mriedem but if we make that work, it's an api change
15:07:51 mriedem even if the request params don't change
15:08:01 jaypipes does it even matter guys? the spec is up.
15:08:01 bauzas I agree it's a signal
15:08:12 mriedem yeah i'm saying it's a spec,
15:08:15 mriedem and likely a microversion bump
15:08:26 bauzas okay, I'll comment that too then
15:09:47 bauzas mriedem: jaypipes: speaking of specs
15:10:03 bauzas jaypipes: now I'm back, can we discuss about my point with vGPU types ?
15:10:17 bauzas looks like you were sad about that
15:11:05 sahid stephenfin, sean-k-mooney[m] I think the issue is on the scheduling, we will have first to select a host so then we could create the ports and query them
15:11:28 stephenfin sahid: I'd be perfectly fine saying we need pre-created ports for this thing
15:11:54 stephenfin That's already a requirement for SR-IOV and afaik is the plan for gibi's bandwidth-aware scheduling spec
15:12:18 sahid stephenfin: the problem is then you need to raise that re-schedule excpetion is the resources are not enough
15:12:27 jaypipes bauzas: "sad" would be one way to put it, yes.
15:12:32 sahid yes i think we do that somewhere but i can't really remember
15:12:47 stephenfin same issue with SR-IOV though, right?
15:12:47 bauzas jaypipes: let's be gentlemen :p
15:13:08 bauzas jaypipes: so, before discussing about a solution, do you understand the problem ?
15:13:18 sahid stephenfin: yes probably you are right
15:13:43 jaypipes bauzas: yes, I fully understand the problem.
15:13:57 bauzas jaypipes: in https://specs.openstack.org/openstack/nova-specs/specs/queens/implemented/add-support-for-vgpu.html we don't mention that a single pGPU can have multiple types and we just suppose a single inventory of VGPU for each
15:14:47 jaypipes bauzas: I had numerous conversations with jianghuaw_ about this.
15:14:48 bauzas the fact is, once *one* mediated device is created, then *all* the others types are not possible for that pGPU
15:15:03 stephenfin sahid, sean-k-mooney[m]: The bigger concern I have is that a vhost-user port is configured to a given PMD based on the guest - not the routes
15:15:04 bauzas jaypipes: and what are your thoughts on that ?
15:15:52 bauzas jaypipes: we already somehow set inventories based on config option thru enabled_gpu_types tho
15:16:00 jaypipes bauzas: the only thing I really don't want is on-the-fly re-configuration of providers based on *what the user requested*.
15:16:15 stephenfin sahid, sean-k-mooney[m]: e.g. all vhost-user ports will be handled by a random PMD until the guest is attached, when that reallocation happens
15:16:17 bauzas jaypipes: it's not that, if I understand correctly your concern
15:16:45 stephenfin sahid, sean-k-mooney[m]: Assuming vhost-user ports that aren't associated with a guest even appear in output of 'ovs-appctl dpif-netdev/pmd-rxq-show' (I can't test it, grrr)
15:17:12 jaypipes bauzas: if we want to pre-define configuration of multiple supported vGPU types using a CONF option, so be it. I would prefer to stop adding yet more CONF options and instead handle inventory of providers using a provider-config YAML file format, but that ain't gonna happen apparently, so be it.
15:17:33 bauzas jaypipes: if we go on a direction where each pGPU has multiple inventories, each for a GPU type, then that said, yes it would be dynamically modified on an instance creation
15:17:48 kashyap alex_xu_: You're right; I can remove that extra test. The main test in test_driver.py takes care of the full config. Thanks for catching.
15:17:52 bauzas jaypipes: I can try to spec it, you know
15:18:05 jaypipes bauzas: that's exactly what I *don't* want.
15:18:05 bauzas jaypipes: the YAML file
15:18:15 bauzas jaypipes: okay cool, so we're aligned
15:18:18 edleafe bauzas: multiple inventories is not going to work
15:18:31 bauzas edleafe: if each inventory is behind a child RP
15:18:38 bauzas either way, sounds we're in agreement
15:18:42 edleafe yeah
15:18:51 bauzas I *don't* want to make things complicated
15:19:03 bauzas the problem is more about the specific config option
15:19:08 bauzas I understand jaypipes on that
15:19:15 jaypipes bauzas: it would be the same inventory of VGPU. But different child providers would be tagged with specific VGPU_TYPE_XXX traits, right?
15:19:37 sean-k-mooney[m] stephenfin: if they are not associate with a guest im not sure. when they do get added to a guest they might also change.
15:19:41 bauzas jaypipes: you're talking of the possibility to have children, each of them being a type ?
15:19:56 bauzas jaypipes: if so, each inventory would be different
15:20:06 jaypipes bauzas: no
15:20:10 bauzas because the total number of vGPUs you can create depends on your type
15:20:20 jaypipes bauzas: I'm saying the resource class would all be "VGPU"
15:20:21 bauzas oh, with the conf opt solution ?
15:20:34 stephenfin sean-k-mooney[m]: Yeah, it seems like a lot of magic (even for ovs-dpdk) to say "we have this route that uses this NIC, therefore the vhost-user port should be processed by this PMD thread"
15:20:39 bauzas yeah, it's still VGPU resource class and a trait for that type
15:20:49 jaypipes bauzas: right.
15:20:58 bauzas so each PGPU will be a leaf

Earlier   Later