Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-25
14:53:56 fried_rice claudiub Well
14:53:57 claudiub hm, haven't tried devname though.
14:54:08 fried_rice claudiub I haven't actually found any reason you can't spoof the vendor/product ID.
14:54:20 fried_rice I don't *think* those are actually checked against anything real, ever.
14:54:54 fried_rice They're just correlated among the passthrough_whitelist (compute conf), the pci_passthrough_devices (compute-provided get_available_resource), and the alias list (api conf).
14:54:57 claudiub the PCI devices reported by the compute drivers have to have those vendor_id / product_ids
14:55:27 fried_rice Yes, have to have them, but I don't think anything cares that those values are really what the dev reports.
14:55:45 claudiub true.
14:56:09 fried_rice I'm... sort of counting on it, actually. Cause I have a vendor ID, but I'm not sure if the thing I'm using as a product ID is really a product ID.
14:56:24 fried_rice https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/host.py@111
14:56:59 claudiub for sr-iov devices, i'm currently doing an md5, if i can't get the actual vendor_id, product_id, which feels wrong. :)
14:57:00 fried_rice Course, if you do that (or really, whatever you do), you have to document how your user needs to glean the right value to put in those spots in the whitelist/alias.
14:57:11 fried_rice claudiub Totally
14:57:16 fried_rice I feel your pain.
14:57:43 fried_rice claudiub I'm doing something similar with PCI addresses: https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/vm.py@884
14:58:05 claudiub *ahem*, if only we could also whitelist via device_id, that would work for me. would it work for you as well?
14:58:25 fried_rice claudiub Where device_id can be an arbitrary string, yes, totally.
14:59:32 fried_rice claudiub In order to do wildcarding, though, you would need to outsource the validation/filtering to the compute driver.
14:59:46 fried_rice Cause he's the only guy who knows how to interpret that "arbitrary" format.
15:00:17 fried_rice Otherwise you would just get a straight string match, which means you have to enumerate every device (or continue to use classes via vendor/prod IDs)
15:02:05 fried_rice claudiub What does the flavor side look like in that case?
15:02:17 claudiub fried_rice: for sr-iov you mean?
15:02:20 fried_rice I don't like the idea of having a flavor that specifies individual devices
15:02:40 claudiub fried_rice: or normal pci devices
15:02:54 fried_rice Let's go normal PCI for now.
15:03:46 claudiub fried_rice: it should still look the same: pci_passthrough:alias=<alias>:<count>
15:04:16 fried_rice so you're adding "device_id" as an alternative to vendor_id/product_id in the alias?
15:04:39 claudiub fried_rice: for sr-iov, the flavor doesn't change, only the neutron port's vnic_type has to bbe direct or something similar, and your ML2 mechanism_driver must be able to support it.
15:04:53 fried_rice Gobackgoback
15:04:59 fried_rice What does that alias look like?
15:05:10 claudiub for sr-iov, you don't need an alias
15:05:32 claudiub you only alias normal pci devices.
15:05:35 fried_rice Yeah, regular PCI devices. Cause if it's just [pci]alias = {"device_id": "1234"} - then you're effectively specifying one device per alias.
15:06:10 fried_rice and you might as well use extra_specs pci_passthrough:device_id:1234 rather than bothering with an alias entry.
15:07:13 stephenfin dansmith: Any chance you could take a look at https://review.openstack.org/#/c/496605/? Think it might be a good backport candidate
15:07:32 stephenfin the follow-up patch probably needs a little more discussion yet
15:07:36 claudiub fried_rice: well, tbh, for normal PCI devices, vendor_id / product_ids are always there on hyper-v. so, that can be used for aliasing. but IMO, a device_id should also be supported.
15:08:09 fried_rice So how would you do that? Perhaps something like [pci]alias = {"name": "foo", "devices": [{"device_id": "abc"}, {"device_id": "123"}, {"vendor_id": "1f2e", "product_id": "a9b8"}]} ?
15:08:10 claudiub fried_rice: yeah, you're right
15:08:31 fried_rice I.e. you can specify a list of stuff to an alias, that allows you to enumerate several devices by ID and/or clasess by prod/vendor?
15:08:35 fried_rice That would be... kinda cool.
15:09:28 fried_rice Even without the ability to specify device IDs in there, if I want to be able to group different prod/vendor types together under a single alias...
15:09:53 fried_rice Like maybe I only care if I get a GPU. Any GPU will do. So make an alias grouping with these vendor/product IDs that all represent GPUs.
15:12:36 dansmith stephenfin: why are you not just doing the conf deprecation now?
15:13:11 stephenfin dansmith: I want to backport the change. Backporting a deprecation seems wrong
15:13:25 stephenfin Same reason I've kept the deprecation timeline so generic
15:13:30 dansmith stephenfin: ah right okay
15:13:49 dansmith stephenfin: why no tests?
15:14:29 dansmith should be pretty easy to write one to make sure we skip that bit...
15:14:33 stephenfin Damn. Because I was cheating :(
15:14:46 kashyap stephenfin: :P
15:15:01 kashyap stephenfin: I noticed the "no tests" thing, but thought you intentionally skipped them with good reason
15:15:49 dansmith stephenfin: I missed the deprecation patch after this, which makes sense..
15:16:23 fried_rice claudiub Would you be interested in co-authoring/sponsoring a bp to allow whitelisting & aliasing by device ID?
15:16:26 stephenfin dansmith: Yup, that one apparently needs a little more discussion. Key mappings are a minefield
15:16:56 claudiub fried_rice: well, tbh, a device_id would work in sr-iov usecase. that device_id would represent the NIC to which the VFs belong to. thus, I'd have N VFs grouped under the same device_id, which can be aliased as well.
15:16:56 stephenfin fried_rice, claudiub: Stick me on the review if you do author such a bp/spec
15:17:16 fried_rice stephenfin ack
15:17:25 claudiub ould work in my sr-iov usecase *
15:17:42 fried_rice claudiub "under the same device_id" ??
15:17:55 dansmith stephenfin: doesn't that have an impact on whether or not just unsetting it is the right path in the bottom patch?
15:18:15 fried_rice claudiub oh, what the concept of parent_addr handles today.
15:18:20 dansmith stephenfin: your bottom patch says it's only for curses-based interaction or whatever, but that reviewer onthe deprecation patch seems to think it matters?
15:18:21 stephenfin Not really. We're not actually unsetting it in the bottom patch. Only allowing it to be unset
15:18:29 dansmith sure, but..
15:18:49 claudiub fried_rice: oh yeah, something like that.
15:19:14 stephenfin ...and that entire first paragraph is essentially taken from danpb's comments on the bug
15:19:20 claudiub fried_rice: then, adding that parent_addr making the pci device whitelist support that parent_addr would work for me
15:19:34 dansmith stephenfin: maybe the reno on the bottom patch should drop the "we'll be deprecating it soon" part and just say "you can unset it if you want now"
15:19:46 fried_rice claudiub Yeah, today you can whitelist the parent... PF, I think. And then the claim matches any VF whose parent_addr is in the whitelist.
15:19:52 stephenfin dansmith: Not a bad call. I'll do that instead
15:20:55 fried_rice claudiub Course, like I said, SR-IOV is a whole different can of worms on Power. Cause we create our "VFs" on the fly, but the thing that gets assigned to the VM is really a virtual-virtual-function
15:21:14 dansmith stephenfin: actually it's the warning log that needs to go I think. I dropped some comments ont here
15:21:41 claudiub fried_rice: according to the docs, the only valid keys in the passthrough_whitelist config option are: vendor_id, product_id, address, devname, physical_network
15:22:05 claudiub fried_rice: well, that sounds the same as hyper-v.
15:22:18 fried_rice claudiub Right, the parent_addr is produced by the pci_passthrough_devices list (get_available_resource)
15:23:17 claudiub fried_rice: i wonder if we can't report the parent_addr, and whitelist the devices using the parent_addr
15:23:22 fried_rice This is where, if the claim is asking for a VF, it matches whitelist entries to pci_passthrough_devices entries based on their parent_addr in the latter, not their actual addr.
15:23:41 fried_rice claudiub Yeah, ^^ that's effectively what already happens.
15:23:43 fried_rice IIUC
15:24:50 fried_rice But... getting the virtual-virtual-function thing to work is a whole different ball of wax. We have a delicate dance between our compute driver and our mech driver.
15:25:40 fried_rice We had to have our compute driver spoof the list of VFs in passthrough_devices (because they don't exist yet; and if they did, they still wouldn't be available/visible/accessible on the compute node)
15:27:12 claudiub fried_rice: same here.
15:27:45 fried_rice So the claim just arbitrarily picks one off, and we kinda ignore that bit; and then it binds the port based on the physnet and the mech driver kinda passes it back and then our compute driver does the virtual-virtual-function (which we call an SR-IOV VNIC, confusingly) construction and assignment to the VM.
15:28:12 fried_rice claudiub Here's where I would really want generic and/or nested resource providers to help me out.
15:28:47 claudiub fried_rice: same here. :)
15:28:51 fried_rice I want the RP to be able to say "I can supply this many VFs" (or VNICs, or whatever) - just like it says "I can supply this many VCPUs".
15:29:04 fried_rice And when I do a claim, it decrements by one and lets my compute driver do the rest.
15:29:39 claudiub fried_rice: that sounds ideal, IMO.
15:29:55 fried_rice claudiub Okay, so same question about blueprint collaboration there.
15:30:55 claudiub fried_rice: I'd help in any shape / form I could, but I am a bit swamped at the moment, so I can't make any promises. :)
15:32:06 fried_rice claudiub Really all I'm asking for is a hearty "me too" if I propose this stuff. If this is just for PowerVM, it's a tough sell. But if there's more than one driver that cares, it carries a lot more weight.
15:33:02 claudiub fried_rice: then yes, that would help me as well. :)
15:33:06 fried_rice claudiub You going to the PTG?
15:35:34 claudiub fried_rice: unfortunately, no
15:36:31 fried_rice Okay. There's a topic queued up (see https://etherpad.openstack.org/p/nova-ptg-queens ~L66). If you want to add a note in there, at least, that'll help when I take the floor with it.
15:37:25 fried_rice I would like to have this same discussion with someone from VMWare - see if this direction is also of interest there. Any idea who would be a good touchpoint there? stephenfin
15:38:03 stephenfin dansmith: I've no idea, unfortunately. mriedem might know though?
15:38:09 stephenfin Oops - fried_rice ^
15:38:17 mriedem cdent is vmware

Earlier   Later