Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-25
14:39:46 efried Your RP (in this case the compute node) inventories the devices it has available, *already* filtered by whitelist.
14:40:25 cdent meaning filtering as a descendant of get_inventory?
14:40:36 jaypipes stephenfin: are you asking whether there'd be a need for a whitelist conf option when we manage PCI devices in placement?
14:40:48 stephenfin jaypipes: Yup
14:41:01 efried cdent Within get_inventory, yeah.
14:41:08 cdent cool, good
14:41:14 jaypipes stephenfin: yes, there would still be a need for a way to filter out host devices that should not be allocatable to a consumer.
14:42:00 jaypipes stephenfin: I just want the pci_passthrough_whitelist CONF option to do a single thing. I'd like to see a separate CONF option (or YAML file even) that contains device inventory information that cannot be auto-discovered.
14:42:05 claudiub jaypipes: isn't that already done at the api level?
14:42:14 jaypipes claudiub: hmm?
14:42:21 jaypipes claudiub: not sure I understand you
14:42:22 claudiub with pci device alias config opion
14:42:55 claudiub afaik, you request a pci device through an alias. if a pci device is not aliased, i don't think it's allocatable
14:43:11 efried Correct. passthrough_whitelist is intersected with alias to get the list of claimable devs.
14:43:12 jaypipes claudiub: well, that's yet another thing that the pci_passthrough_whitelist CONF option does :( alias things...
14:43:32 efried jaypipes Well, [pci]alias aliases things.
14:43:44 efried AFAICT that's in place to make it easier to specify them in the flavor
14:43:51 efried extra_specs pci_passthrough:alias=<alias>:<count>
14:44:10 jaypipes there's still alias stuff in the whitelist option.
14:44:17 efried ?
14:44:35 efried no lo veo
14:44:37 claudiub so, then, you can still request PCI devices through other means, like vendor_id / product_id, or other fields like this? i didn't see this documented.
14:44:58 claudiub not counting sr-iov.
14:45:19 claudiub that's requested whenever a port is vnic_type direct or something similar
14:45:34 efried oh, that.
14:45:54 claudiub that's requested automatically *
14:45:56 efried That's done based on the physical_network tag in the passthrough_whitelist, I believe. Aliases are not involved at all in that flow.
14:46:11 claudiub efried: yeah
14:46:12 efried which is... goofy.
14:46:38 efried Hey, it's Friday.
14:47:19 fried_rice Yeah, handling SR-IOV is a separate can of worms.
14:47:48 fried_rice And btw, I think SR-IOV is where a lot of the /sys/bus business comes into play. Or rather, why it's there in the first place.
14:47:55 claudiub oh yeah. which reminds me. i didn't find a proper way to "associate" different SR-IOV NICs to different "physical_networks"
14:48:28 fried_rice claudiub How so?
14:48:44 fried_rice oh, the thing where you can only actually have one physical network
14:49:20 fried_rice Yeah, there's a limitation in the binding metadata in neutron. We tried to "fix" it in I think ocata, but armax shot us down.
14:49:38 armax fried_rice: what did I do?
14:49:49 fried_rice armax :) Hold on, let me see if I can dig it up.
14:50:24 fried_rice armax I think it was this one: https://review.openstack.org/#/c/358125/
14:50:25 armax I don’t typically shoot people down without proper justification :)
14:50:58 fried_rice https://bugs.launchpad.net/neutron/+bug/1615128
14:50:59 openstack Launchpad bug 1615128 in neutron "Custom binding:profile values not coming through" [Undecided,Invalid] - Assigned to Eric Fried (efried)
14:51:04 fried_rice Oh, you provided justification
14:51:08 fried_rice And we backed off accordingly.
14:51:37 jaypipes fried_rice: apologies, I was wrong on the pci_passthrough + alias thing. :(
14:51:54 armax fried_rice: right, I remember this now…I suppose the next step from the -2 was to follow up with an RFE that never came
14:52:05 armax iirc
14:52:25 fried_rice armax That sounds right.
14:52:44 fried_rice armax Wasn't trying to bust your chops :)
14:52:53 armax fried_rice: so I only injured you, not shot you down mortally :)
14:53:28 armax fried_rice: not a problem, I make mistakes of judment all the time, I was just trying to figure out the context
14:53:35 claudiub fried_rice: hm, simply because the limitation of the passthrough_whitelist config option. I can't reliably get a vendor_id / product_id for the NICs I have, and I can't use the address field either. IMO, if device_id would be allowed, it would work.
14:53:56 fried_rice claudiub Well
14:53:57 claudiub hm, haven't tried devname though.
14:54:08 fried_rice claudiub I haven't actually found any reason you can't spoof the vendor/product ID.
14:54:20 fried_rice I don't *think* those are actually checked against anything real, ever.
14:54:54 fried_rice They're just correlated among the passthrough_whitelist (compute conf), the pci_passthrough_devices (compute-provided get_available_resource), and the alias list (api conf).
14:54:57 claudiub the PCI devices reported by the compute drivers have to have those vendor_id / product_ids
14:55:27 fried_rice Yes, have to have them, but I don't think anything cares that those values are really what the dev reports.
14:55:45 claudiub true.
14:56:09 fried_rice I'm... sort of counting on it, actually. Cause I have a vendor ID, but I'm not sure if the thing I'm using as a product ID is really a product ID.
14:56:24 fried_rice https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/host.py@111
14:56:59 claudiub for sr-iov devices, i'm currently doing an md5, if i can't get the actual vendor_id, product_id, which feels wrong. :)
14:57:00 fried_rice Course, if you do that (or really, whatever you do), you have to document how your user needs to glean the right value to put in those spots in the whitelist/alias.
14:57:11 fried_rice claudiub Totally
14:57:16 fried_rice I feel your pain.
14:57:43 fried_rice claudiub I'm doing something similar with PCI addresses: https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/vm.py@884
14:58:05 claudiub *ahem*, if only we could also whitelist via device_id, that would work for me. would it work for you as well?
14:58:25 fried_rice claudiub Where device_id can be an arbitrary string, yes, totally.
14:59:32 fried_rice claudiub In order to do wildcarding, though, you would need to outsource the validation/filtering to the compute driver.
14:59:46 fried_rice Cause he's the only guy who knows how to interpret that "arbitrary" format.
15:00:17 fried_rice Otherwise you would just get a straight string match, which means you have to enumerate every device (or continue to use classes via vendor/prod IDs)
15:02:05 fried_rice claudiub What does the flavor side look like in that case?
15:02:17 claudiub fried_rice: for sr-iov you mean?
15:02:20 fried_rice I don't like the idea of having a flavor that specifies individual devices
15:02:40 claudiub fried_rice: or normal pci devices
15:02:54 fried_rice Let's go normal PCI for now.
15:03:46 claudiub fried_rice: it should still look the same: pci_passthrough:alias=<alias>:<count>
15:04:16 fried_rice so you're adding "device_id" as an alternative to vendor_id/product_id in the alias?
15:04:39 claudiub fried_rice: for sr-iov, the flavor doesn't change, only the neutron port's vnic_type has to bbe direct or something similar, and your ML2 mechanism_driver must be able to support it.
15:04:53 fried_rice Gobackgoback
15:04:59 fried_rice What does that alias look like?
15:05:10 claudiub for sr-iov, you don't need an alias
15:05:32 claudiub you only alias normal pci devices.
15:05:35 fried_rice Yeah, regular PCI devices. Cause if it's just [pci]alias = {"device_id": "1234"} - then you're effectively specifying one device per alias.
15:06:10 fried_rice and you might as well use extra_specs pci_passthrough:device_id:1234 rather than bothering with an alias entry.
15:07:13 stephenfin dansmith: Any chance you could take a look at https://review.openstack.org/#/c/496605/? Think it might be a good backport candidate
15:07:32 stephenfin the follow-up patch probably needs a little more discussion yet
15:07:36 claudiub fried_rice: well, tbh, for normal PCI devices, vendor_id / product_ids are always there on hyper-v. so, that can be used for aliasing. but IMO, a device_id should also be supported.
15:08:09 fried_rice So how would you do that? Perhaps something like [pci]alias = {"name": "foo", "devices": [{"device_id": "abc"}, {"device_id": "123"}, {"vendor_id": "1f2e", "product_id": "a9b8"}]} ?
15:08:10 claudiub fried_rice: yeah, you're right
15:08:31 fried_rice I.e. you can specify a list of stuff to an alias, that allows you to enumerate several devices by ID and/or clasess by prod/vendor?
15:08:35 fried_rice That would be... kinda cool.
15:09:28 fried_rice Even without the ability to specify device IDs in there, if I want to be able to group different prod/vendor types together under a single alias...
15:09:53 fried_rice Like maybe I only care if I get a GPU. Any GPU will do. So make an alias grouping with these vendor/product IDs that all represent GPUs.
15:12:36 dansmith stephenfin: why are you not just doing the conf deprecation now?
15:13:11 stephenfin dansmith: I want to backport the change. Backporting a deprecation seems wrong
15:13:25 stephenfin Same reason I've kept the deprecation timeline so generic
15:13:30 dansmith stephenfin: ah right okay
15:13:49 dansmith stephenfin: why no tests?

Earlier   Later