| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-25 | |||
| 14:37:51 | efried | Because as soon as I start chucking addresses around, nova tries to get at them under /sys/bus/... | |
| 14:38:13 | stephenfin | efried, jaypipes: Would we still have per-compute-node filtering in the resource provider world? | |
| 14:38:32 | claudiub | efried: ah yes, i see. that would make sense, IMO. the pci resource tracker should be updated to allow other filters to be configured, imo. | |
| 14:39:19 | efried | stephenfin So the way I would think it would work in a RP world is... | |
| 14:39:46 | efried | Your RP (in this case the compute node) inventories the devices it has available, *already* filtered by whitelist. | |
| 14:40:25 | cdent | meaning filtering as a descendant of get_inventory? | |
| 14:40:36 | jaypipes | stephenfin: are you asking whether there'd be a need for a whitelist conf option when we manage PCI devices in placement? | |
| 14:40:48 | stephenfin | jaypipes: Yup | |
| 14:41:01 | efried | cdent Within get_inventory, yeah. | |
| 14:41:08 | cdent | cool, good | |
| 14:41:14 | jaypipes | stephenfin: yes, there would still be a need for a way to filter out host devices that should not be allocatable to a consumer. | |
| 14:42:00 | jaypipes | stephenfin: I just want the pci_passthrough_whitelist CONF option to do a single thing. I'd like to see a separate CONF option (or YAML file even) that contains device inventory information that cannot be auto-discovered. | |
| 14:42:05 | claudiub | jaypipes: isn't that already done at the api level? | |
| 14:42:14 | jaypipes | claudiub: hmm? | |
| 14:42:21 | jaypipes | claudiub: not sure I understand you | |
| 14:42:22 | claudiub | with pci device alias config opion | |
| 14:42:55 | claudiub | afaik, you request a pci device through an alias. if a pci device is not aliased, i don't think it's allocatable | |
| 14:43:11 | efried | Correct. passthrough_whitelist is intersected with alias to get the list of claimable devs. | |
| 14:43:12 | jaypipes | claudiub: well, that's yet another thing that the pci_passthrough_whitelist CONF option does :( alias things... | |
| 14:43:32 | efried | jaypipes Well, [pci]alias aliases things. | |
| 14:43:44 | efried | AFAICT that's in place to make it easier to specify them in the flavor | |
| 14:43:51 | efried | extra_specs pci_passthrough:alias=<alias>:<count> | |
| 14:44:10 | jaypipes | there's still alias stuff in the whitelist option. | |
| 14:44:17 | efried | ? | |
| 14:44:35 | efried | no lo veo | |
| 14:44:37 | claudiub | so, then, you can still request PCI devices through other means, like vendor_id / product_id, or other fields like this? i didn't see this documented. | |
| 14:44:58 | claudiub | not counting sr-iov. | |
| 14:45:19 | claudiub | that's requested whenever a port is vnic_type direct or something similar | |
| 14:45:34 | efried | oh, that. | |
| 14:45:54 | claudiub | that's requested automatically * | |
| 14:45:56 | efried | That's done based on the physical_network tag in the passthrough_whitelist, I believe. Aliases are not involved at all in that flow. | |
| 14:46:11 | claudiub | efried: yeah | |
| 14:46:12 | efried | which is... goofy. | |
| 14:46:38 | efried | Hey, it's Friday. | |
| 14:47:19 | fried_rice | Yeah, handling SR-IOV is a separate can of worms. | |
| 14:47:48 | fried_rice | And btw, I think SR-IOV is where a lot of the /sys/bus business comes into play. Or rather, why it's there in the first place. | |
| 14:47:55 | claudiub | oh yeah. which reminds me. i didn't find a proper way to "associate" different SR-IOV NICs to different "physical_networks" | |
| 14:48:28 | fried_rice | claudiub How so? | |
| 14:48:44 | fried_rice | oh, the thing where you can only actually have one physical network | |
| 14:49:20 | fried_rice | Yeah, there's a limitation in the binding metadata in neutron. We tried to "fix" it in I think ocata, but armax shot us down. | |
| 14:49:38 | armax | fried_rice: what did I do? | |
| 14:49:49 | fried_rice | armax :) Hold on, let me see if I can dig it up. | |
| 14:50:24 | fried_rice | armax I think it was this one: https://review.openstack.org/#/c/358125/ | |
| 14:50:25 | armax | I don’t typically shoot people down without proper justification :) | |
| 14:50:58 | fried_rice | https://bugs.launchpad.net/neutron/+bug/1615128 | |
| 14:50:59 | openstack | Launchpad bug 1615128 in neutron "Custom binding:profile values not coming through" [Undecided,Invalid] - Assigned to Eric Fried (efried) | |
| 14:51:04 | fried_rice | Oh, you provided justification | |
| 14:51:08 | fried_rice | And we backed off accordingly. | |
| 14:51:37 | jaypipes | fried_rice: apologies, I was wrong on the pci_passthrough + alias thing. :( | |
| 14:51:54 | armax | fried_rice: right, I remember this now…I suppose the next step from the -2 was to follow up with an RFE that never came | |
| 14:52:05 | armax | iirc | |
| 14:52:25 | fried_rice | armax That sounds right. | |
| 14:52:44 | fried_rice | armax Wasn't trying to bust your chops :) | |
| 14:52:53 | armax | fried_rice: so I only injured you, not shot you down mortally :) | |
| 14:53:28 | armax | fried_rice: not a problem, I make mistakes of judment all the time, I was just trying to figure out the context | |
| 14:53:35 | claudiub | fried_rice: hm, simply because the limitation of the passthrough_whitelist config option. I can't reliably get a vendor_id / product_id for the NICs I have, and I can't use the address field either. IMO, if device_id would be allowed, it would work. | |
| 14:53:56 | fried_rice | claudiub Well | |
| 14:53:57 | claudiub | hm, haven't tried devname though. | |
| 14:54:08 | fried_rice | claudiub I haven't actually found any reason you can't spoof the vendor/product ID. | |
| 14:54:20 | fried_rice | I don't *think* those are actually checked against anything real, ever. | |
| 14:54:54 | fried_rice | They're just correlated among the passthrough_whitelist (compute conf), the pci_passthrough_devices (compute-provided get_available_resource), and the alias list (api conf). | |
| 14:54:57 | claudiub | the PCI devices reported by the compute drivers have to have those vendor_id / product_ids | |
| 14:55:27 | fried_rice | Yes, have to have them, but I don't think anything cares that those values are really what the dev reports. | |
| 14:55:45 | claudiub | true. | |
| 14:56:09 | fried_rice | I'm... sort of counting on it, actually. Cause I have a vendor ID, but I'm not sure if the thing I'm using as a product ID is really a product ID. | |
| 14:56:24 | fried_rice | https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/host.py@111 | |
| 14:56:59 | claudiub | for sr-iov devices, i'm currently doing an md5, if i can't get the actual vendor_id, product_id, which feels wrong. :) | |
| 14:57:00 | fried_rice | Course, if you do that (or really, whatever you do), you have to document how your user needs to glean the right value to put in those spots in the whitelist/alias. | |
| 14:57:11 | fried_rice | claudiub Totally | |
| 14:57:16 | fried_rice | I feel your pain. | |
| 14:57:43 | fried_rice | claudiub I'm doing something similar with PCI addresses: https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/vm.py@884 | |
| 14:58:05 | claudiub | *ahem*, if only we could also whitelist via device_id, that would work for me. would it work for you as well? | |
| 14:58:25 | fried_rice | claudiub Where device_id can be an arbitrary string, yes, totally. | |
| 14:59:32 | fried_rice | claudiub In order to do wildcarding, though, you would need to outsource the validation/filtering to the compute driver. | |
| 14:59:46 | fried_rice | Cause he's the only guy who knows how to interpret that "arbitrary" format. | |
| 15:00:17 | fried_rice | Otherwise you would just get a straight string match, which means you have to enumerate every device (or continue to use classes via vendor/prod IDs) | |
| 15:02:05 | fried_rice | claudiub What does the flavor side look like in that case? | |
| 15:02:17 | claudiub | fried_rice: for sr-iov you mean? | |
| 15:02:20 | fried_rice | I don't like the idea of having a flavor that specifies individual devices | |
| 15:02:40 | claudiub | fried_rice: or normal pci devices | |
| 15:02:54 | fried_rice | Let's go normal PCI for now. | |
| 15:03:46 | claudiub | fried_rice: it should still look the same: pci_passthrough:alias=<alias>:<count> | |
| 15:04:16 | fried_rice | so you're adding "device_id" as an alternative to vendor_id/product_id in the alias? | |
| 15:04:39 | claudiub | fried_rice: for sr-iov, the flavor doesn't change, only the neutron port's vnic_type has to bbe direct or something similar, and your ML2 mechanism_driver must be able to support it. | |
| 15:04:53 | fried_rice | Gobackgoback | |
| 15:04:59 | fried_rice | What does that alias look like? | |
| 15:05:10 | claudiub | for sr-iov, you don't need an alias | |
| 15:05:32 | claudiub | you only alias normal pci devices. | |
| 15:05:35 | fried_rice | Yeah, regular PCI devices. Cause if it's just [pci]alias = {"device_id": "1234"} - then you're effectively specifying one device per alias. | |
| 15:06:10 | fried_rice | and you might as well use extra_specs pci_passthrough:device_id:1234 rather than bothering with an alias entry. | |
| 15:07:13 | stephenfin | dansmith: Any chance you could take a look at https://review.openstack.org/#/c/496605/? Think it might be a good backport candidate | |
| 15:07:32 | stephenfin | the follow-up patch probably needs a little more discussion yet | |
| 15:07:36 | claudiub | fried_rice: well, tbh, for normal PCI devices, vendor_id / product_ids are always there on hyper-v. so, that can be used for aliasing. but IMO, a device_id should also be supported. | |
| 15:08:09 | fried_rice | So how would you do that? Perhaps something like [pci]alias = {"name": "foo", "devices": [{"device_id": "abc"}, {"device_id": "123"}, {"vendor_id": "1f2e", "product_id": "a9b8"}]} ? | |
| 15:08:10 | claudiub | fried_rice: yeah, you're right | |
| 15:08:31 | fried_rice | I.e. you can specify a list of stuff to an alias, that allows you to enumerate several devices by ID and/or clasess by prod/vendor? | |
| 15:08:35 | fried_rice | That would be... kinda cool. | |
| 15:09:28 | fried_rice | Even without the ability to specify device IDs in there, if I want to be able to group different prod/vendor types together under a single alias... | |
| 15:09:53 | fried_rice | Like maybe I only care if I get a GPU. Any GPU will do. So make an alias grouping with these vendor/product IDs that all represent GPUs. | |
| 15:12:36 | dansmith | stephenfin: why are you not just doing the conf deprecation now? | |