Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-25
14:00:52 efried On Power, our PCI devices don't have 32-bit domain:bus:slot.func PCI addresses.
14:01:24 efried The nova PCI code around whitelisting, passthrough devices, and aliases is *all* set up to assume domain:bus:slot.func
14:02:11 efried At the moment, if I want to make PCI passthrough work in PowerVM, I have to very carefully leave out any reference to addresses in the whitelist & alias specs.
14:02:25 openstackgerrit sahid proposed openstack/nova master: libvirt: slowly live-migration to ensure network is ready https://review.openstack.org/497457
14:02:25 openstackgerrit sahid proposed openstack/nova master: libvirt: add method to configure migration speed https://review.openstack.org/497456
14:02:26 openstackgerrit sahid proposed openstack/nova master: libvirt: bandwidth param should be set in guest migrate https://review.openstack.org/497455
14:02:40 efried And when I create my pci_passthrough_devices list (from get_available_resource), I have to spoof a domain:bus:slot.func PCI address for each dev.
14:02:52 efried Cause each dev has to have a distinct address, or it won't get recognized at all.
14:03:02 efried So this leads to a couple of problems.
14:03:30 efried One is that I have to hack these addresses horribly. (And from a 64-bit source, so possibility of collisions, etc.)
14:04:04 efried The other is that I've given my op the ability to whitelist device *classes* (vendor x prod ID) but not specific devices.
14:04:25 efried Because as soon as I try to specify a device in the whitelist, nova tries to get at /sys/bus/.../<PCI address>/...
14:04:33 efried ...which isn't a thing in PowerVM.
14:05:18 efried btw, I have reason to believe HyperV and VMWare (and maybe xen) suffer from a similar class of problem (the devs not being owned by the node running the compute process).
14:05:55 efried So
14:06:40 stephenfin HyperV too, for that matter. I don't know the specifics though
14:07:52 efried One possible inroads may be the new resource provider stuff.
14:08:12 cdent I was just going to say that nest providers will help with this, but it is a long road before it completely resolves it
14:08:55 efried If my compute node registers a rp that provides the PCI devs as generic resources, they could conceivably have arbitrary names/identifiers.
14:09:17 efried But we would have to replumb the whole whitelist concept if we wanted to support whitelist filtering.
14:10:44 efried I would be happy to reimplement whitelisting in a way that works for PowerVM, but would have to have some way to bypass the whitelisting that nova's doing.
14:10:57 cdent the long term goal (jaypipes can you confirm this?) would be to get rid of the on-node whitelist stuff
14:10:57 efried [pci]whitelist_by_compute_driver = True
14:11:05 stephenfin Yeah, the plan was to move most of the PCI device management code to resource providers
14:11:22 stephenfin Any vGPU stuff would also be done this way
14:11:23 efried Yeah, jaypipes and I talked briefly about that in Boston.
14:11:55 stephenfin Hmm, so how do you define what devices can/can't be used if you don't have unique identifiers like that?
14:12:06 efried oh, I have unique identifiers.
14:12:12 efried They just don't look like domain:bus:slot.func
14:12:19 cdent bigger better unique identifiers!
14:12:35 cdent bbuid
14:12:43 stephenfin Ah
14:13:24 efried My IDs look like U78CB.001.WZS0JZB-P1-C14 or 2105001F
14:13:53 stephenfin So I guess you'd do your whitelisting using those identifiers. If so , would that be done inside nova?
14:14:01 edmondsw (those are 2 different formats... not different values of the same format... either should work)
14:14:44 stephenfin i.e. 'if driver = powervm: do special whitelisting; else: do normal domain:bus:slot.func whitelist'
14:15:15 efried Yeah, in nova right now all the Whitelist and PciDeviceStats and PciAlias stuff is hardwired to expect domain:bus:slot.func format, down to the point of being able to wildcard any of those components.
14:15:25 efried So if you whitelisted *:12:ab.*, you would match a device with address abcd:12:ab.7 but not abcd:13:ab.7
14:16:10 efried And then there are some code paths that actually try to look up attributes of the device under /sys/bus - which is a total nonstarter on Power (and, I assume, HyperV and VMWare, where the devices also aren't on the compute node)
14:16:22 efried ...like when it's trying to figure out if a dev is a physical function.
14:16:44 efried ...or if the whitelist specifies a dev by name
14:17:08 stephenfin claudiub is hardly about, is he? I'm pretty sure the Hyper-V driver does have some PCI device support now
14:17:36 cdent I was just looking at some of that code and wondered if perhaps that it ought to be in the virt drivers: https://review.openstack.org/#/c/476642/
14:17:39 jaypipes cdent, efried: yeah, I remember chatting with you about it in Boston, though I think I said I would be happy to just stop using the whitelist for two purposes (filtering stuff that guests can use vs. inventory management of devices)
14:17:59 cdent it’s rather limiting that it is doing /sys file stuff ...
14:18:37 efried Yeah, those are two examples where I would like to be able to have e.g. "is_physical_function" be a compute driver override whose default impl could be the /sys/bus business, but in the powervm driver I can do it my way.
14:19:36 efried jaypipes Right, so the inventory is actually being done by get_available_resource (which will eventually become get_inventory when all the plumbing is ready).
14:19:47 efried The whitelist is how you limit which devices are allowed to be assigned.
14:20:12 efried And then the alias list is how you nickname devices or device classes so you can specify them easily in a flavor.
14:20:19 jaypipes ya
14:20:20 efried So really, the conf stuff isn't doing inventorying.
14:20:46 jaypipes if we can just have the whitelist do the former thing and not the latter, I'd be happy
14:20:53 jaypipes efried: zactly.
14:21:03 efried which former/latter?
14:21:11 jaypipes efried: my only question is why haven't you gotten it all done yet? WAITING!
14:21:18 mriedem (filtering stuff that guests can use vs. inventory management of devices)
14:21:21 mriedem "(filtering stuff that guests can use vs. inventory management of devices)"
14:21:27 mriedem you turkeys
14:21:30 jaypipes efried: former == filtering for guests.
14:21:54 efried hm, how is the current impl of the whitelist doing inventory management?
14:22:16 jaypipes would be nice if IBM would just hop on the "newfangled" PCI bus.
14:22:23 efried Hah!
14:22:30 claudiub stephenfin: o/
14:22:33 jaypipes efried: :P
14:22:58 claudiub stephenfin: yeah, we do support pci passthrough since ocata.
14:23:22 efried and jaypipes, I actually have a prototype/PoC that works for PowerVM. But I very carefully have to bypass PCI address handling in some interesting ways.
14:23:42 stephenfin claudiub: Cool. How do you manage whitelisting of those devices? I assume you're not indexing devices in domain:bus:slot.func format
14:23:46 jaypipes efried: yes, I can imagine.
14:23:51 efried https://review.openstack.org/#/c/496434/
14:24:01 claudiub stephenfin: we are.
14:24:55 claudiub stephenfin: we're currently only reporting devices which have the domain:bus:slot.func format
14:25:05 efried ...with e.g. [pci]alias = {"name": "USB", "product_id": "8241", "vendor_id": "104c", "device_type": "type-PCI"} and [pci]passthrough_whitelist = {"product_id": "8241", "vendor_id": "104c"}
14:26:08 efried claudiub Those devices show up on the compute node?
14:26:21 claudiub stephenfin: but I've mainly whitelisted devices using product_id / vendor_id
14:26:42 claudiub efried: if they're are passthrough-able, yes.
14:26:56 claudiub efried: they have to be prepared for passthrough first
14:27:06 efried claudiub How?
14:27:38 claudiub efried: https://github.com/openstack/nova/blob/master/releasenotes/notes/hyper-v-pci-passthrough-babf104d6bc2baa6.yaml
14:28:41 jaypipes efried: reviewed.
14:29:02 efried Oh, thanks jaypipes :)
14:29:22 claudiub stephenfin: anyways. it would be nice to be also be able to whitelist devices using other ways than product_id / vendor_id, or domain:bus:slot.func.
14:29:59 efried jaypipes Nice. None of that was worthy of a -1? Really?
14:30:10 stephenfin claudiub: What kind of IDs would you expect, e.g. how else can devices be identified in Hyper-V?
14:30:45 jaypipes efried: :)
14:30:59 jaypipes stephenfin: by green card.
14:31:00 efried jaypipes I really didn't expect you to review it, but I would like to point out the spoof_pci_address method (https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/vm.py@884)
14:31:21 claudiub stephenfin: for example, I've had this issue when working on the sr-iov support. apparenty, different NIC vendors have different PCI device ID formats. for example, Intel NICs contain the vendor_id and product_id, which can be extracted and reported. But other NICs, like Mellanox or Chelsio, do not.
14:31:37 claudiub stephenfin: a simple device_id would do, IMO.
14:31:53 jaypipes efried: oh, trust me, I saw it :)
14:33:22 efried claudiub Right, so I would like my [pci]passthrough_whitelist entry to identify devices in whatever way my compute deems appropriate, and have the logic for doing that whitelist filtering live in my compute driver.
14:36:46 claudiub efried: well, the compute driver is already reporting "all" the PCI devices it sees, right? after which the PCI resource tracker filters those compute driver reported PCI devices according to the configured whitelist. is there anything that doesn't fit your usecase?
14:37:33 efried claudiub Yeah: the ability to specify individual devices in the whitelist (vs. just vendor/prod ID classes)
14:37:51 efried Because as soon as I start chucking addresses around, nova tries to get at them under /sys/bus/...
14:38:13 stephenfin efried, jaypipes: Would we still have per-compute-node filtering in the resource provider world?
14:38:32 claudiub efried: ah yes, i see. that would make sense, IMO. the pci resource tracker should be updated to allow other filters to be configured, imo.
14:39:19 efried stephenfin So the way I would think it would work in a RP world is...
14:39:46 efried Your RP (in this case the compute node) inventories the devices it has available, *already* filtered by whitelist.
14:40:25 cdent meaning filtering as a descendant of get_inventory?
14:40:36 jaypipes stephenfin: are you asking whether there'd be a need for a whitelist conf option when we manage PCI devices in placement?
14:40:48 stephenfin jaypipes: Yup

Earlier   Later