| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-25 | |||
| 14:55:45 | claudiub | true. | |
| 14:56:09 | fried_rice | I'm... sort of counting on it, actually. Cause I have a vendor ID, but I'm not sure if the thing I'm using as a product ID is really a product ID. | |
| 14:56:24 | fried_rice | https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/host.py@111 | |
| 14:56:59 | claudiub | for sr-iov devices, i'm currently doing an md5, if i can't get the actual vendor_id, product_id, which feels wrong. :) | |
| 14:57:00 | fried_rice | Course, if you do that (or really, whatever you do), you have to document how your user needs to glean the right value to put in those spots in the whitelist/alias. | |
| 14:57:11 | fried_rice | claudiub Totally | |
| 14:57:16 | fried_rice | I feel your pain. | |
| 14:57:43 | fried_rice | claudiub I'm doing something similar with PCI addresses: https://review.openstack.org/#/c/496434/3/nova_powervm/virt/powervm/vm.py@884 | |
| 14:58:05 | claudiub | *ahem*, if only we could also whitelist via device_id, that would work for me. would it work for you as well? | |
| 14:58:25 | fried_rice | claudiub Where device_id can be an arbitrary string, yes, totally. | |
| 14:59:32 | fried_rice | claudiub In order to do wildcarding, though, you would need to outsource the validation/filtering to the compute driver. | |
| 14:59:46 | fried_rice | Cause he's the only guy who knows how to interpret that "arbitrary" format. | |
| 15:00:17 | fried_rice | Otherwise you would just get a straight string match, which means you have to enumerate every device (or continue to use classes via vendor/prod IDs) | |
| 15:02:05 | fried_rice | claudiub What does the flavor side look like in that case? | |
| 15:02:17 | claudiub | fried_rice: for sr-iov you mean? | |
| 15:02:20 | fried_rice | I don't like the idea of having a flavor that specifies individual devices | |
| 15:02:40 | claudiub | fried_rice: or normal pci devices | |
| 15:02:54 | fried_rice | Let's go normal PCI for now. | |
| 15:03:46 | claudiub | fried_rice: it should still look the same: pci_passthrough:alias=<alias>:<count> | |
| 15:04:16 | fried_rice | so you're adding "device_id" as an alternative to vendor_id/product_id in the alias? | |
| 15:04:39 | claudiub | fried_rice: for sr-iov, the flavor doesn't change, only the neutron port's vnic_type has to bbe direct or something similar, and your ML2 mechanism_driver must be able to support it. | |
| 15:04:53 | fried_rice | Gobackgoback | |
| 15:04:59 | fried_rice | What does that alias look like? | |
| 15:05:10 | claudiub | for sr-iov, you don't need an alias | |
| 15:05:32 | claudiub | you only alias normal pci devices. | |
| 15:05:35 | fried_rice | Yeah, regular PCI devices. Cause if it's just [pci]alias = {"device_id": "1234"} - then you're effectively specifying one device per alias. | |
| 15:06:10 | fried_rice | and you might as well use extra_specs pci_passthrough:device_id:1234 rather than bothering with an alias entry. | |
| 15:07:13 | stephenfin | dansmith: Any chance you could take a look at https://review.openstack.org/#/c/496605/? Think it might be a good backport candidate | |
| 15:07:32 | stephenfin | the follow-up patch probably needs a little more discussion yet | |
| 15:07:36 | claudiub | fried_rice: well, tbh, for normal PCI devices, vendor_id / product_ids are always there on hyper-v. so, that can be used for aliasing. but IMO, a device_id should also be supported. | |
| 15:08:09 | fried_rice | So how would you do that? Perhaps something like [pci]alias = {"name": "foo", "devices": [{"device_id": "abc"}, {"device_id": "123"}, {"vendor_id": "1f2e", "product_id": "a9b8"}]} ? | |
| 15:08:10 | claudiub | fried_rice: yeah, you're right | |
| 15:08:31 | fried_rice | I.e. you can specify a list of stuff to an alias, that allows you to enumerate several devices by ID and/or clasess by prod/vendor? | |
| 15:08:35 | fried_rice | That would be... kinda cool. | |
| 15:09:28 | fried_rice | Even without the ability to specify device IDs in there, if I want to be able to group different prod/vendor types together under a single alias... | |
| 15:09:53 | fried_rice | Like maybe I only care if I get a GPU. Any GPU will do. So make an alias grouping with these vendor/product IDs that all represent GPUs. | |
| 15:12:36 | dansmith | stephenfin: why are you not just doing the conf deprecation now? | |
| 15:13:11 | stephenfin | dansmith: I want to backport the change. Backporting a deprecation seems wrong | |
| 15:13:25 | stephenfin | Same reason I've kept the deprecation timeline so generic | |
| 15:13:30 | dansmith | stephenfin: ah right okay | |
| 15:13:49 | dansmith | stephenfin: why no tests? | |
| 15:14:29 | dansmith | should be pretty easy to write one to make sure we skip that bit... | |
| 15:14:33 | stephenfin | Damn. Because I was cheating :( | |
| 15:14:46 | kashyap | stephenfin: :P | |
| 15:15:01 | kashyap | stephenfin: I noticed the "no tests" thing, but thought you intentionally skipped them with good reason | |
| 15:15:49 | dansmith | stephenfin: I missed the deprecation patch after this, which makes sense.. | |
| 15:16:23 | fried_rice | claudiub Would you be interested in co-authoring/sponsoring a bp to allow whitelisting & aliasing by device ID? | |
| 15:16:26 | stephenfin | dansmith: Yup, that one apparently needs a little more discussion. Key mappings are a minefield | |
| 15:16:56 | claudiub | fried_rice: well, tbh, a device_id would work in sr-iov usecase. that device_id would represent the NIC to which the VFs belong to. thus, I'd have N VFs grouped under the same device_id, which can be aliased as well. | |
| 15:16:56 | stephenfin | fried_rice, claudiub: Stick me on the review if you do author such a bp/spec | |
| 15:17:16 | fried_rice | stephenfin ack | |
| 15:17:25 | claudiub | ould work in my sr-iov usecase * | |
| 15:17:42 | fried_rice | claudiub "under the same device_id" ?? | |
| 15:17:55 | dansmith | stephenfin: doesn't that have an impact on whether or not just unsetting it is the right path in the bottom patch? | |
| 15:18:15 | fried_rice | claudiub oh, what the concept of parent_addr handles today. | |
| 15:18:20 | dansmith | stephenfin: your bottom patch says it's only for curses-based interaction or whatever, but that reviewer onthe deprecation patch seems to think it matters? | |
| 15:18:21 | stephenfin | Not really. We're not actually unsetting it in the bottom patch. Only allowing it to be unset | |
| 15:18:29 | dansmith | sure, but.. | |
| 15:18:49 | claudiub | fried_rice: oh yeah, something like that. | |
| 15:19:14 | stephenfin | ...and that entire first paragraph is essentially taken from danpb's comments on the bug | |
| 15:19:20 | claudiub | fried_rice: then, adding that parent_addr making the pci device whitelist support that parent_addr would work for me | |
| 15:19:34 | dansmith | stephenfin: maybe the reno on the bottom patch should drop the "we'll be deprecating it soon" part and just say "you can unset it if you want now" | |
| 15:19:46 | fried_rice | claudiub Yeah, today you can whitelist the parent... PF, I think. And then the claim matches any VF whose parent_addr is in the whitelist. | |
| 15:19:52 | stephenfin | dansmith: Not a bad call. I'll do that instead | |
| 15:20:55 | fried_rice | claudiub Course, like I said, SR-IOV is a whole different can of worms on Power. Cause we create our "VFs" on the fly, but the thing that gets assigned to the VM is really a virtual-virtual-function | |
| 15:21:14 | dansmith | stephenfin: actually it's the warning log that needs to go I think. I dropped some comments ont here | |
| 15:21:41 | claudiub | fried_rice: according to the docs, the only valid keys in the passthrough_whitelist config option are: vendor_id, product_id, address, devname, physical_network | |
| 15:22:05 | claudiub | fried_rice: well, that sounds the same as hyper-v. | |
| 15:22:18 | fried_rice | claudiub Right, the parent_addr is produced by the pci_passthrough_devices list (get_available_resource) | |
| 15:23:17 | claudiub | fried_rice: i wonder if we can't report the parent_addr, and whitelist the devices using the parent_addr | |
| 15:23:22 | fried_rice | This is where, if the claim is asking for a VF, it matches whitelist entries to pci_passthrough_devices entries based on their parent_addr in the latter, not their actual addr. | |
| 15:23:41 | fried_rice | claudiub Yeah, ^^ that's effectively what already happens. | |
| 15:23:43 | fried_rice | IIUC | |
| 15:24:50 | fried_rice | But... getting the virtual-virtual-function thing to work is a whole different ball of wax. We have a delicate dance between our compute driver and our mech driver. | |
| 15:25:40 | fried_rice | We had to have our compute driver spoof the list of VFs in passthrough_devices (because they don't exist yet; and if they did, they still wouldn't be available/visible/accessible on the compute node) | |
| 15:27:12 | claudiub | fried_rice: same here. | |
| 15:27:45 | fried_rice | So the claim just arbitrarily picks one off, and we kinda ignore that bit; and then it binds the port based on the physnet and the mech driver kinda passes it back and then our compute driver does the virtual-virtual-function (which we call an SR-IOV VNIC, confusingly) construction and assignment to the VM. | |
| 15:28:12 | fried_rice | claudiub Here's where I would really want generic and/or nested resource providers to help me out. | |
| 15:28:47 | claudiub | fried_rice: same here. :) | |
| 15:28:51 | fried_rice | I want the RP to be able to say "I can supply this many VFs" (or VNICs, or whatever) - just like it says "I can supply this many VCPUs". | |
| 15:29:04 | fried_rice | And when I do a claim, it decrements by one and lets my compute driver do the rest. | |
| 15:29:39 | claudiub | fried_rice: that sounds ideal, IMO. | |
| 15:29:55 | fried_rice | claudiub Okay, so same question about blueprint collaboration there. | |
| 15:30:55 | claudiub | fried_rice: I'd help in any shape / form I could, but I am a bit swamped at the moment, so I can't make any promises. :) | |
| 15:32:06 | fried_rice | claudiub Really all I'm asking for is a hearty "me too" if I propose this stuff. If this is just for PowerVM, it's a tough sell. But if there's more than one driver that cares, it carries a lot more weight. | |
| 15:33:02 | claudiub | fried_rice: then yes, that would help me as well. :) | |
| 15:33:06 | fried_rice | claudiub You going to the PTG? | |
| 15:35:34 | claudiub | fried_rice: unfortunately, no | |
| 15:36:31 | fried_rice | Okay. There's a topic queued up (see https://etherpad.openstack.org/p/nova-ptg-queens ~L66). If you want to add a note in there, at least, that'll help when I take the floor with it. | |
| 15:37:25 | fried_rice | I would like to have this same discussion with someone from VMWare - see if this direction is also of interest there. Any idea who would be a good touchpoint there? stephenfin | |
| 15:38:03 | stephenfin | dansmith: I've no idea, unfortunately. mriedem might know though? | |
| 15:38:09 | stephenfin | Oops - fried_rice ^ | |
| 15:38:17 | mriedem | cdent is vmware | |
| 15:38:30 | fried_rice | oh, okay, cool. | |
| 15:38:35 | stephenfin | dansmith: Should I remove that deprecation warning entirely or simply soften the language? | |
| 15:38:50 | dansmith | stephenfin: I think you should remove the log statement entirely | |
| 15:39:00 | stephenfin | There's still some issues. A debug-level warning and pointer to the bug might be helpful | |
| 15:39:05 | stephenfin | K. I'll do that so | |
| 15:40:09 | mriedem | WOOT http://logs.openstack.org/44/497944/1/check/gate-tempest-dsvm-py35-ubuntu-xenial/ba85bc2/logs/etc/nova/nova_cell1.conf.txt.gz | |
| 15:40:14 | mriedem | log formatting ftw | |