Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-25
15:04:39 claudiub fried_rice: for sr-iov, the flavor doesn't change, only the neutron port's vnic_type has to bbe direct or something similar, and your ML2 mechanism_driver must be able to support it.
15:04:53 fried_rice Gobackgoback
15:04:59 fried_rice What does that alias look like?
15:05:10 claudiub for sr-iov, you don't need an alias
15:05:32 claudiub you only alias normal pci devices.
15:05:35 fried_rice Yeah, regular PCI devices. Cause if it's just [pci]alias = {"device_id": "1234"} - then you're effectively specifying one device per alias.
15:06:10 fried_rice and you might as well use extra_specs pci_passthrough:device_id:1234 rather than bothering with an alias entry.
15:07:13 stephenfin dansmith: Any chance you could take a look at https://review.openstack.org/#/c/496605/? Think it might be a good backport candidate
15:07:32 stephenfin the follow-up patch probably needs a little more discussion yet
15:07:36 claudiub fried_rice: well, tbh, for normal PCI devices, vendor_id / product_ids are always there on hyper-v. so, that can be used for aliasing. but IMO, a device_id should also be supported.
15:08:09 fried_rice So how would you do that? Perhaps something like [pci]alias = {"name": "foo", "devices": [{"device_id": "abc"}, {"device_id": "123"}, {"vendor_id": "1f2e", "product_id": "a9b8"}]} ?
15:08:10 claudiub fried_rice: yeah, you're right
15:08:31 fried_rice I.e. you can specify a list of stuff to an alias, that allows you to enumerate several devices by ID and/or clasess by prod/vendor?
15:08:35 fried_rice That would be... kinda cool.
15:09:28 fried_rice Even without the ability to specify device IDs in there, if I want to be able to group different prod/vendor types together under a single alias...
15:09:53 fried_rice Like maybe I only care if I get a GPU. Any GPU will do. So make an alias grouping with these vendor/product IDs that all represent GPUs.
15:12:36 dansmith stephenfin: why are you not just doing the conf deprecation now?
15:13:11 stephenfin dansmith: I want to backport the change. Backporting a deprecation seems wrong
15:13:25 stephenfin Same reason I've kept the deprecation timeline so generic
15:13:30 dansmith stephenfin: ah right okay
15:13:49 dansmith stephenfin: why no tests?
15:14:29 dansmith should be pretty easy to write one to make sure we skip that bit...
15:14:33 stephenfin Damn. Because I was cheating :(
15:14:46 kashyap stephenfin: :P
15:15:01 kashyap stephenfin: I noticed the "no tests" thing, but thought you intentionally skipped them with good reason
15:15:49 dansmith stephenfin: I missed the deprecation patch after this, which makes sense..
15:16:23 fried_rice claudiub Would you be interested in co-authoring/sponsoring a bp to allow whitelisting & aliasing by device ID?
15:16:26 stephenfin dansmith: Yup, that one apparently needs a little more discussion. Key mappings are a minefield
15:16:56 claudiub fried_rice: well, tbh, a device_id would work in sr-iov usecase. that device_id would represent the NIC to which the VFs belong to. thus, I'd have N VFs grouped under the same device_id, which can be aliased as well.
15:16:56 stephenfin fried_rice, claudiub: Stick me on the review if you do author such a bp/spec
15:17:16 fried_rice stephenfin ack
15:17:25 claudiub ould work in my sr-iov usecase *
15:17:42 fried_rice claudiub "under the same device_id" ??
15:17:55 dansmith stephenfin: doesn't that have an impact on whether or not just unsetting it is the right path in the bottom patch?
15:18:15 fried_rice claudiub oh, what the concept of parent_addr handles today.
15:18:20 dansmith stephenfin: your bottom patch says it's only for curses-based interaction or whatever, but that reviewer onthe deprecation patch seems to think it matters?
15:18:21 stephenfin Not really. We're not actually unsetting it in the bottom patch. Only allowing it to be unset
15:18:29 dansmith sure, but..
15:18:49 claudiub fried_rice: oh yeah, something like that.
15:19:14 stephenfin ...and that entire first paragraph is essentially taken from danpb's comments on the bug
15:19:20 claudiub fried_rice: then, adding that parent_addr making the pci device whitelist support that parent_addr would work for me
15:19:34 dansmith stephenfin: maybe the reno on the bottom patch should drop the "we'll be deprecating it soon" part and just say "you can unset it if you want now"
15:19:46 fried_rice claudiub Yeah, today you can whitelist the parent... PF, I think. And then the claim matches any VF whose parent_addr is in the whitelist.
15:19:52 stephenfin dansmith: Not a bad call. I'll do that instead
15:20:55 fried_rice claudiub Course, like I said, SR-IOV is a whole different can of worms on Power. Cause we create our "VFs" on the fly, but the thing that gets assigned to the VM is really a virtual-virtual-function
15:21:14 dansmith stephenfin: actually it's the warning log that needs to go I think. I dropped some comments ont here
15:21:41 claudiub fried_rice: according to the docs, the only valid keys in the passthrough_whitelist config option are: vendor_id, product_id, address, devname, physical_network
15:22:05 claudiub fried_rice: well, that sounds the same as hyper-v.
15:22:18 fried_rice claudiub Right, the parent_addr is produced by the pci_passthrough_devices list (get_available_resource)
15:23:17 claudiub fried_rice: i wonder if we can't report the parent_addr, and whitelist the devices using the parent_addr
15:23:22 fried_rice This is where, if the claim is asking for a VF, it matches whitelist entries to pci_passthrough_devices entries based on their parent_addr in the latter, not their actual addr.
15:23:41 fried_rice claudiub Yeah, ^^ that's effectively what already happens.
15:23:43 fried_rice IIUC
15:24:50 fried_rice But... getting the virtual-virtual-function thing to work is a whole different ball of wax. We have a delicate dance between our compute driver and our mech driver.
15:25:40 fried_rice We had to have our compute driver spoof the list of VFs in passthrough_devices (because they don't exist yet; and if they did, they still wouldn't be available/visible/accessible on the compute node)
15:27:12 claudiub fried_rice: same here.
15:27:45 fried_rice So the claim just arbitrarily picks one off, and we kinda ignore that bit; and then it binds the port based on the physnet and the mech driver kinda passes it back and then our compute driver does the virtual-virtual-function (which we call an SR-IOV VNIC, confusingly) construction and assignment to the VM.
15:28:12 fried_rice claudiub Here's where I would really want generic and/or nested resource providers to help me out.
15:28:47 claudiub fried_rice: same here. :)
15:28:51 fried_rice I want the RP to be able to say "I can supply this many VFs" (or VNICs, or whatever) - just like it says "I can supply this many VCPUs".
15:29:04 fried_rice And when I do a claim, it decrements by one and lets my compute driver do the rest.
15:29:39 claudiub fried_rice: that sounds ideal, IMO.
15:29:55 fried_rice claudiub Okay, so same question about blueprint collaboration there.
15:30:55 claudiub fried_rice: I'd help in any shape / form I could, but I am a bit swamped at the moment, so I can't make any promises. :)
15:32:06 fried_rice claudiub Really all I'm asking for is a hearty "me too" if I propose this stuff. If this is just for PowerVM, it's a tough sell. But if there's more than one driver that cares, it carries a lot more weight.
15:33:02 claudiub fried_rice: then yes, that would help me as well. :)
15:33:06 fried_rice claudiub You going to the PTG?
15:35:34 claudiub fried_rice: unfortunately, no
15:36:31 fried_rice Okay. There's a topic queued up (see https://etherpad.openstack.org/p/nova-ptg-queens ~L66). If you want to add a note in there, at least, that'll help when I take the floor with it.
15:37:25 fried_rice I would like to have this same discussion with someone from VMWare - see if this direction is also of interest there. Any idea who would be a good touchpoint there? stephenfin
15:38:03 stephenfin dansmith: I've no idea, unfortunately. mriedem might know though?
15:38:09 stephenfin Oops - fried_rice ^
15:38:17 mriedem cdent is vmware
15:38:30 fried_rice oh, okay, cool.
15:38:35 stephenfin dansmith: Should I remove that deprecation warning entirely or simply soften the language?
15:38:50 dansmith stephenfin: I think you should remove the log statement entirely
15:39:00 stephenfin There's still some issues. A debug-level warning and pointer to the bug might be helpful
15:39:05 stephenfin K. I'll do that so
15:40:09 mriedem WOOT http://logs.openstack.org/44/497944/1/check/gate-tempest-dsvm-py35-ubuntu-xenial/ba85bc2/logs/etc/nova/nova_cell1.conf.txt.gz
15:40:14 mriedem log formatting ftw
15:40:38 mriedem clarkb: ^
15:40:45 mriedem now to find a devstack core
15:41:01 kashyap dansmith: Reading your comment on the unset 'keymap' patch, where you refer to the feedback on this - https://review.openstack.org/#/c/483994/4
15:41:30 kashyap dansmith: Did you also read my comment? Also, we don't know what version of noVNC did the tester try it with
15:41:37 openstackgerrit Stephen Finucane proposed openstack/nova master: conf: Allow users to unset 'keymap' options https://review.openstack.org/496605
15:41:51 stephenfin dansmith: ^
15:41:56 kashyap But I agree, we shouldn't hurry to deprecate it
15:42:30 stephenfin kashyap: Yeah, this is a good interim (and backportable) step. We can look into the issue in more detail over Queens
15:42:37 dansmith kashyap: I'm not sure what your comment has to do with my comment
15:42:52 kashyap "We probably should drop the warning here given the feedback on the following deprecation patch."
15:43:16 kashyap I think you were referring to the feedback by Tushar Patil?
15:44:08 kashyap Ah, I missed to note in _what_ scenario one can unset the '-k' option
15:44:47 kashyap *If* the noVNC client supports (from 0.6.1 & above it seems) the "QEMU RFB extension" (https://github.com/novnc/noVNC/pull/596), then the '-k' is not needed at all.
15:45:06 edmondsw claudiub: fried_rice: problem with spoofing vendor/product id would be that if they're not real how does the operator figure them out to put in the conf? Same issue as address
15:45:16 edmondsw sorry, finally got off my calls and catching up
15:45:27 kashyap Anyway, for now, stephenfin's allowing to unset the option is good.
15:46:27 dansmith stephenfin: see my comment just now.. I might be missing something
15:46:59 fried_rice edmondsw Yeah, I think that's why we want to be able to support alternative mechanisms for identifying the devices.
15:47:06 edmondsw +1
15:47:09 claudiub edmondsw: yep, that's what i'm thinking about. although, there is a way, at least for my scenario. since i'm doing an md5 of the NIC's PCI device_id, a small script can be provided to "figure out" the vendor_id and product_id

Earlier   Later