Earlier  
Posted Nick Remark
#openstack-nova - 2021-11-16
13:55:07 sean-k-mooney huh interesting
13:55:51 sean-k-mooney the VFs for the connectx-6 on the bluefiled 2 have the same vendor and product id as a normal connectx-6
13:59:59 bauzas elodilles: sure, please do it, I'll just update the wikipage after you
14:06:46 dmitriis sean-k-mooney: ack on the address field usage.
14:07:57 dmitriis sean-k-mooney: I don't have a separate ConnectX-6 at hand but BF2 has ConnectX-6 in it. Let me check the PCI ID DB - I think I've seen different ids but maybe that's for something else.
14:11:10 elodilles bauzas: done, thanks (i might overused the info and link markers o:) feel free to edit :))
14:11:29 bauzas elodilles: ack, thanks
14:13:07 dmitriis sean-k-mooney: so the PF is different but VFs look like the ones from a "regular" ConnectX-6.
14:13:07 dmitriis PF:
14:13:07 dmitriis 82:00.0 Ethernet controller: Mellanox Technologies MT42822 BlueField-2 integrated ConnectX-6 Dx network controller (rev 01)
14:13:07 dmitriis 82:00.0 0200: 15b3:a2d6 (rev 01)
14:13:07 dmitriis Subsystem: 15b3:0061
14:13:23 dmitriis VF:
14:13:23 dmitriis 82:00.3 Ethernet controller: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function (rev 01)
14:13:24 dmitriis 82:00.3 0200: 15b3:101e (rev 01)
14:13:24 dmitriis Subsystem: 15b3:0061
14:14:15 dmitriis so careful "remote_managed" tagging is needed
14:14:44 sean-k-mooney ack good to know
14:29:21 sean-k-mooney dmitriis: you can use the adress of the pf and vendor id of the VF to whitelist all the VFs that belog to that pf
14:29:24 sean-k-mooney just so you know
14:29:44 sean-k-mooney dmitriis: that behavior is not well known
14:32:37 dmitriis sean-k-mooney: didn't know that (not surprisingly), thanks for the info.
14:33:14 dmitriis sean-k-mooney: btw, BF2 does bonding at the ARM CPU side transparently to the hypervisor
14:34:37 dmitriis and there's an option to hide the inactive PF for the hypervisor side: https://docs.mellanox.com/display/BlueFieldSWv24011082/BlueField%20Link%20Aggregation
14:35:27 dmitriis that makes it easier for OpenStack deployers/operators since only one PF needs to be taken into account
14:42:16 dmitriis sean-k-mooney: so this is the place where I had to use a string (instead of bool, not None so my reference was not correct) https://review.opendev.org/c/openstack/nova/+/812111/3/nova/network/neutron.py#2295 - that's where a device spec is dynamically generated (not based on flavor or image properties).
14:45:41 sean-k-mooney im on a call but ill look it up after thanks
14:47:31 dmitriis ack
15:03:22 Adri2000 hi, I've got a race condition issue on ussuri and victoria when resizing an instance... specifically this is with /var/lib/nova/instances on NFS, and the following happens sometimes when resizing an instance where a cold migration is triggered: `qemu-img resize` will be run on the new compute node, before the old compute node has fully released the lock on the disk file; this
15:03:25 Adri2000 will put the instance in ERROR state. does that ring a bell to anyone?
15:03:36 Adri2000 ERROR nova.compute.manager [req-...] [instance: 6ca672fd-8746-441f-bbca-6baa3234bb5e] Setting instance vm_state to ERROR: oslo_concurrency.processutils.ProcessExecutionError: Unexpected error while running command. Command: qemu-img resize /var/lib/nova/instances/6ca672fd-8746-441f-bbca-6baa3234bb5e/disk Exit code: 1 Stdout: ''
15:03:43 Adri2000 Stderr: "qemu-img: Could not open '/var/lib/nova/instances/6ca672fd-8746-441f-bbca-6baa3234bb5e/disk': Could not open '/var/lib/nova/instances/6ca672fd-8746-441f-bbca-6baa3234bb5e/disk': Permission denied\n"
15:04:09 sean-k-mooney Adri2000: are you using nfsv3
15:04:37 Adri2000 sean-k-mooney: `/var/lib/nova/instances type nfs4 (rw,relatime,vers=4.1...`
15:05:10 sean-k-mooney ok nfsv3 has locking issues v4.1 improves the situration but recommend v4.2+
15:05:21 sean-k-mooney lyarwood: does ^ seem familar to you
15:07:54 sean-k-mooney Adri2000: i belive there are some tunabel in the mount option that can be used to help resolve this
15:11:37 sean-k-mooney Adri2000: are you using raw images?
15:11:57 bauzas gibi: I'll have to hardstop our meeting by 5:50pm our TZ
15:12:15 gibi bauzas: ack
15:12:27 bauzas in case we have to continue discussing, could you be chairing it ?
15:13:32 sean-k-mooney dmitriis: oh there
15:13:42 sean-k-mooney str(self._is_remote_managed(vnic_type)),
15:13:53 sean-k-mooney dmitriis: ya that makes sense
15:13:57 Adri2000 sean-k-mooney: qcow3 images. one nfs option I have currently is local_lock=none, maybe I should look into this one.
15:14:18 sean-k-mooney dmitriis: technicaly the tags are defiend to be of type String
15:14:31 sean-k-mooney so its a dict of string to string
15:15:12 sean-k-mooney Adri2000: ack the reason i asked is apparently the lockign behavior is differnt in qemu for raw vs qcow
15:15:16 dmitriis sean-k-mooney: ack
15:17:26 sean-k-mooney dmitriis: https://github.com/openstack/nova/blob/master/nova/pci/devspec.py#L262-L263
15:18:27 dmitriis sean-k-mooney: yep, makes sense
15:18:51 dmitriis sean-k-mooney: btw, the ovn-vif repo is now up under ovn-org https://github.com/ovn-org/ovn-vif
15:19:06 sean-k-mooney yes i saw your comment
15:19:10 sean-k-mooney just looking at the code
15:19:17 dmitriis ack
15:19:44 sean-k-mooney am i right in assuming we do not want to allow these device to be used for flavor based pci passthouhg
15:20:38 dmitriis sean-k-mooney: yes, they won't be of much use without being plugged appropriately. Not the VFs at least.
15:21:00 sean-k-mooney ya
15:21:11 sean-k-mooney im wondering if we should explictly block that
15:21:44 sean-k-mooney unfortunetly i dont see a trivial way to do that
15:22:05 sean-k-mooney although we might already do that
15:22:20 dmitriis sean-k-mooney: I guess we could exclude devices from search results if remote_managed is present but not requestedd
15:22:46 sean-k-mooney yep
15:22:58 sean-k-mooney i was just going to provide an exmaple
15:23:04 sean-k-mooney we do this in other cases already
15:23:33 sean-k-mooney https://github.com/openstack/nova/blob/master/nova/pci/stats.py#L411-L433
15:23:46 sean-k-mooney That filteres out PF if you did not ask for one
15:24:08 sean-k-mooney dmitriis: so you can copy past https://github.com/openstack/nova/blob/master/nova/pci/stats.py#L520-L535
15:24:27 sean-k-mooney and then add a new function that will filter out remote managed device if not requested
15:24:49 dmitriis sean-k-mooney: https://review.opendev.org/c/openstack/nova/+/812111/3/nova/network/neutron.py#2295
15:25:00 sean-k-mooney i did this recently when i added support for vdpa https://github.com/openstack/nova/blob/master/nova/pci/stats.py#L540
15:25:01 dmitriis actually, I'm explicitly passing remote_managed=False
15:25:28 sean-k-mooney that wont work
15:25:55 dmitriis sean-k-mooney: even with this? https://review.opendev.org/c/openstack/nova/+/812111/3/nova/pci/stats.py#111
15:25:58 sean-k-mooney it will break existing deploymetn on upgrade as there existing device wont have remote_managed=False
15:26:07 sean-k-mooney and it woudl only apply for pci recuts form ports
15:26:52 sean-k-mooney dmitriis: that would work but we might end up doign a data migration of all existing rows
15:27:34 sean-k-mooney dmitriis: ok we can review this as part of the code review rather then the spec.
15:27:51 sean-k-mooney im just finsihing reading it now and ill approve it shortly
15:27:52 dmitriis sean-k-mooney: ack, I am open to adding a filter as you suggested
15:28:35 sean-k-mooney either would work but one involes updateing every row in the pci device tabel with remote_managed=false :)
15:28:48 dmitriis right, I would certainly like to avoid introducing a change that would break with a stale state in PCI stats
15:28:50 sean-k-mooney the important thing is there is not a gap in the design
15:28:55 dmitriis agreed
15:43:38 sean-k-mooney dmitriis: ok i captured some of my tought in this converstaion in the spec but +2 +w from me
15:45:12 sean-k-mooney dmitriis: feel free to ping me to review the implemetaion too. i do not have +2 right on the code repo but ill try and spend some time reviewing it end to end next week
15:46:14 dmitriis sean-k-mooney: tyvm. I'll try to get the code updated with some of the latest changes by then. Still have to extend func tests to cover more cases but there are some already.
15:47:03 dmitriis sean-k-mooney: speaking of other lifecycle operations, I've spent some time looking at the recent VF hot-plug/unplug changes so I may revisit some of the unsupported operations at a later point
15:47:14 bauzas reminder : nova weekly meeting starts in 13 mins here in this #chan
15:47:33 dmitriis maybe we can actually make things like cold migration work, just need to review that further
15:47:45 sean-k-mooney ack
15:48:09 sean-k-mooney dmitriis: it might just work
15:48:29 sean-k-mooney there is very littel in the nova side that will need to be updated
15:48:37 sean-k-mooney also for the live migration
15:48:46 dmitriis sean-k-mooney: yes, we might need to document the need for extra slots to be added via the new config
15:49:05 dmitriis https://review.opendev.org/c/openstack/nova/+/545034/16/nova/conf/libvirt.py
15:49:13 sean-k-mooney dmitriis: that is really only need for q35
15:49:26 sean-k-mooney and we already have a config to add extra slots in that case
15:49:47 sean-k-mooney yep that one
15:49:50 dmitriis ack

Earlier   Later