| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-06 | |||
| 15:08:57 | sean-k-mooney | dansmith: im currently trying to fix some os-vif tox issues related to ubuntu 22.04 | |
| 15:09:26 | sean-k-mooney | 4.0 is after that on my list | |
| 15:16:24 | sean-k-mooney | dansmith: we had a pin in place until recently to prevent the gate block | |
| 15:16:51 | dansmith | yeah and we're pinning on stable, I'm just saying I don't think _temporarily_ pinning and expecting things to stabilize is realistic | |
| 15:16:57 | sean-k-mooney | i understand why they remvoed it but i dont think we shoudl block the gate while we are fixing it | |
| 15:17:12 | dansmith | it's not like they broke a bunch of backwards compat in 4.0 and now things are stable.. they *keep* breaking things | |
| 15:18:03 | dansmith | also for the reason that tox will auto-upgrade itself in certain scenarios (which is like ....) | |
| 15:18:09 | sean-k-mooney | i havent really had issue wiht tox but also havnt been using it much in the last while as i have not been really coding in python for a few months | |
| 15:18:38 | sean-k-mooney | dansmith: apprently you can force it to install iseslf in a venv and use that version to run things | |
| 15:18:53 | dansmith | sean-k-mooney: it will do that itself if it decides to | |
| 15:19:02 | dansmith | but only the latest, not a specific version | |
| 15:19:34 | sean-k-mooney | not according to Brian Rosmaita's latest email | |
| 15:19:56 | sean-k-mooney | you can force the version via requires in tox.ini | |
| 15:20:19 | dansmith | " it doesn't ensure that the available tox is that version." | |
| 15:20:54 | dansmith | oh, there's two pins, with different behaviors | |
| 15:21:13 | dansmith | he's talking about requires, but there are projects with ensure | |
| 15:22:17 | dansmith | it's really a mess | |
| 15:23:08 | sean-k-mooney | yep | |
| 15:23:32 | sean-k-mooney | the reason i was suggestign we wait a week is at that poitn we would be 4 weeks form FF | |
| 15:23:45 | sean-k-mooney | and dont really wnat to still have the gates blocked by this at that point | |
| 15:24:15 | sean-k-mooney | i.e. lets see if we can fix it next week and if not pin it so we can continue merging things and work on it in parallel | |
| 18:02:23 | sean-k-mooney | gmann: stephenfin i have got the os-vif fucntional test workign locally | |
| 18:03:06 | sean-k-mooney | it looks like we need CAP_DAC_OVERRIDE on ubuntu 22.04 | |
| 18:03:36 | sean-k-mooney | without that vsctl and some other commands fail | |
| 18:03:46 | sean-k-mooney | CAP_NET_ADMIN used to work | |
| 18:04:00 | sean-k-mooney | i have some other chagne locally so im going to see if they are required or not | |
| 18:04:51 | sean-k-mooney | its proably because i am not a meber of the openvswitch group but it also fails for ip link commands | |
| 18:05:16 | sean-k-mooney | so i think this has to do with disto packaging and how the goups are configured | |
| 18:05:37 | sean-k-mooney | so while i dont like adding CAP_DAC_OVERRIED that is proably what we will need to do | |
| 18:06:44 | sean-k-mooney | what im less happy about is this is only required when using the vsctl ovs backend which is deprecated | |
| 18:07:07 | sean-k-mooney | so i might us a diffferent privsep context based on the driver to limit the scope of the change. | |
| 18:20:38 | sean-k-mooney | actully i think the cahnge is less in vasive then that and only in the test code | |
| 18:34:20 | opendevreview | sean mooney proposed openstack/os-vif master: add CAP_DAC_OVERRIDE to test privsep contexts https://review.opendev.org/c/openstack/os-vif/+/869500 | |
| 18:35:59 | sean-k-mooney | gibi: stephenfin gmann ^ i think that will fix the functional job and unblock https://review.opendev.org/c/openstack/os-vif/+/868420 and https://review.opendev.org/c/openstack/os-vif/+/861468 | |
| 18:36:29 | sean-k-mooney | once we have those 3 commits merged we may want to consider an os-vif release | |
| 18:54:32 | gmann | sean-k-mooney: thanks, will keep eyes on gate result | |
| 19:09:25 | darkhorse | Hi team, I would like to resize an shelved_offloaded instance. The use case is that when an instance with pci device is shelved_offloaded and the device is broken, it fails to unshelve. However, as a user, I would like to recover my data in the instance so I would like to change the flavor with a new one that does not have pci cards. | |
| 19:09:49 | darkhorse | Is there a quick workaround to this? | |
| 19:10:05 | darkhorse | Thank you in advance for any help! | |
| 19:19:35 | opendevreview | Danylo Vodopianov proposed openstack/nova master: Napatech SmartNIC support https://review.opendev.org/c/openstack/nova/+/859577 | |
| 19:24:52 | dansmith | melwitt: around? | |
| 19:35:01 | melwitt | dansmith: o/ | |
| 19:35:17 | dansmith | hey so, | |
| 19:35:32 | dansmith | I don't really know what I was going to say | |
| 19:35:46 | dansmith | part of it was that now that I've implemented compute undelete later in the series, | |
| 19:35:51 | dansmith | I'm failing a couple more regression tests, | |
| 19:36:09 | dansmith | but around ironic because of all the hash ring rebalance weirdness | |
| 19:36:24 | dansmith | unfortunately I think those are going to have to change a bit as well, which makes me nervous, | |
| 19:36:56 | dansmith | but they're asserting things like "this shouldn't create another thing, but it does, so assert that it happens, and then assert that it goes away later" sort of stuff | |
| 19:37:12 | dansmith | which is kinda expected and kinda why we're doing this, and the ironic hash-ring-ectomy thing | |
| 19:37:42 | dansmith | but I dunno, I guess I just want to say... I hope you're going to check all my work :P | |
| 19:40:09 | melwitt | dansmith: ack, that's the intent (to check everything) :) thanks for the heads up | |
| 19:40:48 | dansmith | I know | |
| 19:41:10 | dansmith | but this rabbit hole is deep, the tea tastes funny, and everyone is wearing strange hats | |
| 19:42:50 | melwitt | haha, I hear you (and am not surprised) | |
| 19:49:58 | opendevreview | Danylo Vodopianov proposed openstack/nova master: Napatech SmartNIC support https://review.opendev.org/c/openstack/nova/+/859577 | |
| 19:52:51 | sean-k-mooney | darkhorse: no quick workaround. resize is currently not supproted while shelve_offloaded but we have discussed that it could be supported in the future | |
| 19:53:24 | sean-k-mooney | darkhorse: i think artom has already fixed the issue with pci device shelve however | |
| 19:53:37 | sean-k-mooney | so i dont think that happens any more | |
| 20:06:13 | darkhorse | sean-k-mooney: do you have a link to the patchset? when you say the issue is fixed, does that mean you can unshelve instances even if pci card is broken or unavailable? | |
| 20:12:52 | opendevreview | Danylo Vodopianov proposed openstack/nova master: Napatech SmartNIC support https://review.opendev.org/c/openstack/nova/+/859577 | |
| 20:14:05 | artom | sean-k-mooney, IIUC darkhorse wants to resize a shelved_offloaded instance | |
| 20:14:10 | artom | Which IIUC is... not a thingÉ | |
| 20:14:10 | artom | Which IIUC is... not a thingÉ | |
| 20:14:12 | artom | ? | |
| 20:14:18 | artom | As in, you have to unshelve first, and then resize? | |
| 20:14:34 | artom | And yeah, unshelve with PCI has been backported to... I want to say Ussuri? | |
| 20:14:49 | artom | Or maybe wallaby | |
| 20:14:54 | darkhorse | sean-k-mooney: If you can share the link of the discussion of the resize support for shelved instances, it would be helpful. I will take a look and work on it. | |
| 20:15:44 | darkhorse | artom: Do you mean you can unshelve instance when pci card is broken/unavailable in Ussuri or Wallaby? | |
| 20:16:56 | artom | darkhorse, https://review.opendev.org/q/Icfa8c1d6e84eab758af6223a2870078685584aaa | |
| 20:16:57 | artom | wallaby | |
| 20:19:03 | darkhorse | artom: We are operating on xena. So if I understood you correct, all I need to do to allow users to unshelve pci instance even if card is broken/unavailable is to backport this patch to xena, is that correct? | |
| 20:19:51 | artom | darkhorse, no, you should be set. Xena is after wallaby :) | |
| 20:20:00 | artom | When the master patch merged, master was xena | |
| 20:20:26 | artom | darkhorse, hold on though - define "card is broken/unavailable"? | |
| 20:20:52 | artom | The unselve will attempt to find a PCI card that fits the port (if it's Neutron SRIOV)/flavor | |
| 20:21:30 | artom | But... if no such cards are available, then it will (legitimately) fails to schedule | |
| 20:21:39 | darkhorse | artom: not neutron SRIOV but fpga device. | |
| 20:21:47 | artom | So flavor PCI passthrough... | |
| 20:21:59 | darkhorse | right! | |
| 20:22:16 | artom | That should just... work. Off the top of my head I don't recall any issues with PCI and unshelve | |
| 20:23:01 | darkhorse | artom: no it fails to unshelve because the pci device is unavailable. | |
| 20:23:35 | artom | Unavailable how? It got pulled from the server? :) | |
| 20:23:45 | darkhorse | in that case, i would like to either snapshot or resize the instance so that I don't lose the data inside it. | |
| 20:24:21 | darkhorse | either because the card is occupied by another instances or physically broken | |
| 20:32:33 | darkhorse | artom: did i answer your question? | |
| 20:34:18 | artom | darkhorse, ah, I think I see. If you can't unshelve the instance because the cloud lacks the resources the instance needs (in this case, a PCI card), you'd like to be able to boot it regardless with its disk intact, just without the PCI device | |
| 20:34:39 | artom | So a shelved_offloaded instance lives as an image in Glance | |
| 20:34:55 | artom | IIRC you should just be able to boot a new instance from that image? | |
| 20:35:12 | artom | If keeping the same UUID is important to you though, you're out of luck I believe :( | |
| 20:36:32 | darkhorse | artom: the point is i want to recover the data inside the instance. if i boot a new instance, i think i am not able to get the data? | |
| 20:37:00 | artom | If it's been shelved offloaded, its disk has been uploaded to Glance as an image. | |
| 20:37:27 | artom | But... if you want data to persist, the "real" solution is to use volumes | |
| 20:38:38 | darkhorse | artom: will you elaborate? i was thinking of snapshotting or resizing with a new flavor that does not have pci so that i can unshelve. | |
| 20:39:38 | artom | darkhorse, elaborate on which aspect? Volumes, or booting from the Glance image? | |
| 20:40:22 | darkhorse | 1. if booting from glance image will save the data 2. volumes | |
| 20:41:02 | darkhorse | artom:1. if booting from glance image will save the data 2. volumes | |
| 20:42:07 | artom | darkhorse, it's been a while since I've done this, but a shelved_offloaded image will have its disk uploaded as image in Glance | |
| 20:42:28 | artom | I believe you can just boot from that image with `openstack server create --image <image uuid> <etc>` | |