| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-08 | |||
| 23:00:12 | clarkb | Hello nova. I've been testing server rescue behavior recently and run into two different interesting behaviors. The first is when rescuing a non bfv instance if the rootfs label configured in grub and fstab is the same between the rescue image and the instance being recsued you can boot into the rescue image kernel but have the rescued instance / mounted | |
| 23:00:57 | clarkb | This is problematic because if the problem is say in systemd init steps you'd still be broken in the rescued setup. Additionally it may not always be the case that the kernel and your rescued filesystem are compatbile enough to boot | |
| 23:01:26 | clarkb | That said I'm not sure if nova can do anything to deal with this problem direclty. I suspect one of the best options is for clouds to have a purpose built rescue image that isn't likely to collide in this way | |
| 23:02:22 | clarkb | The other issue is when doing rescue on a bfv instance the api accepts the request (as long as yo uset the api version high enough) but then the instance promptly goes into an error state with Driver Error: Cannot access storage file and no such file or directory with a path. | |
| 23:02:47 | clarkb | It almost looks like nova / libvirt are looking for a disk file rather than looking for the volume via whatever the volume provider mechansim is | |
| 23:03:11 | clarkb | I don't have enough access on the cloud side to debug this further so I'm not sure if this is a nova issue or a cloud configuration issue etc. | |
| 23:03:50 | clarkb | It does make me wonder if we've got any rescue testing to ensure this generally works? Also, should we try to write more docs on how to ensure rescues work? | |
| 23:04:26 | clarkb | I'm coming at this as an end user wishing this functioned bette rand wondering what I/we can do to get there. But I lack a lot of background and knowledge on the inner working here :) | |
| 23:05:11 | clarkb | separately, I do wonder if it makes more sense for bfv rescue to be a process more like "stop instnace, detach volume from instance, attach volume to another instance, make changes, reattach to original instance and start original instance" but this fails because you can't detach a root device | |
| 23:29:32 | opendevreview | melanie witt proposed openstack/nova master: DNM testing images_type = raw with resize enabled https://review.opendev.org/c/openstack/nova/+/862416 | |
| #openstack-nova - 2022-11-09 | |||
| 01:37:45 | melwitt | clarkb: this is the test coverage we have for rescue https://github.com/openstack/tempest/blob/master/tempest/api/compute/servers/test_server_rescue.py and it's enabled in the tempest-integrated-compute job for example https://zuul.opendev.org/t/openstack/build/c35d560c76a24e45959aa609ac372d67/log/controller/logs/tempest_conf.txt#70 | |
| 06:23:09 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806 | |
| 06:23:10 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055 | |
| 07:48:03 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806 | |
| 07:48:04 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055 | |
| 08:20:39 | opendevreview | Nobuhiro MIKI proposed openstack/nova master: libvirt: add maxphysaddr support https://review.opendev.org/c/openstack/nova/+/864091 | |
| 10:06:50 | samuelkunkel[m] | Good morning,... (full message at <https://matrix.org/_matrix/media/r0/download/matrix.org/rADMLssdKgBpMiHEvywbsOpx>) | |
| 10:21:50 | frickler | samuelkunkel[m]: your message has been truncated by the matrix bridge. I suggest not to use matrix in order to join IRC. if you think that this is still the right solution for you, make sure your messages are not too long | |
| 10:22:37 | frickler | in particular avoiding to send multiline messages may be helpful | |
| 10:23:20 | samuelkunkel[m] | ah sure, sorry. I can try to make it single line. Links still should work? gonna look for a different client... | |
| 10:23:51 | samuelkunkel[m] | we are currently facing an issue in yoga with libvirt 8.0 for reporting mdev devices | |
| 10:23:57 | samuelkunkel[m] | in particular https://review.opendev.org/c/openstack/nova/+/838976 | |
| 10:24:03 | samuelkunkel[m] | is this still being worked on? | |
| 10:24:10 | samuelkunkel[m] | (hope it is readable now) | |
| 10:25:18 | frickler | seem bauzas was the last one working on it | |
| 10:25:51 | samuelkunkel[m] | currently I will use the quick fix provided https://review.opendev.org/c/openstack/nova/+/838976 | |
| 10:26:43 | bauzas | frickler: yup, I need to update my change | |
| 10:26:53 | bauzas | it's a priority I have | |
| 10:38:36 | samuelkunkel[m] | that sounds nice, if you need somebody to test - feel free to reach out to me, have some nodes with mdevs to play on | |
| 11:56:41 | ygk_12345 | HI all | |
| 12:04:39 | sean-k-mooney | samuelkunkel[m]: we not only plan to fix that but backport the fix to wallaby as we require it for our downstream product that far and there is no point doing it downstream only since the fix is backpoartable | |
| 12:05:19 | sean-k-mooney | so given your on yoga that shoudl hopefully also adress your usecase | |
| 12:05:33 | samuelkunkel[m] | yes, that sounds great | |
| 12:05:48 | samuelkunkel[m] | I assume there is currently no estimation possible on a timeframe? | |
| 12:06:17 | sean-k-mooney | well the patch thats propsoed actully works we just need a few comments adressed | |
| 12:06:45 | auniyal | Hi sean-k-mooney | |
| 12:07:02 | sean-k-mooney | downstream we have a dealine of mid decemebr to adress this so i am stongly hoping that we can adress this upstream before then so our product team does not start asking me about it | |
| 12:07:11 | auniyal | how can we run tox functional locally in train branch | |
| 12:07:20 | auniyal | tox -e functional fails | |
| 12:07:27 | sean-k-mooney | use the python3 version | |
| 12:07:51 | sean-k-mooney | or a vm/container based on ubutu 18.04? | |
| 12:07:52 | samuelkunkel[m] | I can second that, it also works on my yoga setup on a non productive cluster. Thanks for the clarification. Until the fix is backported I just use the patch | |
| 12:07:56 | samuelkunkel[m] | thanks for all the information | |
| 12:09:00 | sean-k-mooney | auniyal: so on tain you can use tox -e functional-py36 or tox -e functional-py37 | |
| 12:10:37 | sean-k-mooney | auniyal: i would either use ubuntu 18.04/ubuntu-bionic or centos 8 stream to run the tests | |
| 12:11:14 | sean-k-mooney | we use 18.04 in teh ci https://github.com/openstack/nova/blob/stable/train/.zuul.yaml#L72-L119 | |
| 12:11:19 | auniyal | got same error, I think its trying to need some package/module | |
| 12:11:20 | auniyal | https://paste.opendev.org/show/bsE4F25vNPl8BaGh7I6I/ | |
| 12:11:45 | sean-k-mooney | you are trying to use 3.8 | |
| 12:12:24 | auniyal | oh in here - /usr/lib/python3.8/runpy.py | |
| 12:12:34 | sean-k-mooney | do you have 3.6 avaiable | |
| 12:12:53 | auniyal | no right now 3.6 | |
| 12:12:57 | auniyal | 3.8 | |
| 12:13:13 | sean-k-mooney | ya 3.8 was not released/supported by train | |
| 12:13:33 | auniyal | if I create venv of 3.6 and install test-requirements.txt in it | |
| 12:13:36 | auniyal | will it work | |
| 12:13:47 | sean-k-mooney | so if you want to run these you need to use an operating system that was support hence why i said centos 8 stream or ubuntu 18.04 | |
| 12:14:14 | auniyal | ack, will go with ubuntu 18, | |
| 12:14:23 | auniyal | thanks Sean | |
| 12:14:42 | sean-k-mooney | if you host os is too new thing liek sqlight might have issues | |
| 12:15:05 | sean-k-mooney | basically where we have python modules that wrap c libs | |
| 12:15:23 | sean-k-mooney | if your host os lib is too new then the old python bindign might now work | |
| 12:15:53 | sean-k-mooney | so if your currently using say the latest fedora you are likely to have issues with old releases like train | |
| 12:16:16 | sean-k-mooney | i generally use vms or contaienr to work around that if i hit that | |
| 12:16:54 | auniyal | yeah I am using vm , devstack on ubuntu 20 | |
| 12:17:31 | auniyal | fo this tests, will go with ubuntu 18 | |
| 12:17:57 | sean-k-mooney | ack i used to keep a few vms around for backporting | |
| 12:18:31 | sean-k-mooney | i do that less now just because its rare that i need older then 3.8 | |
| 12:19:42 | auniyal | ack | |
| 12:19:45 | sean-k-mooney | i think we added 3.8 in ussuri so train is really the only release that does not supprot 3.8 officall now | |
| 12:20:18 | sean-k-mooney | on it was victoria | |
| 12:21:22 | auniyal | for ussuri also I was dependent on zuul, but as there less conflict so it need less tests | |
| 12:23:53 | sean-k-mooney | frickler: by the way i have been using matrix on and off via the element client pretty seamlessly for irc | |
| 12:24:07 | sean-k-mooney | frickler: i still use weechat as my main irc client | |
| 12:24:48 | sean-k-mooney | but if im not at my work laptop i somethimes use teh eleemnt client form my personal laptop or ipad to chat via teh matrix.org bridge | |
| 12:25:44 | sean-k-mooney | so ya if you keep messanges relitivly short (3-4 lines) it works fine i havent hit the lenght limit personlly | |
| 12:26:10 | sean-k-mooney | of the irc alternivies i have used matrix is really the only one i tollerate | |
| 12:27:16 | sean-k-mooney | if the element desktop clinet ever get the ablity to sign into two matix accounts at once it might even be something i would consider as a replacemnt for weechat | |
| 12:27:22 | frickler | sean-k-mooney: there have also been issues where the bridge disconnects but you do not notice on the matrix side, so my personal suggest is still to not use this, ymmv | |
| 12:29:54 | sean-k-mooney | ya i have not had that issue but i still use irc as my primary interface and matix as what i use when im traveling or not working from my normal location | |
| 12:30:12 | sean-k-mooney | so i porably would not notice if there were tempoiry issues | |
| 12:33:36 | admin1 | i have a vm which is always in a pause state in the hypervisor .. trying to unpause using virsh gives error: Timed out during operation: cannot acquire state change lock (held by monitor=remoteDispatchDomainCreateWithFlags) .. the vm is backed by volume on ceph, but ceph is fine and there are no locks | |
| 12:33:45 | admin1 | what can i do to check/troubleshoot this issue | |
| 12:33:56 | admin1 | i rebooted the hypervisor as well, no luck | |
| 12:34:50 | sean-k-mooney | this might be a lock crated by qemu | |
| 12:35:06 | sean-k-mooney | have you tried stopping the vm and staring it | |
| 12:35:14 | sean-k-mooney | e.g. via a hard reboot | |
| 12:35:47 | admin1 | when i do a vrish destroy, it disappears from virsh list --all | |
| 12:36:06 | admin1 | when i start again (horizon/cli) appears back | |
| 12:36:09 | admin1 | with a paused state | |
| 12:36:55 | sean-k-mooney | ack | |
| 12:37:19 | sean-k-mooney | did you check the qemu instance log for any errors | |
| 12:37:35 | sean-k-mooney | this does not sound like a nova issue by the way | |
| 12:37:55 | sean-k-mooney | this sound like an issue at the qemu/libvirt level and or perhaps the ceph interaction | |
| 12:38:19 | sean-k-mooney | you dont happen to have a kvm error in the instance log do you? | |
| 12:38:59 | sean-k-mooney | we hit an issue with ubutu 22.04 where libvirt incorrectly detected the cpu model | |
| 12:39:11 | sean-k-mooney | it enabled amd cpu flags in the domain on an intel host | |
| 12:39:25 | sean-k-mooney | that left teh vm in a paused state | |
| 12:39:59 | sean-k-mooney | although that would not expaling the lock message but i woudl check the qemu instance log in anycase | |
| 12:40:13 | admin1 | sean-k-mooney thanks . i know what to check for now | |