Earlier  
Posted Nick Remark
#openstack-nova - 2022-11-08
19:05:56 opendevreview Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806
19:05:57 opendevreview Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055
20:08:25 opendevreview Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806
20:08:26 opendevreview Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055
23:00:12 clarkb Hello nova. I've been testing server rescue behavior recently and run into two different interesting behaviors. The first is when rescuing a non bfv instance if the rootfs label configured in grub and fstab is the same between the rescue image and the instance being recsued you can boot into the rescue image kernel but have the rescued instance / mounted
23:00:57 clarkb This is problematic because if the problem is say in systemd init steps you'd still be broken in the rescued setup. Additionally it may not always be the case that the kernel and your rescued filesystem are compatbile enough to boot
23:01:26 clarkb That said I'm not sure if nova can do anything to deal with this problem direclty. I suspect one of the best options is for clouds to have a purpose built rescue image that isn't likely to collide in this way
23:02:22 clarkb The other issue is when doing rescue on a bfv instance the api accepts the request (as long as yo uset the api version high enough) but then the instance promptly goes into an error state with Driver Error: Cannot access storage file and no such file or directory with a path.
23:02:47 clarkb It almost looks like nova / libvirt are looking for a disk file rather than looking for the volume via whatever the volume provider mechansim is
23:03:11 clarkb I don't have enough access on the cloud side to debug this further so I'm not sure if this is a nova issue or a cloud configuration issue etc.
23:03:50 clarkb It does make me wonder if we've got any rescue testing to ensure this generally works? Also, should we try to write more docs on how to ensure rescues work?
23:04:26 clarkb I'm coming at this as an end user wishing this functioned bette rand wondering what I/we can do to get there. But I lack a lot of background and knowledge on the inner working here :)
23:05:11 clarkb separately, I do wonder if it makes more sense for bfv rescue to be a process more like "stop instnace, detach volume from instance, attach volume to another instance, make changes, reattach to original instance and start original instance" but this fails because you can't detach a root device
23:29:32 opendevreview melanie witt proposed openstack/nova master: DNM testing images_type = raw with resize enabled https://review.opendev.org/c/openstack/nova/+/862416
#openstack-nova - 2022-11-09
01:37:45 melwitt clarkb: this is the test coverage we have for rescue https://github.com/openstack/tempest/blob/master/tempest/api/compute/servers/test_server_rescue.py and it's enabled in the tempest-integrated-compute job for example https://zuul.opendev.org/t/openstack/build/c35d560c76a24e45959aa609ac372d67/log/controller/logs/tempest_conf.txt#70
06:23:09 opendevreview Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806
06:23:10 opendevreview Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055
07:48:03 opendevreview Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806
07:48:04 opendevreview Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055
08:20:39 opendevreview Nobuhiro MIKI proposed openstack/nova master: libvirt: add maxphysaddr support https://review.opendev.org/c/openstack/nova/+/864091
10:06:50 samuelkunkel[m] Good morning,... (full message at <https://matrix.org/_matrix/media/r0/download/matrix.org/rADMLssdKgBpMiHEvywbsOpx>)
10:21:50 frickler samuelkunkel[m]: your message has been truncated by the matrix bridge. I suggest not to use matrix in order to join IRC. if you think that this is still the right solution for you, make sure your messages are not too long
10:22:37 frickler in particular avoiding to send multiline messages may be helpful
10:23:20 samuelkunkel[m] ah sure, sorry. I can try to make it single line. Links still should work? gonna look for a different client...
10:23:51 samuelkunkel[m] we are currently facing an issue in yoga with libvirt 8.0 for reporting mdev devices
10:23:57 samuelkunkel[m] in particular https://review.opendev.org/c/openstack/nova/+/838976
10:24:03 samuelkunkel[m] is this still being worked on?
10:24:10 samuelkunkel[m] (hope it is readable now)
10:25:18 frickler seem bauzas was the last one working on it
10:25:51 samuelkunkel[m] currently I will use the quick fix provided https://review.opendev.org/c/openstack/nova/+/838976
10:26:43 bauzas frickler: yup, I need to update my change
10:26:53 bauzas it's a priority I have
10:38:36 samuelkunkel[m] that sounds nice, if you need somebody to test - feel free to reach out to me, have some nodes with mdevs to play on
11:56:41 ygk_12345 HI all
12:04:39 sean-k-mooney samuelkunkel[m]: we not only plan to fix that but backport the fix to wallaby as we require it for our downstream product that far and there is no point doing it downstream only since the fix is backpoartable
12:05:19 sean-k-mooney so given your on yoga that shoudl hopefully also adress your usecase
12:05:33 samuelkunkel[m] yes, that sounds great
12:05:48 samuelkunkel[m] I assume there is currently no estimation possible on a timeframe?
12:06:17 sean-k-mooney well the patch thats propsoed actully works we just need a few comments adressed
12:06:45 auniyal Hi sean-k-mooney
12:07:02 sean-k-mooney downstream we have a dealine of mid decemebr to adress this so i am stongly hoping that we can adress this upstream before then so our product team does not start asking me about it
12:07:11 auniyal how can we run tox functional locally in train branch
12:07:20 auniyal tox -e functional fails
12:07:27 sean-k-mooney use the python3 version
12:07:51 sean-k-mooney or a vm/container based on ubutu 18.04?
12:07:52 samuelkunkel[m] I can second that, it also works on my yoga setup on a non productive cluster. Thanks for the clarification. Until the fix is backported I just use the patch
12:07:56 samuelkunkel[m] thanks for all the information
12:09:00 sean-k-mooney auniyal: so on tain you can use tox -e functional-py36 or tox -e functional-py37
12:10:37 sean-k-mooney auniyal: i would either use ubuntu 18.04/ubuntu-bionic or centos 8 stream to run the tests
12:11:14 sean-k-mooney we use 18.04 in teh ci https://github.com/openstack/nova/blob/stable/train/.zuul.yaml#L72-L119
12:11:19 auniyal got same error, I think its trying to need some package/module
12:11:20 auniyal https://paste.opendev.org/show/bsE4F25vNPl8BaGh7I6I/
12:11:45 sean-k-mooney you are trying to use 3.8
12:12:24 auniyal oh in here - /usr/lib/python3.8/runpy.py
12:12:34 sean-k-mooney do you have 3.6 avaiable
12:12:53 auniyal no right now 3.6
12:12:57 auniyal 3.8
12:13:13 sean-k-mooney ya 3.8 was not released/supported by train
12:13:33 auniyal if I create venv of 3.6 and install test-requirements.txt in it
12:13:36 auniyal will it work
12:13:47 sean-k-mooney so if you want to run these you need to use an operating system that was support hence why i said centos 8 stream or ubuntu 18.04
12:14:14 auniyal ack, will go with ubuntu 18,
12:14:23 auniyal thanks Sean
12:14:42 sean-k-mooney if you host os is too new thing liek sqlight might have issues
12:15:05 sean-k-mooney basically where we have python modules that wrap c libs
12:15:23 sean-k-mooney if your host os lib is too new then the old python bindign might now work
12:15:53 sean-k-mooney so if your currently using say the latest fedora you are likely to have issues with old releases like train
12:16:16 sean-k-mooney i generally use vms or contaienr to work around that if i hit that
12:16:54 auniyal yeah I am using vm , devstack on ubuntu 20
12:17:31 auniyal fo this tests, will go with ubuntu 18
12:17:57 sean-k-mooney ack i used to keep a few vms around for backporting
12:18:31 sean-k-mooney i do that less now just because its rare that i need older then 3.8
12:19:42 auniyal ack
12:19:45 sean-k-mooney i think we added 3.8 in ussuri so train is really the only release that does not supprot 3.8 officall now
12:20:18 sean-k-mooney on it was victoria
12:21:22 auniyal for ussuri also I was dependent on zuul, but as there less conflict so it need less tests
12:23:53 sean-k-mooney frickler: by the way i have been using matrix on and off via the element client pretty seamlessly for irc
12:24:07 sean-k-mooney frickler: i still use weechat as my main irc client
12:24:48 sean-k-mooney but if im not at my work laptop i somethimes use teh eleemnt client form my personal laptop or ipad to chat via teh matrix.org bridge
12:25:44 sean-k-mooney so ya if you keep messanges relitivly short (3-4 lines) it works fine i havent hit the lenght limit personlly
12:26:10 sean-k-mooney of the irc alternivies i have used matrix is really the only one i tollerate
12:27:16 sean-k-mooney if the element desktop clinet ever get the ablity to sign into two matix accounts at once it might even be something i would consider as a replacemnt for weechat
12:27:22 frickler sean-k-mooney: there have also been issues where the bridge disconnects but you do not notice on the matrix side, so my personal suggest is still to not use this, ymmv
12:29:54 sean-k-mooney ya i have not had that issue but i still use irc as my primary interface and matix as what i use when im traveling or not working from my normal location
12:30:12 sean-k-mooney so i porably would not notice if there were tempoiry issues
12:33:36 admin1 i have a vm which is always in a pause state in the hypervisor .. trying to unpause using virsh gives error: Timed out during operation: cannot acquire state change lock (held by monitor=remoteDispatchDomainCreateWithFlags) .. the vm is backed by volume on ceph, but ceph is fine and there are no locks
12:33:45 admin1 what can i do to check/troubleshoot this issue
12:33:56 admin1 i rebooted the hypervisor as well, no luck
12:34:50 sean-k-mooney this might be a lock crated by qemu
12:35:06 sean-k-mooney have you tried stopping the vm and staring it
12:35:14 sean-k-mooney e.g. via a hard reboot
12:35:47 admin1 when i do a vrish destroy, it disappears from virsh list --all
12:36:06 admin1 when i start again (horizon/cli) appears back
12:36:09 admin1 with a paused state
12:36:55 sean-k-mooney ack
12:37:19 sean-k-mooney did you check the qemu instance log for any errors
12:37:35 sean-k-mooney this does not sound like a nova issue by the way
12:37:55 sean-k-mooney this sound like an issue at the qemu/libvirt level and or perhaps the ceph interaction
12:38:19 sean-k-mooney you dont happen to have a kvm error in the instance log do you?
12:38:59 sean-k-mooney we hit an issue with ubutu 22.04 where libvirt incorrectly detected the cpu model

Earlier   Later