Earlier  
Posted Nick Remark
#openstack-nova - 2022-02-11
10:05:18 kashyap gibi: Thanks for the link to the smaller repro; also check out Peter's response on that thread
10:05:32 kashyap He points out two possibilities:
10:05:32 kashyap 1) the guest OS didn't confirm the detach
10:05:33 kashyap 2) there was a recent bug in qemu triggered by using JSON syntax for -device
10:10:16 kashyap gibi: That's it: this looks like it --
10:10:17 kashyap "DEVICE_DELETED event is not delivered for device frontend if -device is configured via JSON"
10:10:21 kashyap https://bugzilla.redhat.com/show_bug.cgi?id=2036669
10:19:45 kashyap But based on the versions in the CI job, they should already have the fix:
10:19:48 kashyap - libvirt version: 8.0.0, package: 2.el9
10:19:51 kashyap - qemu-kvm-6.2.0-5.el9
10:27:42 kashyap gibi: When you're around, to rule out the above bug, I wonder if we could try this workaround:
10:28:18 kashyap On compute nodes, in /etc/libvit/qemu.conf:
10:28:27 kashyap capability_filters = [ "device.json" ]
10:34:47 gibi kashyap: hi!
10:35:22 gibi kashyap: sure, I will try to make that config change via devstack
10:35:31 gibi kashyap: does it require a libvirtd restart?
10:35:34 kashyap gibi: See my latest comment: https://bugs.launchpad.net/nova/+bug/1960346/
10:35:47 kashyap gibi: Yeah, it is required, afraid
10:36:09 gibi OK
10:36:10 gibi thanks
11:01:08 gibi kashyap: pushed new PS to https://review.opendev.org/c/openstack/devstack/+/828705 with the WA, lets see if it helps or not
11:01:29 kashyap gibi: Thank you! It will at least rule out the 2nd possibility above for sure.
11:01:46 gibi I can try to look at the first
11:01:55 gibi we can grab the console log
11:02:01 gibi after the failed detach
11:03:26 gibi hm, we already grabbing it in tempest
11:03:29 gibi let me find it
11:06:31 kashyap I see, I need to be AFK for an hour-ish; will come back and check
11:15:36 gibi added the consol log to the bug https://paste.opendev.org/show/bXXn63wbTPOwiCGC5xDI/
11:15:56 gibi nothing obviously wrong there
11:16:10 gibi but the guest is still in a state to getting IP from DHCP
11:16:26 gibi so maybe it is not fully boot when the detach was requested
11:25:27 gibi chateaulav: left some suggestions inline about the ovo backports
11:35:46 opendevreview Manuel Bentele proposed openstack/nova master: libvirt: Add properties to set advanced QXL video RAM settings https://review.opendev.org/c/openstack/nova/+/828674
12:03:44 opendevreview Manuel Bentele proposed openstack/nova master: libvirt: Add configuration options to set SPICE compression settings https://review.opendev.org/c/openstack/nova/+/828675
12:24:58 chateaulav gibi: thanks for the follow up, that makes sense i was doing research and reading last night and had found references to `obj_relationships`. but your comments align to what i was trying yesterday, i just had brought in the exception aspect. appreciated
12:25:33 gibi chateaulav: cool
12:37:44 opendevreview Manuel Bentele proposed openstack/nova master: libvirt: Add property to set number of screens per video adapter https://review.opendev.org/c/openstack/nova/+/828676
12:40:45 erlon sean-k-mooney: hey Sean, I believe I have finished implementing all suggested changes in the live migration rollback fix: https://review.opendev.org/q/topic:bug%252F1944619
12:41:15 erlon when you have a chance to give a look ill appreciate
13:22:56 rosmaita bauzas: fyi, i will be raising the minima in requirements for os-brick release: http://lists.openstack.org/pipermail/openstack-discuss/2022-February/027192.html
14:16:21 gibi bauzas, rosmaita: I quickly checked the os_brick requirements patch I see no major bump in any deps so I think it is not a risky change.
14:17:08 rosmaita gibi: ty
14:17:37 gibi and tempest is green so nova is co-installable with the new os_brick deps
14:36:43 gibi gmann, frickler, bauzas: about the centos-9-steam job failure https://bugs.launchpad.net/nova/+bug/1960346/ I conculded that the cirros guest is not fully booted when the volume detach happens and the guest OS does not release the device. We need https://review.opendev.org/q/topic:wait_until_sshable_pingable to solve this in general
14:49:08 kashyap gibi: So cracked the prob! It's the guest OS indeed - adding a delay helps here?
14:49:15 kashyap s/So/So you/
14:53:17 kashyap I think for now, going with the extra delay before the detach happens is fine. That saves more time here, before the big Tempest series gets merged
14:55:49 gibi kashyap: I don't have brains any more today but next week I can put up a tempest patch with some selective delays. I'm not sure how well QA will appreciate it
14:56:39 gibi also I can take a look at lyarwood's series and try to move that forward
14:56:46 gibi kashyap: thanks you for your help!
15:04:40 frickler gibi: thx for the update, do you know why this only occurs on c9s? is booting slower or did previous libvirts not care whether that release actually happens?
15:19:27 gibi frickler: I think older libvirt let nova to restart the detach process but newer libvirt simply rejectes the retry as the original detach is still ongoing
15:22:29 opendevreview Merged openstack/nova stable/victoria: Add a WA flag waiting for vif-plugged event during reboot https://review.opendev.org/c/openstack/nova/+/818559
15:39:10 sean-k-mooney gibi: older qemu did not support restart the detach but did not raise an error then qemu started enforcing it
15:40:15 gibi sean-k-mooney: yeah, sorry, s/libvirt/qemu.
15:45:10 elodilles melwitt: whenever you have time, could you review this patch? https://review.opendev.org/c/openstack/nova/+/805628
15:45:33 frickler sean-k-mooney: gibi: but then it sounds to me that the real fix would still be to make nova not retry the detach, just wait longer?
15:45:47 elodilles melwitt: i think it would reduce the number of rechecks in wallaby and victoria if it merges (and its devstack part)
15:46:56 gibi frickler: if the detach happens while the cirros is booting then the guest OS never releases the device
15:47:06 gibi so right now waiting more is not an option
15:47:15 gibi but in general I agree to remove the retry loop from nova
15:47:28 gibi as it is pointless after qemu starts rejecting the retry
15:47:49 sean-k-mooney really the jobs should wait for the instance to be pingable/sshable
15:47:53 sean-k-mooney and only detach then
15:48:11 sean-k-mooney and nova shoudl not retry if the detach fails and just have the client retry
15:48:35 sean-k-mooney clieht beign tempest or enduser if a retry is needed
15:50:04 gibi sean-k-mooney: yepp
15:50:40 gibi on the other hand if ever a detach is issue by the client right after a boot then that detach will time out in nova, but I'm not sure it ever time outs in qemu
15:51:15 gibi so in that case a client retry will not help either
15:51:22 sean-k-mooney we might be abel to expicitly cancel the job in qemu
15:51:24 gibi as qemu will say that the detach is in progress
15:51:26 sean-k-mooney when we time out
15:51:57 opendevreview Jonathan Race proposed openstack/nova master: object/notification for Adds Pick guest CPU architecture based on host arch in libvirt driver support https://review.opendev.org/c/openstack/nova/+/828369
15:51:57 opendevreview Jonathan Race proposed openstack/nova master: driver/secheduler/docs for Adds Pick guest CPU architecture based on host arch in libvirt driver support https://review.opendev.org/c/openstack/nova/+/822053
15:51:58 opendevreview Jonathan Race proposed openstack/nova master: zuul-job for Adds Pick guest CPU architecture based on host arch in libvirt driver support https://review.opendev.org/c/openstack/nova/+/828372
15:54:11 gibi sean-k-mooney: at least the doc did not mention a way to cancel via libvirt https://libvirt.org/html/libvirt-libvirt-domain.html#virDomainDetachDeviceFlags
15:55:47 sean-k-mooney we could always issue a detach again :) since that will do an abort in qemu
15:56:22 sean-k-mooney but ya we shoudl aske the libirt folks although im on PTO today so im going to drop off irc again soon
15:56:47 sean-k-mooney so kashyap maybe you could folow up and see if there is a way to abort the detach or pass a timeout to qemu via libfirt
15:57:58 gibi sean-k-mooney: if we attach the detach again qemu reject it
15:58:04 gibi sean-k-mooney: if we issue the detach again qemu reject it
15:58:29 gibi as the pervious one is still ongoing
15:59:08 sean-k-mooney yep
15:59:38 sean-k-mooney but we could catch the error
15:59:58 gibi yepp, but that does not make the device actually detached :D
16:00:01 sean-k-mooney if it say devcice not found well presumabel it finsihed before we sent the detach after the time out
16:00:27 gibi sean-k-mooney: it say detach is ongoing
16:00:44 sean-k-mooney right but the second detach will abort the detach
16:00:56 sean-k-mooney that is the new behavior in qemu
16:01:15 gibi really?
16:01:24 gibi I've only checked the first two detach
16:01:33 gibi so you say the 3rd returns device not found?
16:02:34 sean-k-mooney no
16:02:52 sean-k-mooney im saying the second detach that return "detach is ongoing" cause qemu to abort the detach
16:03:12 sean-k-mooney at lest that is what i was told was the new behavior
16:09:10 gmann gibi: ack, thanks. I will check that tempest patches.
16:10:27 gibi sean-k-mooney: qemu reject each 7 retries with the message the the unplug is in progress https://paste.opendev.org/show/bW5wXCyH5em5tNI34zwV/
16:10:52 gibi sean-k-mooney: I don't think the first detach job was abborted by the second detach
16:11:06 gibi gmann: they are WIP

Earlier   Later