Earlier  
Posted Nick Remark
#openstack-nova - 2021-10-27
13:43:02 kashyap Hmm, not sure if it'll be that same w/ accel=kvm. The details would change quite a bit. The "host" (guest hypervisor) setup can determine the guest behaviour here a lot.
13:43:19 kashyap That's one of the questions I'd expect from a TCG dev
13:45:38 frickler kashyap: o.k., I'll try running cirros locally without devstack in between, that would give the simplest setup in the end
13:47:06 kashyap frickler: Sure; yeah, that'd be the best. The shorter the route to the reproducer, the more likely we can get to the root cause
13:47:13 kashyap frickler: Thanks for all the testing! It's a pain, I Know
13:47:24 kashyap s/K/k/
14:14:01 frickler kashyap: that went easier than I expected, updated the issue
14:19:49 bauzas gibi: sean-k-mooney: I think we said https://bugs.launchpad.net/nova/+bug/1947753 is valid during our PTG, right?
14:21:22 gibi bauzas: I don't remember discussing this
14:21:35 bauzas gibi: sorry you're maybe right
14:21:42 gibi what I see is that in the bug they evacuate instances without restarting the compute node
14:21:57 bauzas I originally thought this was about evacuate/evacuateback/evacuate
14:22:33 bauzas adding a comment
14:23:00 gibi so far we said that you can only evacuate if you make sure that the compute is dead
14:23:52 gibi in the bug case the compute was halted / stuck, the heartbeat was missing so the service was considered down, nova allowed evacuation, then the compute recovered without the nova-compute service restarted
14:24:10 gibi nova-compute only cleans up evacuated instance during init_host but does not do it periodically
14:24:45 gibi so in this case the evacuated instance was not cleaned up on the source leading to duplicated instances causing corruption
14:25:55 gibi option a) change nova-compute to clean up evacuated instance in a periodic
14:26:37 gibi option b) change the evac API to only allow evacuation if the compute is forced down (meaning the admin mades sure the host is fenced)
14:27:14 gibi option c) declare the current bug as user error as the nova-compute was not restarted as part of the compute node recovery
14:33:29 bauzas gibi: I wrote a large comment on the bvug
14:33:51 bauzas I think CERN is triggering evacuations before verifying the host status
14:34:32 sean-k-mooney bauzas: i dont think we talked baout this either
14:34:33 bauzas as I said, I feel healthchecks can help them getting a better decision-making about whether they need to evacuate or not
14:34:49 sean-k-mooney bauzas: we talk about a related issue with allcoations tha talso impact evacuate
14:35:03 bauzas sean-k-mooney: yeah, my confusion, I originally thought it was about the back-and-forth about evacuate we discussed for pain points
14:35:05 sean-k-mooney e.g. if for any reason we oversubscie the allcoations then we cant evacuate
14:36:15 sean-k-mooney i have not read it fully but it sound like they are not properly fencing the node an ensuring the vm is not running
14:36:26 sean-k-mooney before evacuating if they are having apllciation data currption
14:41:46 bauzas that's literrally what I wrote.
14:41:58 bauzas anyway, moving to a new bug.
14:42:43 sean-k-mooney ack as i said have not finsihed reading the bug description or comments so glad we agree :)
14:51:50 kashyap frickler: Ah-ha, noted. Good news: there's already some response from two QEMU devs, with a patch in a newer version :)
14:55:27 frickler kashyap: yeah I just responded, but I didn't see the patch reference. going from 32M to 1G really sounds a bit excessive, would be good to be able to tune that
14:55:29 kashyap (Well, I don't quite think it's "good news" ...)
14:55:47 kashyap frickler: Sorry, I was referring to the commit that DanPB pointed out - https://gitlab.com/qemu-project/qemu/-/commit/600e17b26
14:56:39 frickler kashyap: ah, yes, that seems to be the patch that triggers this, I though you were referring to a fix in a recent commit
14:56:45 kashyap frickler: Yeah, that increase is a tad too much.
14:56:53 kashyap frickler: Yes, poor phrasing on my part.
14:56:59 opendevreview Balazs Gibizer proposed openstack/nova master: Remove SESSION_CONFIGURED global from DB fixture https://review.opendev.org/c/openstack/nova/+/815689
14:57:42 frickler kashyap: otoh that also is likely to explain why tests seemed to be going faster on Bullseye than on Focal
14:58:18 opendevreview Balazs Gibizer proposed openstack/nova master: Refactor Database fixture https://review.opendev.org/c/openstack/nova/+/815690
14:59:22 kashyap frickler: Interesting; what tests are going faster?
14:59:35 opendevreview Balazs Gibizer proposed openstack/nova master: Fix interference in db unit test https://review.opendev.org/c/openstack/nova/+/814735
15:00:06 gibi stephenfin: ^^ here is the removal of the global SESSION_CONFIGURED flag from the DB fixture and some extra :D
15:02:37 frickler kashyap: I didn't check in particular, but the whole tempest-full job with --serial on Debian doesn't take much longer than with the default (-c 4 I think) on Focal
15:02:54 kashyap I see.
15:05:17 gibi melwitt: thanks a lot for the help exlaning the global db transaction factory situation. I used your info to actually remove SESSION_CONFIGURED from our fixture along the the unit test fixes
15:05:21 kashyap frickler: So, it is tunable via command-line, but it's not wired up in libvirt yet, though.
15:05:57 kashyap frickler: See the option: -accel=tcg,tb-size=$value_in_MiB
15:06:10 kashyap "tb-size" in the man page
15:10:32 frickler kashyap: as long as libvirt doesn't support it, I fear that won't help much. might be good to cap it to something like 50% of the VM memory
15:11:06 kashyap frickler: Right; libvirt just didn't wire it up ... we can meanwhile do a nasty hack of uploading a QEMU binary to the CI system w/ this param tweaked
15:11:52 kashyap frickler: Do you have the appetite to file a libvirt upstream RFE? (Then I can clone it downstream, and get it triaged)
15:13:15 frickler kashyap: I think I'll do a local test with a reduced default tb-size first in order to be certain that that's the cause. but not before tomorrow
15:13:25 kashyap Right, no rush at all
15:14:49 gibi bauzas: replied in https://bugs.launchpad.net/nova/+bug/1947753 I think _destroy_evacuated_instances is not called periodically
15:16:45 kashyap frickler: So I see that someone else has raised this upstream last year: https://lists.gnu.org/archive/html/qemu-devel/2020-07/msg05235.html (TB Cache size grows out of control with qemu 5.0)
15:22:08 bauzas gibi: indeed, only when restarting
15:22:16 bauzas did I said the other way ?
15:22:57 kashyap frickler: So, this worked for me:
15:22:57 kashyap -machine q35
15:22:57 kashyap -accel tcg,tb-size=256
15:23:03 kashyap (As an example)
15:25:07 gibi bauzas: at least I understood this sentence that way "Either way, if the service continues to run, it verifies the evacuation status periodically and deletes the host."
15:27:04 gibi bauzas: btw, about https://bugs.launchpad.net/nova/+bug/1947687 I cannot formulate a logstash signature it seems that this error happens in a lot of cases when no test cases are failing so I get a lot of false positives
15:27:21 bauzas gibi: okay, then my brain fucked
15:27:33 kashyap frickler: For reference, a minimal command-line:
15:27:35 kashyap $> qemu-kvm -display none -cpu Nehalem -no-user-config \ -machine q35 \ -accel tcg,tb-size=256 \ -nodefaults -m 2048 -serial stdio \ -drive file=/export/vm1.qcow2,format=qcow2,if=virtio
15:27:48 kashyap (Ugh, line-breaks are broken, but you see what I mean)
15:28:49 bauzas gibi: ack for the logstash thing, no worries
15:34:57 frickler kashyap: thx, added a comment to the issue, seems the libvirt path is really the most promising one
15:39:22 kashyap frickler: Definitely. Please file the RFE (and post a link to me, Bz Ccs will take me slower to process) when you can
15:39:29 kashyap Thanks for the patience :)
15:49:31 melwitt bauzas: hi, could you pls take a look at these train backports when you get a chance? someone posted a comment on the top patch yesterday indicating they are awaiting merge of the fixes https://review.opendev.org/q/topic:%2522bug/1927677%2522+branch:stable/train+status:open
15:49:36 opendevreview Balazs Gibizer proposed openstack/nova stable/pike: Add a WA flag waiting for vif-plugged event during reboot https://review.opendev.org/c/openstack/nova/+/813437
15:49:50 bauzas melwitt: ack, doing it now
15:50:03 melwitt thanks!
15:51:55 bauzas melwitt: I already reviewed them but forgot to submit, my bad
15:52:05 bauzas now this is fixed.
15:59:21 melwitt bauzas: a-ha, thank you
16:13:36 stephenfin gibi: question on https://review.opendev.org/c/openstack/nova/+/815690
16:13:43 stephenfin please excuse my ignorance
16:14:18 gibi lookgin
16:23:38 gibi stephenfin: you are right something is fishy there
16:23:53 gibi I have to go back and poke that test to understand what is happening
16:32:40 opendevreview Artom Lifshitz proposed openstack/nova master: DNM:goat https://review.opendev.org/c/openstack/nova/+/815705
16:32:41 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: goat 2 https://review.opendev.org/c/openstack/nova/+/815706
16:32:41 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: goat3 https://review.opendev.org/c/openstack/nova/+/815707
16:34:08 gibi gmann: I did the change you requested in https://review.opendev.org/c/openstack/tempest/+/809168/comment/35477e85_10754ba5/ but I wondering why we need that indirection
16:36:01 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: goat 2 https://review.opendev.org/c/openstack/nova/+/815706
16:36:01 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: goat3 https://review.opendev.org/c/openstack/nova/+/815707
17:16:06 em_ are there currently issues with xena nova and (debian) cloud images? Neither my ssh keys nor the admin password seems to get applied. Any open bugs (maybe libvirt issues or kernel related?) using 5.10 debian bullseye as host, kolla xena (ubuntu/source) as libvirt
17:18:47 opendevreview Balazs Gibizer proposed openstack/nova master: Refactor Database fixture https://review.opendev.org/c/openstack/nova/+/815690
17:19:16 gibi stephenfin: you had a valid point, fixed it ^^
17:20:05 opendevreview Balazs Gibizer proposed openstack/nova master: Fix interference in db unit test https://review.opendev.org/c/openstack/nova/+/814735
17:21:00 gmann gibi: replied, basically Tempest test the services with what is configured to test instead of 'test what cloud/service APIs return'
17:22:09 gmann autodetecting service features/extensions to what to test can hide the error.
17:26:41 gibi gmann: OK, I think I got it. Does devstack needs to be changed to generate the extension name to the tempest config/
17:26:44 gibi ?

Earlier   Later