Earlier  
Posted Nick Remark
#openstack-nova - 2021-10-27
11:13:43 stephenfin sure, will look now
11:14:05 sean-k-mooney thanks :)
11:19:44 sean-k-mooney stephenfin: based on the ptg discussion woudl you mind removing your -2 on https://review.opendev.org/c/openstack/nova/+/804292 im going to rebase that and the autopep8 one shortly
13:05:29 frickler kashyap: I didn't make progress with reproduction without nova yet, so I created https://gitlab.com/qemu-project/qemu/-/issues/693 for now. let me know if you need additional data there
13:08:38 kashyap frickler: Thanks for the report. A quick one is: were you using nested setup, or was this DevStack instance on a baremetal host (<shudder>)?
13:09:07 kashyap A thumb-rule is to always explicitly state so if you're using a nested setup
13:09:47 kashyap frickler: Can you edit the report to state that "deploy DevStack in a VM?" So that an unsuspecting dev won't run it on their baremetal laptop and wreak havoc...
13:10:21 kashyap I'll add a quick comment there, actualy
13:17:19 kashyap Done
13:17:52 kashyap frickler: I'll check about it w/ a TCG dev
13:25:41 frickler kashyap: yes, nested is correct, I added that to the description. though I could also duplicate on a baremetal host if you assume that it would behave differently
13:37:59 kashyap frickler: No, no need for baremetal. VMs are best. Can you also post the QEMU command-line of the DevStack VM itself? (The level-1 VM)
13:39:42 frickler kashyap: no, I have no admin access to the cloud it is running on. I'm assuming it will essentially look the same, though, just with accel=kvm
13:43:02 kashyap Hmm, not sure if it'll be that same w/ accel=kvm. The details would change quite a bit. The "host" (guest hypervisor) setup can determine the guest behaviour here a lot.
13:43:19 kashyap That's one of the questions I'd expect from a TCG dev
13:45:38 frickler kashyap: o.k., I'll try running cirros locally without devstack in between, that would give the simplest setup in the end
13:47:06 kashyap frickler: Sure; yeah, that'd be the best. The shorter the route to the reproducer, the more likely we can get to the root cause
13:47:13 kashyap frickler: Thanks for all the testing! It's a pain, I Know
13:47:24 kashyap s/K/k/
14:14:01 frickler kashyap: that went easier than I expected, updated the issue
14:19:49 bauzas gibi: sean-k-mooney: I think we said https://bugs.launchpad.net/nova/+bug/1947753 is valid during our PTG, right?
14:21:22 gibi bauzas: I don't remember discussing this
14:21:35 bauzas gibi: sorry you're maybe right
14:21:42 gibi what I see is that in the bug they evacuate instances without restarting the compute node
14:21:57 bauzas I originally thought this was about evacuate/evacuateback/evacuate
14:22:33 bauzas adding a comment
14:23:00 gibi so far we said that you can only evacuate if you make sure that the compute is dead
14:23:52 gibi in the bug case the compute was halted / stuck, the heartbeat was missing so the service was considered down, nova allowed evacuation, then the compute recovered without the nova-compute service restarted
14:24:10 gibi nova-compute only cleans up evacuated instance during init_host but does not do it periodically
14:24:45 gibi so in this case the evacuated instance was not cleaned up on the source leading to duplicated instances causing corruption
14:25:55 gibi option a) change nova-compute to clean up evacuated instance in a periodic
14:26:37 gibi option b) change the evac API to only allow evacuation if the compute is forced down (meaning the admin mades sure the host is fenced)
14:27:14 gibi option c) declare the current bug as user error as the nova-compute was not restarted as part of the compute node recovery
14:33:29 bauzas gibi: I wrote a large comment on the bvug
14:33:51 bauzas I think CERN is triggering evacuations before verifying the host status
14:34:32 sean-k-mooney bauzas: i dont think we talked baout this either
14:34:33 bauzas as I said, I feel healthchecks can help them getting a better decision-making about whether they need to evacuate or not
14:34:49 sean-k-mooney bauzas: we talk about a related issue with allcoations tha talso impact evacuate
14:35:03 bauzas sean-k-mooney: yeah, my confusion, I originally thought it was about the back-and-forth about evacuate we discussed for pain points
14:35:05 sean-k-mooney e.g. if for any reason we oversubscie the allcoations then we cant evacuate
14:36:15 sean-k-mooney i have not read it fully but it sound like they are not properly fencing the node an ensuring the vm is not running
14:36:26 sean-k-mooney before evacuating if they are having apllciation data currption
14:41:46 bauzas that's literrally what I wrote.
14:41:58 bauzas anyway, moving to a new bug.
14:42:43 sean-k-mooney ack as i said have not finsihed reading the bug description or comments so glad we agree :)
14:51:50 kashyap frickler: Ah-ha, noted. Good news: there's already some response from two QEMU devs, with a patch in a newer version :)
14:55:27 frickler kashyap: yeah I just responded, but I didn't see the patch reference. going from 32M to 1G really sounds a bit excessive, would be good to be able to tune that
14:55:29 kashyap (Well, I don't quite think it's "good news" ...)
14:55:47 kashyap frickler: Sorry, I was referring to the commit that DanPB pointed out - https://gitlab.com/qemu-project/qemu/-/commit/600e17b26
14:56:39 frickler kashyap: ah, yes, that seems to be the patch that triggers this, I though you were referring to a fix in a recent commit
14:56:45 kashyap frickler: Yeah, that increase is a tad too much.
14:56:53 kashyap frickler: Yes, poor phrasing on my part.
14:56:59 opendevreview Balazs Gibizer proposed openstack/nova master: Remove SESSION_CONFIGURED global from DB fixture https://review.opendev.org/c/openstack/nova/+/815689
14:57:42 frickler kashyap: otoh that also is likely to explain why tests seemed to be going faster on Bullseye than on Focal
14:58:18 opendevreview Balazs Gibizer proposed openstack/nova master: Refactor Database fixture https://review.opendev.org/c/openstack/nova/+/815690
14:59:22 kashyap frickler: Interesting; what tests are going faster?
14:59:35 opendevreview Balazs Gibizer proposed openstack/nova master: Fix interference in db unit test https://review.opendev.org/c/openstack/nova/+/814735
15:00:06 gibi stephenfin: ^^ here is the removal of the global SESSION_CONFIGURED flag from the DB fixture and some extra :D
15:02:37 frickler kashyap: I didn't check in particular, but the whole tempest-full job with --serial on Debian doesn't take much longer than with the default (-c 4 I think) on Focal
15:02:54 kashyap I see.
15:05:17 gibi melwitt: thanks a lot for the help exlaning the global db transaction factory situation. I used your info to actually remove SESSION_CONFIGURED from our fixture along the the unit test fixes
15:05:21 kashyap frickler: So, it is tunable via command-line, but it's not wired up in libvirt yet, though.
15:05:57 kashyap frickler: See the option: -accel=tcg,tb-size=$value_in_MiB
15:06:10 kashyap "tb-size" in the man page
15:10:32 frickler kashyap: as long as libvirt doesn't support it, I fear that won't help much. might be good to cap it to something like 50% of the VM memory
15:11:06 kashyap frickler: Right; libvirt just didn't wire it up ... we can meanwhile do a nasty hack of uploading a QEMU binary to the CI system w/ this param tweaked
15:11:52 kashyap frickler: Do you have the appetite to file a libvirt upstream RFE? (Then I can clone it downstream, and get it triaged)
15:13:15 frickler kashyap: I think I'll do a local test with a reduced default tb-size first in order to be certain that that's the cause. but not before tomorrow
15:13:25 kashyap Right, no rush at all
15:14:49 gibi bauzas: replied in https://bugs.launchpad.net/nova/+bug/1947753 I think _destroy_evacuated_instances is not called periodically
15:16:45 kashyap frickler: So I see that someone else has raised this upstream last year: https://lists.gnu.org/archive/html/qemu-devel/2020-07/msg05235.html (TB Cache size grows out of control with qemu 5.0)
15:22:08 bauzas gibi: indeed, only when restarting
15:22:16 bauzas did I said the other way ?
15:22:57 kashyap -accel tcg,tb-size=256
15:22:57 kashyap -machine q35
15:22:57 kashyap frickler: So, this worked for me:
15:23:03 kashyap (As an example)
15:25:07 gibi bauzas: at least I understood this sentence that way "Either way, if the service continues to run, it verifies the evacuation status periodically and deletes the host."
15:27:04 gibi bauzas: btw, about https://bugs.launchpad.net/nova/+bug/1947687 I cannot formulate a logstash signature it seems that this error happens in a lot of cases when no test cases are failing so I get a lot of false positives
15:27:21 bauzas gibi: okay, then my brain fucked
15:27:33 kashyap frickler: For reference, a minimal command-line:
15:27:35 kashyap $> qemu-kvm -display none -cpu Nehalem -no-user-config \ -machine q35 \ -accel tcg,tb-size=256 \ -nodefaults -m 2048 -serial stdio \ -drive file=/export/vm1.qcow2,format=qcow2,if=virtio
15:27:48 kashyap (Ugh, line-breaks are broken, but you see what I mean)
15:28:49 bauzas gibi: ack for the logstash thing, no worries
15:34:57 frickler kashyap: thx, added a comment to the issue, seems the libvirt path is really the most promising one
15:39:22 kashyap frickler: Definitely. Please file the RFE (and post a link to me, Bz Ccs will take me slower to process) when you can
15:39:29 kashyap Thanks for the patience :)
15:49:31 melwitt bauzas: hi, could you pls take a look at these train backports when you get a chance? someone posted a comment on the top patch yesterday indicating they are awaiting merge of the fixes https://review.opendev.org/q/topic:%2522bug/1927677%2522+branch:stable/train+status:open
15:49:36 opendevreview Balazs Gibizer proposed openstack/nova stable/pike: Add a WA flag waiting for vif-plugged event during reboot https://review.opendev.org/c/openstack/nova/+/813437
15:49:50 bauzas melwitt: ack, doing it now
15:50:03 melwitt thanks!
15:51:55 bauzas melwitt: I already reviewed them but forgot to submit, my bad
15:52:05 bauzas now this is fixed.
15:59:21 melwitt bauzas: a-ha, thank you
16:13:36 stephenfin gibi: question on https://review.opendev.org/c/openstack/nova/+/815690
16:13:43 stephenfin please excuse my ignorance
16:14:18 gibi lookgin
16:23:38 gibi stephenfin: you are right something is fishy there
16:23:53 gibi I have to go back and poke that test to understand what is happening
16:32:40 opendevreview Artom Lifshitz proposed openstack/nova master: DNM:goat https://review.opendev.org/c/openstack/nova/+/815705

Earlier   Later