Earlier  
Posted Nick Remark
#openstack-nova - 2021-10-26
18:34:25 sean-k-mooney anyway going to call it a night o/
19:24:41 opendevreview Ilya Popov proposed openstack/nova master: Fix to use NUMA cell with more free memory first https://review.opendev.org/c/openstack/nova/+/805649
20:22:56 opendevreview Ilya Popov proposed openstack/nova master: Fix to use NUMA cell with more free memory first https://review.opendev.org/c/openstack/nova/+/805649
20:31:30 opendevreview Ilya Popov proposed openstack/nova master: Fix to use NUMA cell with free resources first https://review.opendev.org/c/openstack/nova/+/805649
#openstack-nova - 2021-10-27
08:56:40 gibi good day nova
09:17:47 bauzas hola folks
09:20:41 gibi o/
09:20:54 gibi -another day in downstream land-
09:34:50 opendevreview Balazs Gibizer proposed openstack/nova master: Add a WA flag waiting for vif-plugged event during reboot https://review.opendev.org/c/openstack/nova/+/813419
10:58:44 gibi stephenfin: do you recall why you needed to configure both db in a single Database fixture in https://review.opendev.org/c/openstack/nova/+/799526/5/nova/tests/fixtures/nova.py#612 ?
10:59:29 gibi I think our tests are using a separate fixture instance for each db (api, main)
11:00:18 stephenfin gibi: I'm not 100% sure but my guess is that I don't, and that was simply for expediency/laziness :) If SESSION_CONFIGURED became a mapping of DB type to "is configured" bool, we probably wouldn't need that
11:00:28 stephenfin *we don't
11:01:00 gibi I see so we had a single global but with two dbs to configure
11:01:07 stephenfin Yeah, I think so
11:02:48 gibi OK, if that is the only reason then I think I have a way to remove that global flag (based on melwitt's idea) with patch_factory from oslo_db
11:03:15 gibi it is no pretty confusing that we have to Database fixture intantiated one for main and one for api but the first one configures both db
11:03:38 gibi s/no/now
11:03:47 gibi /to/two
11:03:50 gibi /o\
11:04:20 stephenfin yeah, tbc it could be more complicated than that but I really doubt it
11:08:07 gibi yeah, lets see if my idea works
11:12:21 sean-k-mooney stephenfin: since your about here an easy one for you https://review.opendev.org/c/openstack/nova/+/811947 think we can get that over the line
11:13:43 stephenfin sure, will look now
11:14:05 sean-k-mooney thanks :)
11:19:44 sean-k-mooney stephenfin: based on the ptg discussion woudl you mind removing your -2 on https://review.opendev.org/c/openstack/nova/+/804292 im going to rebase that and the autopep8 one shortly
13:05:29 frickler kashyap: I didn't make progress with reproduction without nova yet, so I created https://gitlab.com/qemu-project/qemu/-/issues/693 for now. let me know if you need additional data there
13:08:38 kashyap frickler: Thanks for the report. A quick one is: were you using nested setup, or was this DevStack instance on a baremetal host (<shudder>)?
13:09:07 kashyap A thumb-rule is to always explicitly state so if you're using a nested setup
13:09:47 kashyap frickler: Can you edit the report to state that "deploy DevStack in a VM?" So that an unsuspecting dev won't run it on their baremetal laptop and wreak havoc...
13:10:21 kashyap I'll add a quick comment there, actualy
13:17:19 kashyap Done
13:17:52 kashyap frickler: I'll check about it w/ a TCG dev
13:25:41 frickler kashyap: yes, nested is correct, I added that to the description. though I could also duplicate on a baremetal host if you assume that it would behave differently
13:37:59 kashyap frickler: No, no need for baremetal. VMs are best. Can you also post the QEMU command-line of the DevStack VM itself? (The level-1 VM)
13:39:42 frickler kashyap: no, I have no admin access to the cloud it is running on. I'm assuming it will essentially look the same, though, just with accel=kvm
13:43:02 kashyap Hmm, not sure if it'll be that same w/ accel=kvm. The details would change quite a bit. The "host" (guest hypervisor) setup can determine the guest behaviour here a lot.
13:43:19 kashyap That's one of the questions I'd expect from a TCG dev
13:45:38 frickler kashyap: o.k., I'll try running cirros locally without devstack in between, that would give the simplest setup in the end
13:47:06 kashyap frickler: Sure; yeah, that'd be the best. The shorter the route to the reproducer, the more likely we can get to the root cause
13:47:13 kashyap frickler: Thanks for all the testing! It's a pain, I Know
13:47:24 kashyap s/K/k/
14:14:01 frickler kashyap: that went easier than I expected, updated the issue
14:19:49 bauzas gibi: sean-k-mooney: I think we said https://bugs.launchpad.net/nova/+bug/1947753 is valid during our PTG, right?
14:21:22 gibi bauzas: I don't remember discussing this
14:21:35 bauzas gibi: sorry you're maybe right
14:21:42 gibi what I see is that in the bug they evacuate instances without restarting the compute node
14:21:57 bauzas I originally thought this was about evacuate/evacuateback/evacuate
14:22:33 bauzas adding a comment
14:23:00 gibi so far we said that you can only evacuate if you make sure that the compute is dead
14:23:52 gibi in the bug case the compute was halted / stuck, the heartbeat was missing so the service was considered down, nova allowed evacuation, then the compute recovered without the nova-compute service restarted
14:24:10 gibi nova-compute only cleans up evacuated instance during init_host but does not do it periodically
14:24:45 gibi so in this case the evacuated instance was not cleaned up on the source leading to duplicated instances causing corruption
14:25:55 gibi option a) change nova-compute to clean up evacuated instance in a periodic
14:26:37 gibi option b) change the evac API to only allow evacuation if the compute is forced down (meaning the admin mades sure the host is fenced)
14:27:14 gibi option c) declare the current bug as user error as the nova-compute was not restarted as part of the compute node recovery
14:33:29 bauzas gibi: I wrote a large comment on the bvug
14:33:51 bauzas I think CERN is triggering evacuations before verifying the host status
14:34:32 sean-k-mooney bauzas: i dont think we talked baout this either
14:34:33 bauzas as I said, I feel healthchecks can help them getting a better decision-making about whether they need to evacuate or not
14:34:49 sean-k-mooney bauzas: we talk about a related issue with allcoations tha talso impact evacuate
14:35:03 bauzas sean-k-mooney: yeah, my confusion, I originally thought it was about the back-and-forth about evacuate we discussed for pain points
14:35:05 sean-k-mooney e.g. if for any reason we oversubscie the allcoations then we cant evacuate
14:36:15 sean-k-mooney i have not read it fully but it sound like they are not properly fencing the node an ensuring the vm is not running
14:36:26 sean-k-mooney before evacuating if they are having apllciation data currption
14:41:46 bauzas that's literrally what I wrote.
14:41:58 bauzas anyway, moving to a new bug.
14:42:43 sean-k-mooney ack as i said have not finsihed reading the bug description or comments so glad we agree :)
14:51:50 kashyap frickler: Ah-ha, noted. Good news: there's already some response from two QEMU devs, with a patch in a newer version :)
14:55:27 frickler kashyap: yeah I just responded, but I didn't see the patch reference. going from 32M to 1G really sounds a bit excessive, would be good to be able to tune that
14:55:29 kashyap (Well, I don't quite think it's "good news" ...)
14:55:47 kashyap frickler: Sorry, I was referring to the commit that DanPB pointed out - https://gitlab.com/qemu-project/qemu/-/commit/600e17b26
14:56:39 frickler kashyap: ah, yes, that seems to be the patch that triggers this, I though you were referring to a fix in a recent commit
14:56:45 kashyap frickler: Yeah, that increase is a tad too much.
14:56:53 kashyap frickler: Yes, poor phrasing on my part.
14:56:59 opendevreview Balazs Gibizer proposed openstack/nova master: Remove SESSION_CONFIGURED global from DB fixture https://review.opendev.org/c/openstack/nova/+/815689
14:57:42 frickler kashyap: otoh that also is likely to explain why tests seemed to be going faster on Bullseye than on Focal
14:58:18 opendevreview Balazs Gibizer proposed openstack/nova master: Refactor Database fixture https://review.opendev.org/c/openstack/nova/+/815690
14:59:22 kashyap frickler: Interesting; what tests are going faster?
14:59:35 opendevreview Balazs Gibizer proposed openstack/nova master: Fix interference in db unit test https://review.opendev.org/c/openstack/nova/+/814735
15:00:06 gibi stephenfin: ^^ here is the removal of the global SESSION_CONFIGURED flag from the DB fixture and some extra :D
15:02:37 frickler kashyap: I didn't check in particular, but the whole tempest-full job with --serial on Debian doesn't take much longer than with the default (-c 4 I think) on Focal
15:02:54 kashyap I see.
15:05:17 gibi melwitt: thanks a lot for the help exlaning the global db transaction factory situation. I used your info to actually remove SESSION_CONFIGURED from our fixture along the the unit test fixes
15:05:21 kashyap frickler: So, it is tunable via command-line, but it's not wired up in libvirt yet, though.
15:05:57 kashyap frickler: See the option: -accel=tcg,tb-size=$value_in_MiB
15:06:10 kashyap "tb-size" in the man page
15:10:32 frickler kashyap: as long as libvirt doesn't support it, I fear that won't help much. might be good to cap it to something like 50% of the VM memory
15:11:06 kashyap frickler: Right; libvirt just didn't wire it up ... we can meanwhile do a nasty hack of uploading a QEMU binary to the CI system w/ this param tweaked
15:11:52 kashyap frickler: Do you have the appetite to file a libvirt upstream RFE? (Then I can clone it downstream, and get it triaged)
15:13:15 frickler kashyap: I think I'll do a local test with a reduced default tb-size first in order to be certain that that's the cause. but not before tomorrow
15:13:25 kashyap Right, no rush at all
15:14:49 gibi bauzas: replied in https://bugs.launchpad.net/nova/+bug/1947753 I think _destroy_evacuated_instances is not called periodically
15:16:45 kashyap frickler: So I see that someone else has raised this upstream last year: https://lists.gnu.org/archive/html/qemu-devel/2020-07/msg05235.html (TB Cache size grows out of control with qemu 5.0)
15:22:08 bauzas gibi: indeed, only when restarting
15:22:16 bauzas did I said the other way ?
15:22:57 kashyap frickler: So, this worked for me:
15:22:57 kashyap -machine q35
15:22:57 kashyap -accel tcg,tb-size=256
15:23:03 kashyap (As an example)

Earlier   Later