Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
13:13:52 gibi dansmith: thanks for looking into it
13:15:04 dansmith gibi: yeah, some of the other failures are clearly after we have already attached and detached once, so I'm sure they're not sshable related,
13:15:16 dansmith but I haven't dug into the rescue ones yet
13:16:13 opendevreview Sylvain Bauza proposed openstack/nova master: Add a new policy for cold-migrate with host https://review.opendev.org/c/openstack/nova/+/881562
13:16:53 bauzas dansmith: I can also help for the rescue action
13:17:01 bauzas so you can look at the ceph failures
13:17:15 bauzas I'm done with my small feature now
13:18:02 dansmith bauzas: no, they're related
13:18:30 dansmith bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job
13:19:07 bauzas dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6
13:19:16 bauzas -ETOOMANYFAILURES
13:19:36 dansmith bauzas: right that's what I'm talking about
13:20:21 bauzas hmmm, then I'm lost out of context
13:20:36 sean-k-mooney so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point
13:20:55 sean-k-mooney basically im almost at a point where context switing for a day or two would be ok
13:20:57 bauzas I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob
13:21:19 dansmith my patch for the ceph job has gotten us almost all of the way there,
13:21:52 dansmith just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on
13:22:09 dansmith if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases
13:23:27 bauzas dansmith: ok, you're prioritary (because of the blocking gate), so cool with me
13:24:10 dansmith the gate is unblocked, so I think the pressure is off for the moment, but we obviously have to get this resolved ASAP
13:24:35 dansmith bauzas: maybe you could re+W these: https://review.opendev.org/c/openstack/nova/+/880632/5
13:24:45 dansmith the bottom patch already merged, they just need a kick to get into the gate
13:25:26 bauzas dansmith: I did this kickass
13:26:45 dansmith thanks
13:27:27 bauzas fwiw, I'll send a status email saying then that the gate is back
13:28:06 bauzas are we sure all the py38-removal releases were reverted ?
13:29:45 dansmith well, we're passing jobs so, I think that's all we need to know
13:29:56 frickler pretty sure oslo bumps were not reverted yet. but only oslo.db was released
13:30:02 dansmith there have been other releases but they are not yet in requirements, so we're kinda on the edge
13:30:06 dansmith yeah that^
14:44:59 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
15:01:48 dansmith yeah, so the rescue tests work for me locally, which I guess is good
15:01:56 dansmith however, the teardown does fail for me
15:02:29 dansmith which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation
15:03:50 dansmith sorry, s/rescue/rebuild/
15:07:13 bauzas dansmith: bravo
15:09:47 dansmith locally that includes the validations, which is disabled in the job (for some reason) so I pushed up a patch to set those to enabled,
15:10:02 dansmith which will both likely slow it down but also maybe catch some other issues
15:10:53 dansmith that causes us to at least wait for the server to be created and sshable in some of the other cases, before we might try to attach a volume, but like I say, the failures I'm seeing don't seem like they'd be impacted there
15:11:15 dansmith and the regular jobs have those enabled, and we see detach failures there still
15:11:17 dansmith albeit not 100%
15:23:34 dansmith bauzas: this needs re-+W as well: https://review.opendev.org/c/openstack/nova/+/880633
15:23:39 dansmith it's about to pass its recheck
15:24:07 bauzas dansmith: done
15:24:14 dansmith thanks
15:37:44 opendevreview Merged openstack/nova master: Remove silent failure to find a node on rebuild https://review.opendev.org/c/openstack/nova/+/880632
15:38:15 bauzas :)
15:53:56 dansmith well, the distro-based ceph job just passed with validations turned on
15:54:20 dansmith the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that
16:09:36 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:19:19 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:21:17 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:24:28 dansmith okay it timed out, but legitimately, it was making progress just taking much longer because of the validations I think
16:24:48 dansmith so we might need to trim that job down again or we can try with the concurrency upped to 2
16:26:41 dansmith ah, the base job is not concurrency=1, so that's probably the difference there.. might be getting close then!
16:29:18 dansmith gouthamr: ^
16:52:31 opendevreview ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773
17:01:43 opendevreview Merged openstack/nova master: Stop ignoring missing compute nodes in claims https://review.opendev.org/c/openstack/nova/+/880633
17:46:12 dansmith eharney: gouthamr: the cephadm run just passed.. the distro package one failed with errors that look very different than before, but related to ImageBusy type things
17:46:52 dansmith the difference here is enabling validations on both, which means we try to ssh to instances before we do certain things to them, which helps a lot with race conditions. I dunno why those were disabled on these jobs before, but we thought they were enabled
17:47:47 dansmith so my plan is to: (1) switch the cephadm job to be the voting primary job and make the distro one non-voting, (2) set the nova DNM to run against cephadm, and get another run of everything
17:48:31 dansmith is there anything else you want to change before we could/should merge this? the cephfs job is still on focal (because it's defined elsewhere) and thus failing because of packages, but I could hard-code that to jammy here to "fix" that if you want
17:48:44 dansmith it's non-voting so we should be able to switch it externally separately, but let me know what you want
17:50:22 dansmith oh another change on the cephadm job was removing the concurrency limit, which was causing it to run too slow to finish when validations was turned on. We need that to run in parallel anyway, so it's good that it passed in that form
17:53:58 dansmith the cephadm job also finished pretty quickly, which is also a good sign that it was pretty healthy
18:15:11 dansmith ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think
18:16:37 sean-k-mooney that would do it
18:16:55 sean-k-mooney it makes snese why you are siing rbd busy messages
18:17:21 dansmith yeah
18:17:48 sean-k-mooney is this with the mariadb and ceph options to reduce mememory set
18:18:03 sean-k-mooney and idealy the job configred for 8GB of swap?
18:18:31 dansmith it does include the mysql and ceph tweaks, not sure about swap,
18:18:45 sean-k-mooney i think we default to 2GB instead of 8
18:18:49 dansmith but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else
18:19:01 sean-k-mooney ack
18:19:03 dansmith I know
18:19:27 dansmith I'm waiting for the nova one to finish and see if it's the same,
18:19:38 dansmith but also to switch it to cephadm and get another set of data
18:19:49 dansmith but this is the first time I've ever seen the cephadm job pass, so I'm encouraged
18:20:46 sean-k-mooney if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm
18:21:29 dansmith we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now
18:21:46 dansmith the cephadm job ran at =4 at top speed with no failures this time
18:22:00 sean-k-mooney ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph
18:22:16 sean-k-mooney cool
18:23:09 sean-k-mooney your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log
18:23:19 dansmith yeah
18:23:30 dansmith I want the logs so I can look while I run it against cephadm as the base
18:23:47 dansmith hoping it's also the same issue as the other one because they look very similar, just watching them
18:23:58 sean-k-mooney an hour and 19 minuts is not bad
18:24:08 sean-k-mooney https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4
18:24:22 dansmith right, it's super fast actually
18:24:33 sean-k-mooney its also singel node which helps
18:24:44 dansmith so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :)
18:24:50 dansmith compared to the distro jobs
18:25:12 sean-k-mooney well going to python 3.10 wil add about 30% to the openstack performace
18:25:20 sean-k-mooney vs 3.8 i think
18:25:29 sean-k-mooney oh
18:25:44 sean-k-mooney you are comparing to disto packages on 22.04 not focal
18:25:51 dansmith yeah, so that might be helping the speed when it doesn't fail

Earlier   Later