Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
07:56:33 bauzas gibi: ack, but I'm atm still digesting dansmith's and gouthamr efforts on ceph/Jammy jobs
08:38:25 gibi sure. I just wanted to be explict about the fact that I won't push a solution
08:51:41 bauzas gibi: I'll try to help this afternoon then
08:51:49 bauzas thanks for the findings again
11:31:16 ykarel Hi bauzas
11:31:29 ykarel reported https://blueprints.launchpad.net/nova/+spec/libvirt-tb-cache-size as discussed yesterday, can you please check
11:32:34 bauzas ykarel: ack, approving it then
11:34:00 bauzas and done
11:34:41 ykarel Thanks bauzas
11:34:55 bauzas np
11:37:23 sean-k-mooney ykarel: once we get the ceph/py38 issues resolved im fine with prioritsing getting this landed too
11:37:59 ykarel sean-k-mooney, ack
11:38:42 opendevreview Elod Illes proposed openstack/nova master: Drop py38 based zuul jobs https://review.opendev.org/c/openstack/nova/+/881339
12:53:06 dansmith gibi: I'm going to look into that today
13:09:22 dansmith gibi: we're failing that rescue test 100% of the time with the new ceph job I'm trying to get working, so perhaps we're losing that race more often there
13:09:48 sean-k-mooney odd
13:09:56 sean-k-mooney i wonder what makes rescue speciel
13:10:19 dansmith oh we're failing plenty of other things 100% of the time
13:10:27 dansmith but yeah, I don't know
13:10:31 dansmith I'm working on it today
13:13:33 gibi I think something is wrongly set up in the rescuce tests regarding the sshable waiter
13:13:43 gibi hence the more frequent failure there
13:13:52 gibi dansmith: thanks for looking into it
13:15:04 dansmith gibi: yeah, some of the other failures are clearly after we have already attached and detached once, so I'm sure they're not sshable related,
13:15:16 dansmith but I haven't dug into the rescue ones yet
13:16:13 opendevreview Sylvain Bauza proposed openstack/nova master: Add a new policy for cold-migrate with host https://review.opendev.org/c/openstack/nova/+/881562
13:16:53 bauzas dansmith: I can also help for the rescue action
13:17:01 bauzas so you can look at the ceph failures
13:17:15 bauzas I'm done with my small feature now
13:18:02 dansmith bauzas: no, they're related
13:18:30 dansmith bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job
13:19:07 bauzas dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6
13:19:16 bauzas -ETOOMANYFAILURES
13:19:36 dansmith bauzas: right that's what I'm talking about
13:20:21 bauzas hmmm, then I'm lost out of context
13:20:36 sean-k-mooney so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point
13:20:55 sean-k-mooney basically im almost at a point where context switing for a day or two would be ok
13:20:57 bauzas I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob
13:21:19 dansmith my patch for the ceph job has gotten us almost all of the way there,
13:21:52 dansmith just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on
13:22:09 dansmith if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases
13:23:27 bauzas dansmith: ok, you're prioritary (because of the blocking gate), so cool with me
13:24:10 dansmith the gate is unblocked, so I think the pressure is off for the moment, but we obviously have to get this resolved ASAP
13:24:35 dansmith bauzas: maybe you could re+W these: https://review.opendev.org/c/openstack/nova/+/880632/5
13:24:45 dansmith the bottom patch already merged, they just need a kick to get into the gate
13:25:26 bauzas dansmith: I did this kickass
13:26:45 dansmith thanks
13:27:27 bauzas fwiw, I'll send a status email saying then that the gate is back
13:28:06 bauzas are we sure all the py38-removal releases were reverted ?
13:29:45 dansmith well, we're passing jobs so, I think that's all we need to know
13:29:56 frickler pretty sure oslo bumps were not reverted yet. but only oslo.db was released
13:30:02 dansmith there have been other releases but they are not yet in requirements, so we're kinda on the edge
13:30:06 dansmith yeah that^
14:44:59 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
15:01:48 dansmith yeah, so the rescue tests work for me locally, which I guess is good
15:01:56 dansmith however, the teardown does fail for me
15:02:29 dansmith which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation
15:03:50 dansmith sorry, s/rescue/rebuild/
15:07:13 bauzas dansmith: bravo
15:09:47 dansmith locally that includes the validations, which is disabled in the job (for some reason) so I pushed up a patch to set those to enabled,
15:10:02 dansmith which will both likely slow it down but also maybe catch some other issues
15:10:53 dansmith that causes us to at least wait for the server to be created and sshable in some of the other cases, before we might try to attach a volume, but like I say, the failures I'm seeing don't seem like they'd be impacted there
15:11:15 dansmith and the regular jobs have those enabled, and we see detach failures there still
15:11:17 dansmith albeit not 100%
15:23:34 dansmith bauzas: this needs re-+W as well: https://review.opendev.org/c/openstack/nova/+/880633
15:23:39 dansmith it's about to pass its recheck
15:24:07 bauzas dansmith: done
15:24:14 dansmith thanks
15:37:44 opendevreview Merged openstack/nova master: Remove silent failure to find a node on rebuild https://review.opendev.org/c/openstack/nova/+/880632
15:38:15 bauzas :)
15:53:56 dansmith well, the distro-based ceph job just passed with validations turned on
15:54:20 dansmith the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that
16:09:36 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:19:19 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:21:17 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:24:28 dansmith okay it timed out, but legitimately, it was making progress just taking much longer because of the validations I think
16:24:48 dansmith so we might need to trim that job down again or we can try with the concurrency upped to 2
16:26:41 dansmith ah, the base job is not concurrency=1, so that's probably the difference there.. might be getting close then!
16:29:18 dansmith gouthamr: ^
16:52:31 opendevreview ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773
17:01:43 opendevreview Merged openstack/nova master: Stop ignoring missing compute nodes in claims https://review.opendev.org/c/openstack/nova/+/880633
17:46:12 dansmith eharney: gouthamr: the cephadm run just passed.. the distro package one failed with errors that look very different than before, but related to ImageBusy type things
17:46:52 dansmith the difference here is enabling validations on both, which means we try to ssh to instances before we do certain things to them, which helps a lot with race conditions. I dunno why those were disabled on these jobs before, but we thought they were enabled
17:47:47 dansmith so my plan is to: (1) switch the cephadm job to be the voting primary job and make the distro one non-voting, (2) set the nova DNM to run against cephadm, and get another run of everything
17:48:31 dansmith is there anything else you want to change before we could/should merge this? the cephfs job is still on focal (because it's defined elsewhere) and thus failing because of packages, but I could hard-code that to jammy here to "fix" that if you want
17:48:44 dansmith it's non-voting so we should be able to switch it externally separately, but let me know what you want
17:50:22 dansmith oh another change on the cephadm job was removing the concurrency limit, which was causing it to run too slow to finish when validations was turned on. We need that to run in parallel anyway, so it's good that it passed in that form
17:53:58 dansmith the cephadm job also finished pretty quickly, which is also a good sign that it was pretty healthy
18:15:11 dansmith ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think
18:16:37 sean-k-mooney that would do it
18:16:55 sean-k-mooney it makes snese why you are siing rbd busy messages
18:17:21 dansmith yeah
18:17:48 sean-k-mooney is this with the mariadb and ceph options to reduce mememory set
18:18:03 sean-k-mooney and idealy the job configred for 8GB of swap?
18:18:31 dansmith it does include the mysql and ceph tweaks, not sure about swap,
18:18:45 sean-k-mooney i think we default to 2GB instead of 8
18:18:49 dansmith but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else
18:19:01 sean-k-mooney ack
18:19:03 dansmith I know
18:19:27 dansmith I'm waiting for the nova one to finish and see if it's the same,

Earlier   Later