Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
13:29:56 frickler pretty sure oslo bumps were not reverted yet. but only oslo.db was released
13:30:02 dansmith there have been other releases but they are not yet in requirements, so we're kinda on the edge
13:30:06 dansmith yeah that^
14:44:59 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
15:01:48 dansmith yeah, so the rescue tests work for me locally, which I guess is good
15:01:56 dansmith however, the teardown does fail for me
15:02:29 dansmith which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation
15:03:50 dansmith sorry, s/rescue/rebuild/
15:07:13 bauzas dansmith: bravo
15:09:47 dansmith locally that includes the validations, which is disabled in the job (for some reason) so I pushed up a patch to set those to enabled,
15:10:02 dansmith which will both likely slow it down but also maybe catch some other issues
15:10:53 dansmith that causes us to at least wait for the server to be created and sshable in some of the other cases, before we might try to attach a volume, but like I say, the failures I'm seeing don't seem like they'd be impacted there
15:11:15 dansmith and the regular jobs have those enabled, and we see detach failures there still
15:11:17 dansmith albeit not 100%
15:23:34 dansmith bauzas: this needs re-+W as well: https://review.opendev.org/c/openstack/nova/+/880633
15:23:39 dansmith it's about to pass its recheck
15:24:07 bauzas dansmith: done
15:24:14 dansmith thanks
15:37:44 opendevreview Merged openstack/nova master: Remove silent failure to find a node on rebuild https://review.opendev.org/c/openstack/nova/+/880632
15:38:15 bauzas :)
15:53:56 dansmith well, the distro-based ceph job just passed with validations turned on
15:54:20 dansmith the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that
16:09:36 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:19:19 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:21:17 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:24:28 dansmith okay it timed out, but legitimately, it was making progress just taking much longer because of the validations I think
16:24:48 dansmith so we might need to trim that job down again or we can try with the concurrency upped to 2
16:26:41 dansmith ah, the base job is not concurrency=1, so that's probably the difference there.. might be getting close then!
16:29:18 dansmith gouthamr: ^
16:52:31 opendevreview ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773
17:01:43 opendevreview Merged openstack/nova master: Stop ignoring missing compute nodes in claims https://review.opendev.org/c/openstack/nova/+/880633
17:46:12 dansmith eharney: gouthamr: the cephadm run just passed.. the distro package one failed with errors that look very different than before, but related to ImageBusy type things
17:46:52 dansmith the difference here is enabling validations on both, which means we try to ssh to instances before we do certain things to them, which helps a lot with race conditions. I dunno why those were disabled on these jobs before, but we thought they were enabled
17:47:47 dansmith so my plan is to: (1) switch the cephadm job to be the voting primary job and make the distro one non-voting, (2) set the nova DNM to run against cephadm, and get another run of everything
17:48:31 dansmith is there anything else you want to change before we could/should merge this? the cephfs job is still on focal (because it's defined elsewhere) and thus failing because of packages, but I could hard-code that to jammy here to "fix" that if you want
17:48:44 dansmith it's non-voting so we should be able to switch it externally separately, but let me know what you want
17:50:22 dansmith oh another change on the cephadm job was removing the concurrency limit, which was causing it to run too slow to finish when validations was turned on. We need that to run in parallel anyway, so it's good that it passed in that form
17:53:58 dansmith the cephadm job also finished pretty quickly, which is also a good sign that it was pretty healthy
18:15:11 dansmith ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think
18:16:37 sean-k-mooney that would do it
18:16:55 sean-k-mooney it makes snese why you are siing rbd busy messages
18:17:21 dansmith yeah
18:17:48 sean-k-mooney is this with the mariadb and ceph options to reduce mememory set
18:18:03 sean-k-mooney and idealy the job configred for 8GB of swap?
18:18:31 dansmith it does include the mysql and ceph tweaks, not sure about swap,
18:18:45 sean-k-mooney i think we default to 2GB instead of 8
18:18:49 dansmith but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else
18:19:01 sean-k-mooney ack
18:19:03 dansmith I know
18:19:27 dansmith I'm waiting for the nova one to finish and see if it's the same,
18:19:38 dansmith but also to switch it to cephadm and get another set of data
18:19:49 dansmith but this is the first time I've ever seen the cephadm job pass, so I'm encouraged
18:20:46 sean-k-mooney if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm
18:21:29 dansmith we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now
18:21:46 dansmith the cephadm job ran at =4 at top speed with no failures this time
18:22:00 sean-k-mooney ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph
18:22:16 sean-k-mooney cool
18:23:09 sean-k-mooney your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log
18:23:19 dansmith yeah
18:23:30 dansmith I want the logs so I can look while I run it against cephadm as the base
18:23:47 dansmith hoping it's also the same issue as the other one because they look very similar, just watching them
18:23:58 sean-k-mooney an hour and 19 minuts is not bad
18:24:08 sean-k-mooney https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4
18:24:22 dansmith right, it's super fast actually
18:24:33 sean-k-mooney its also singel node which helps
18:24:44 dansmith so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :)
18:24:50 dansmith compared to the distro jobs
18:25:12 sean-k-mooney well going to python 3.10 wil add about 30% to the openstack performace
18:25:20 sean-k-mooney vs 3.8 i think
18:25:29 sean-k-mooney oh
18:25:44 sean-k-mooney you are comparing to disto packages on 22.04 not focal
18:25:51 dansmith yeah, so that might be helping the speed when it doesn't fail
18:25:57 dansmith sean-k-mooney: correct
18:26:15 dansmith the cephadm job takes ceph via podman containers from ceph directly, the nova and -py3 jobs are using jammy's ceph packages
18:26:33 sean-k-mooney cool
18:26:50 sean-k-mooney so wait are you finding containerisation useful for something :P
18:27:08 sean-k-mooney i like cephadm for what its worth
18:27:23 dansmith I have no opposition to containerization in general :)
18:27:28 sean-k-mooney of the 4 ways i have installed cpeh it was the lest janky
18:27:55 sean-k-mooney its also the way i have the least experince with
18:29:10 dansmith okay nova job finished, 11 fails
18:29:16 sean-k-mooney yep
18:29:27 sean-k-mooney but thats not bad it proably detach related
18:29:37 sean-k-mooney im intrest to see the memory stats
18:29:39 dansmith so gouthamr eharney I'm going to push up patches for my above proposal so I can get another run started, just FYI
18:29:53 gouthamr dansmith: o/ yep
18:30:10 sean-k-mooney dansmith: wait a sec
18:30:17 sean-k-mooney the job results are not uploaded yet
18:30:26 dansmith sean-k-mooney: I'm not stupid I know :)
18:30:35 gouthamr dansmith: thanks for the explanation on the tempest validations, there were some more things that fpantano was attempting here: https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/834223
18:30:53 sean-k-mooney :) i just didnt wnat you to push and thing crap i wasted my time
18:30:59 gouthamr dansmith: we can compare notes, and i can comment on the patch you're working on
18:31:03 dansmith gouthamr: ack
18:31:45 sean-k-mooney enabling the validation by defautl will slow thing down but it does indeed work around test tha tshoudl be waiting an arnt
18:32:02 sean-k-mooney dansmith: i assume thats what you did untilll we can fix them
18:32:12 dansmith sean-k-mooney: until we can fix what?
18:32:34 dansmith the sshable checks don't even get run if validations is disabled, that's why I enabled them
18:32:44 sean-k-mooney ah yes
18:32:50 dansmith especially for a job where we're testing a different storage backend for the instances, we definitely should have those on,
18:33:03 dansmith so we're not passing when instances are totally dead because their disk didn't come up

Earlier   Later