Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
15:38:15 bauzas :)
15:53:56 dansmith well, the distro-based ceph job just passed with validations turned on
15:54:20 dansmith the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that
16:09:36 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:19:19 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:21:17 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
16:24:28 dansmith okay it timed out, but legitimately, it was making progress just taking much longer because of the validations I think
16:24:48 dansmith so we might need to trim that job down again or we can try with the concurrency upped to 2
16:26:41 dansmith ah, the base job is not concurrency=1, so that's probably the difference there.. might be getting close then!
16:29:18 dansmith gouthamr: ^
16:52:31 opendevreview ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773
17:01:43 opendevreview Merged openstack/nova master: Stop ignoring missing compute nodes in claims https://review.opendev.org/c/openstack/nova/+/880633
17:46:12 dansmith eharney: gouthamr: the cephadm run just passed.. the distro package one failed with errors that look very different than before, but related to ImageBusy type things
17:46:52 dansmith the difference here is enabling validations on both, which means we try to ssh to instances before we do certain things to them, which helps a lot with race conditions. I dunno why those were disabled on these jobs before, but we thought they were enabled
17:47:47 dansmith so my plan is to: (1) switch the cephadm job to be the voting primary job and make the distro one non-voting, (2) set the nova DNM to run against cephadm, and get another run of everything
17:48:31 dansmith is there anything else you want to change before we could/should merge this? the cephfs job is still on focal (because it's defined elsewhere) and thus failing because of packages, but I could hard-code that to jammy here to "fix" that if you want
17:48:44 dansmith it's non-voting so we should be able to switch it externally separately, but let me know what you want
17:50:22 dansmith oh another change on the cephadm job was removing the concurrency limit, which was causing it to run too slow to finish when validations was turned on. We need that to run in parallel anyway, so it's good that it passed in that form
17:53:58 dansmith the cephadm job also finished pretty quickly, which is also a good sign that it was pretty healthy
18:15:11 dansmith ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think
18:16:37 sean-k-mooney that would do it
18:16:55 sean-k-mooney it makes snese why you are siing rbd busy messages
18:17:21 dansmith yeah
18:17:48 sean-k-mooney is this with the mariadb and ceph options to reduce mememory set
18:18:03 sean-k-mooney and idealy the job configred for 8GB of swap?
18:18:31 dansmith it does include the mysql and ceph tweaks, not sure about swap,
18:18:45 sean-k-mooney i think we default to 2GB instead of 8
18:18:49 dansmith but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else
18:19:01 sean-k-mooney ack
18:19:03 dansmith I know
18:19:27 dansmith I'm waiting for the nova one to finish and see if it's the same,
18:19:38 dansmith but also to switch it to cephadm and get another set of data
18:19:49 dansmith but this is the first time I've ever seen the cephadm job pass, so I'm encouraged
18:20:46 sean-k-mooney if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm
18:21:29 dansmith we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now
18:21:46 dansmith the cephadm job ran at =4 at top speed with no failures this time
18:22:00 sean-k-mooney ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph
18:22:16 sean-k-mooney cool
18:23:09 sean-k-mooney your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log
18:23:19 dansmith yeah
18:23:30 dansmith I want the logs so I can look while I run it against cephadm as the base
18:23:47 dansmith hoping it's also the same issue as the other one because they look very similar, just watching them
18:23:58 sean-k-mooney an hour and 19 minuts is not bad
18:24:08 sean-k-mooney https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4
18:24:22 dansmith right, it's super fast actually
18:24:33 sean-k-mooney its also singel node which helps
18:24:44 dansmith so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :)
18:24:50 dansmith compared to the distro jobs
18:25:12 sean-k-mooney well going to python 3.10 wil add about 30% to the openstack performace
18:25:20 sean-k-mooney vs 3.8 i think
18:25:29 sean-k-mooney oh
18:25:44 sean-k-mooney you are comparing to disto packages on 22.04 not focal
18:25:51 dansmith yeah, so that might be helping the speed when it doesn't fail
18:25:57 dansmith sean-k-mooney: correct
18:26:15 dansmith the cephadm job takes ceph via podman containers from ceph directly, the nova and -py3 jobs are using jammy's ceph packages
18:26:33 sean-k-mooney cool
18:26:50 sean-k-mooney so wait are you finding containerisation useful for something :P
18:27:08 sean-k-mooney i like cephadm for what its worth
18:27:23 dansmith I have no opposition to containerization in general :)
18:27:28 sean-k-mooney of the 4 ways i have installed cpeh it was the lest janky
18:27:55 sean-k-mooney its also the way i have the least experince with
18:29:10 dansmith okay nova job finished, 11 fails
18:29:16 sean-k-mooney yep
18:29:27 sean-k-mooney but thats not bad it proably detach related
18:29:37 sean-k-mooney im intrest to see the memory stats
18:29:39 dansmith so gouthamr eharney I'm going to push up patches for my above proposal so I can get another run started, just FYI
18:29:53 gouthamr dansmith: o/ yep
18:30:10 sean-k-mooney dansmith: wait a sec
18:30:17 sean-k-mooney the job results are not uploaded yet
18:30:26 dansmith sean-k-mooney: I'm not stupid I know :)
18:30:35 gouthamr dansmith: thanks for the explanation on the tempest validations, there were some more things that fpantano was attempting here: https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/834223
18:30:53 sean-k-mooney :) i just didnt wnat you to push and thing crap i wasted my time
18:30:59 gouthamr dansmith: we can compare notes, and i can comment on the patch you're working on
18:31:03 dansmith gouthamr: ack
18:31:45 sean-k-mooney enabling the validation by defautl will slow thing down but it does indeed work around test tha tshoudl be waiting an arnt
18:32:02 sean-k-mooney dansmith: i assume thats what you did untilll we can fix them
18:32:12 dansmith sean-k-mooney: until we can fix what?
18:32:34 dansmith the sshable checks don't even get run if validations is disabled, that's why I enabled them
18:32:44 sean-k-mooney ah yes
18:32:50 dansmith especially for a job where we're testing a different storage backend for the instances, we definitely should have those on,
18:33:03 dansmith so we're not passing when instances are totally dead because their disk didn't come up
18:33:16 sean-k-mooney so i tough tempest had a way to also add validation to test that did not request them
18:33:22 dansmith they're on by default in devstack too, AFAICT
18:33:36 dansmith all my local testing has them on and I realized late these jobs opted out of them
18:34:13 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
18:35:34 dansmith I'll please ask everyone here to cross all crossable appendages
18:45:14 sean-k-mooney so the nova job with package that had the failing test had 4G of swap and its memory low point was 130MB with all swapp used
18:45:48 dansmith ah I hadn't started looking yet, but yeah, interesting
18:45:50 dansmith did it oom?
18:45:58 sean-k-mooney checkign now
18:46:10 sean-k-mooney yep
18:46:14 dansmith yeah
18:46:21 sean-k-mooney Apr 26 17:18:11 np0033857622 kernel: /usr/bin/python invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
18:46:23 dansmith also ceph-osd
18:46:30 dansmith so man, let's hope there's a memory leak that is fixed now
18:46:51 dansmith 1.5GiB resident at OOM time
18:46:54 sean-k-mooney it was right at the end too
18:47:35 sean-k-mooney actuly it OOM'd a python proce the the osd
18:48:08 dansmith ah at basically the same time
18:48:16 sean-k-mooney one thing that si true of both josb is i dont think we collect any ceph logs

Earlier   Later