Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
18:16:37 sean-k-mooney that would do it
18:16:55 sean-k-mooney it makes snese why you are siing rbd busy messages
18:17:21 dansmith yeah
18:17:48 sean-k-mooney is this with the mariadb and ceph options to reduce mememory set
18:18:03 sean-k-mooney and idealy the job configred for 8GB of swap?
18:18:31 dansmith it does include the mysql and ceph tweaks, not sure about swap,
18:18:45 sean-k-mooney i think we default to 2GB instead of 8
18:18:49 dansmith but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else
18:19:01 sean-k-mooney ack
18:19:03 dansmith I know
18:19:27 dansmith I'm waiting for the nova one to finish and see if it's the same,
18:19:38 dansmith but also to switch it to cephadm and get another set of data
18:19:49 dansmith but this is the first time I've ever seen the cephadm job pass, so I'm encouraged
18:20:46 sean-k-mooney if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm
18:21:29 dansmith we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now
18:21:46 dansmith the cephadm job ran at =4 at top speed with no failures this time
18:22:00 sean-k-mooney ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph
18:22:16 sean-k-mooney cool
18:23:09 sean-k-mooney your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log
18:23:19 dansmith yeah
18:23:30 dansmith I want the logs so I can look while I run it against cephadm as the base
18:23:47 dansmith hoping it's also the same issue as the other one because they look very similar, just watching them
18:23:58 sean-k-mooney an hour and 19 minuts is not bad
18:24:08 sean-k-mooney https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4
18:24:22 dansmith right, it's super fast actually
18:24:33 sean-k-mooney its also singel node which helps
18:24:44 dansmith so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :)
18:24:50 dansmith compared to the distro jobs
18:25:12 sean-k-mooney well going to python 3.10 wil add about 30% to the openstack performace
18:25:20 sean-k-mooney vs 3.8 i think
18:25:29 sean-k-mooney oh
18:25:44 sean-k-mooney you are comparing to disto packages on 22.04 not focal
18:25:51 dansmith yeah, so that might be helping the speed when it doesn't fail
18:25:57 dansmith sean-k-mooney: correct
18:26:15 dansmith the cephadm job takes ceph via podman containers from ceph directly, the nova and -py3 jobs are using jammy's ceph packages
18:26:33 sean-k-mooney cool
18:26:50 sean-k-mooney so wait are you finding containerisation useful for something :P
18:27:08 sean-k-mooney i like cephadm for what its worth
18:27:23 dansmith I have no opposition to containerization in general :)
18:27:28 sean-k-mooney of the 4 ways i have installed cpeh it was the lest janky
18:27:55 sean-k-mooney its also the way i have the least experince with
18:29:10 dansmith okay nova job finished, 11 fails
18:29:16 sean-k-mooney yep
18:29:27 sean-k-mooney but thats not bad it proably detach related
18:29:37 sean-k-mooney im intrest to see the memory stats
18:29:39 dansmith so gouthamr eharney I'm going to push up patches for my above proposal so I can get another run started, just FYI
18:29:53 gouthamr dansmith: o/ yep
18:30:10 sean-k-mooney dansmith: wait a sec
18:30:17 sean-k-mooney the job results are not uploaded yet
18:30:26 dansmith sean-k-mooney: I'm not stupid I know :)
18:30:35 gouthamr dansmith: thanks for the explanation on the tempest validations, there were some more things that fpantano was attempting here: https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/834223
18:30:53 sean-k-mooney :) i just didnt wnat you to push and thing crap i wasted my time
18:30:59 gouthamr dansmith: we can compare notes, and i can comment on the patch you're working on
18:31:03 dansmith gouthamr: ack
18:31:45 sean-k-mooney enabling the validation by defautl will slow thing down but it does indeed work around test tha tshoudl be waiting an arnt
18:32:02 sean-k-mooney dansmith: i assume thats what you did untilll we can fix them
18:32:12 dansmith sean-k-mooney: until we can fix what?
18:32:34 dansmith the sshable checks don't even get run if validations is disabled, that's why I enabled them
18:32:44 sean-k-mooney ah yes
18:32:50 dansmith especially for a job where we're testing a different storage backend for the instances, we definitely should have those on,
18:33:03 dansmith so we're not passing when instances are totally dead because their disk didn't come up
18:33:16 sean-k-mooney so i tough tempest had a way to also add validation to test that did not request them
18:33:22 dansmith they're on by default in devstack too, AFAICT
18:33:36 dansmith all my local testing has them on and I realized late these jobs opted out of them
18:34:13 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
18:35:34 dansmith I'll please ask everyone here to cross all crossable appendages
18:45:14 sean-k-mooney so the nova job with package that had the failing test had 4G of swap and its memory low point was 130MB with all swapp used
18:45:48 dansmith ah I hadn't started looking yet, but yeah, interesting
18:45:50 dansmith did it oom?
18:45:58 sean-k-mooney checkign now
18:46:10 sean-k-mooney yep
18:46:14 dansmith yeah
18:46:21 sean-k-mooney Apr 26 17:18:11 np0033857622 kernel: /usr/bin/python invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
18:46:23 dansmith also ceph-osd
18:46:30 dansmith so man, let's hope there's a memory leak that is fixed now
18:46:51 dansmith 1.5GiB resident at OOM time
18:46:54 sean-k-mooney it was right at the end too
18:47:35 sean-k-mooney actuly it OOM'd a python proce the the osd
18:48:08 dansmith ah at basically the same time
18:48:16 sean-k-mooney one thing that si true of both josb is i dont think we collect any ceph logs
18:48:22 dansmith wait, no,
18:48:35 dansmith the python process invoked the killer but it didn't kill the python thing first right?
18:49:05 sean-k-mooney i need to look again but maybe
18:49:33 sean-k-mooney ya so python triggered it
18:49:53 sean-k-mooney and they ya it killed the osd
18:50:04 sean-k-mooney because it presumabel had a higher omm score
18:50:09 dansmith right
18:50:48 sean-k-mooney there are caches that can be tunned in cpeh to reduce the osd memroy usage just an fyi
18:51:35 sean-k-mooney i think bluestore is the (only?) backend format now but it has caches that can be reduced to reduce the osd memory usage if i remmeber coreectly so we can also try that if need
18:51:45 sean-k-mooney i wonder if cephadm uses diffent defualt the ubuntu
18:52:25 dansmith yeah, could be
18:53:05 sean-k-mooney memory_tracker low_point: 341952 so the cephadm job used less memroy over all
18:53:31 sean-k-mooney about 200mb less at its low point
18:53:42 dansmith that could also be affected by performance of the node too.. if one is much faster cpu than the other, it could have meant more activity at a single point or something
18:53:47 sean-k-mooney oh and it had swap free too
18:53:55 dansmith but I agree, that's a strong indicator that it's doing better
18:54:59 sean-k-mooney the ceph.confs are diffent between the too
18:55:46 sean-k-mooney i also found where the ceph logs are we have them for both jobs so no regression there
18:56:14 dansmith they're starting tempest now
18:56:35 dansmith well, one is

Earlier   Later