| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 18:15:11 | dansmith | ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think | |
| 18:16:37 | sean-k-mooney | that would do it | |
| 18:16:55 | sean-k-mooney | it makes snese why you are siing rbd busy messages | |
| 18:17:21 | dansmith | yeah | |
| 18:17:48 | sean-k-mooney | is this with the mariadb and ceph options to reduce mememory set | |
| 18:18:03 | sean-k-mooney | and idealy the job configred for 8GB of swap? | |
| 18:18:31 | dansmith | it does include the mysql and ceph tweaks, not sure about swap, | |
| 18:18:45 | sean-k-mooney | i think we default to 2GB instead of 8 | |
| 18:18:49 | dansmith | but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else | |
| 18:19:01 | sean-k-mooney | ack | |
| 18:19:03 | dansmith | I know | |
| 18:19:27 | dansmith | I'm waiting for the nova one to finish and see if it's the same, | |
| 18:19:38 | dansmith | but also to switch it to cephadm and get another set of data | |
| 18:19:49 | dansmith | but this is the first time I've ever seen the cephadm job pass, so I'm encouraged | |
| 18:20:46 | sean-k-mooney | if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm | |
| 18:21:29 | dansmith | we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now | |
| 18:21:46 | dansmith | the cephadm job ran at =4 at top speed with no failures this time | |
| 18:22:00 | sean-k-mooney | ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph | |
| 18:22:16 | sean-k-mooney | cool | |
| 18:23:09 | sean-k-mooney | your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log | |
| 18:23:19 | dansmith | yeah | |
| 18:23:30 | dansmith | I want the logs so I can look while I run it against cephadm as the base | |
| 18:23:47 | dansmith | hoping it's also the same issue as the other one because they look very similar, just watching them | |
| 18:23:58 | sean-k-mooney | an hour and 19 minuts is not bad | |
| 18:24:08 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4 | |
| 18:24:22 | dansmith | right, it's super fast actually | |
| 18:24:33 | sean-k-mooney | its also singel node which helps | |
| 18:24:44 | dansmith | so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :) | |
| 18:24:50 | dansmith | compared to the distro jobs | |
| 18:25:12 | sean-k-mooney | well going to python 3.10 wil add about 30% to the openstack performace | |
| 18:25:20 | sean-k-mooney | vs 3.8 i think | |
| 18:25:29 | sean-k-mooney | oh | |
| 18:25:44 | sean-k-mooney | you are comparing to disto packages on 22.04 not focal | |
| 18:25:51 | dansmith | yeah, so that might be helping the speed when it doesn't fail | |
| 18:25:57 | dansmith | sean-k-mooney: correct | |
| 18:26:15 | dansmith | the cephadm job takes ceph via podman containers from ceph directly, the nova and -py3 jobs are using jammy's ceph packages | |
| 18:26:33 | sean-k-mooney | cool | |
| 18:26:50 | sean-k-mooney | so wait are you finding containerisation useful for something :P | |
| 18:27:08 | sean-k-mooney | i like cephadm for what its worth | |
| 18:27:23 | dansmith | I have no opposition to containerization in general :) | |
| 18:27:28 | sean-k-mooney | of the 4 ways i have installed cpeh it was the lest janky | |
| 18:27:55 | sean-k-mooney | its also the way i have the least experince with | |
| 18:29:10 | dansmith | okay nova job finished, 11 fails | |
| 18:29:16 | sean-k-mooney | yep | |
| 18:29:27 | sean-k-mooney | but thats not bad it proably detach related | |
| 18:29:37 | sean-k-mooney | im intrest to see the memory stats | |
| 18:29:39 | dansmith | so gouthamr eharney I'm going to push up patches for my above proposal so I can get another run started, just FYI | |
| 18:29:53 | gouthamr | dansmith: o/ yep | |
| 18:30:10 | sean-k-mooney | dansmith: wait a sec | |
| 18:30:17 | sean-k-mooney | the job results are not uploaded yet | |
| 18:30:26 | dansmith | sean-k-mooney: I'm not stupid I know :) | |
| 18:30:35 | gouthamr | dansmith: thanks for the explanation on the tempest validations, there were some more things that fpantano was attempting here: https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/834223 | |
| 18:30:53 | sean-k-mooney | :) i just didnt wnat you to push and thing crap i wasted my time | |
| 18:30:59 | gouthamr | dansmith: we can compare notes, and i can comment on the patch you're working on | |
| 18:31:03 | dansmith | gouthamr: ack | |
| 18:31:45 | sean-k-mooney | enabling the validation by defautl will slow thing down but it does indeed work around test tha tshoudl be waiting an arnt | |
| 18:32:02 | sean-k-mooney | dansmith: i assume thats what you did untilll we can fix them | |
| 18:32:12 | dansmith | sean-k-mooney: until we can fix what? | |
| 18:32:34 | dansmith | the sshable checks don't even get run if validations is disabled, that's why I enabled them | |
| 18:32:44 | sean-k-mooney | ah yes | |
| 18:32:50 | dansmith | especially for a job where we're testing a different storage backend for the instances, we definitely should have those on, | |
| 18:33:03 | dansmith | so we're not passing when instances are totally dead because their disk didn't come up | |
| 18:33:16 | sean-k-mooney | so i tough tempest had a way to also add validation to test that did not request them | |
| 18:33:22 | dansmith | they're on by default in devstack too, AFAICT | |
| 18:33:36 | dansmith | all my local testing has them on and I realized late these jobs opted out of them | |
| 18:34:13 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| 18:35:34 | dansmith | I'll please ask everyone here to cross all crossable appendages | |
| 18:45:14 | sean-k-mooney | so the nova job with package that had the failing test had 4G of swap and its memory low point was 130MB with all swapp used | |
| 18:45:48 | dansmith | ah I hadn't started looking yet, but yeah, interesting | |
| 18:45:50 | dansmith | did it oom? | |
| 18:45:58 | sean-k-mooney | checkign now | |
| 18:46:10 | sean-k-mooney | yep | |
| 18:46:14 | dansmith | yeah | |
| 18:46:21 | sean-k-mooney | Apr 26 17:18:11 np0033857622 kernel: /usr/bin/python invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0 | |
| 18:46:23 | dansmith | also ceph-osd | |
| 18:46:30 | dansmith | so man, let's hope there's a memory leak that is fixed now | |
| 18:46:51 | dansmith | 1.5GiB resident at OOM time | |
| 18:46:54 | sean-k-mooney | it was right at the end too | |
| 18:47:35 | sean-k-mooney | actuly it OOM'd a python proce the the osd | |
| 18:48:08 | dansmith | ah at basically the same time | |
| 18:48:16 | sean-k-mooney | one thing that si true of both josb is i dont think we collect any ceph logs | |
| 18:48:22 | dansmith | wait, no, | |
| 18:48:35 | dansmith | the python process invoked the killer but it didn't kill the python thing first right? | |
| 18:49:05 | sean-k-mooney | i need to look again but maybe | |
| 18:49:33 | sean-k-mooney | ya so python triggered it | |
| 18:49:53 | sean-k-mooney | and they ya it killed the osd | |
| 18:50:04 | sean-k-mooney | because it presumabel had a higher omm score | |
| 18:50:09 | dansmith | right | |
| 18:50:48 | sean-k-mooney | there are caches that can be tunned in cpeh to reduce the osd memroy usage just an fyi | |
| 18:51:35 | sean-k-mooney | i think bluestore is the (only?) backend format now but it has caches that can be reduced to reduce the osd memory usage if i remmeber coreectly so we can also try that if need | |
| 18:51:45 | sean-k-mooney | i wonder if cephadm uses diffent defualt the ubuntu | |
| 18:52:25 | dansmith | yeah, could be | |
| 18:53:05 | sean-k-mooney | memory_tracker low_point: 341952 so the cephadm job used less memroy over all | |
| 18:53:31 | sean-k-mooney | about 200mb less at its low point | |
| 18:53:42 | dansmith | that could also be affected by performance of the node too.. if one is much faster cpu than the other, it could have meant more activity at a single point or something | |
| 18:53:47 | sean-k-mooney | oh and it had swap free too | |
| 18:53:55 | dansmith | but I agree, that's a strong indicator that it's doing better | |
| 18:54:59 | sean-k-mooney | the ceph.confs are diffent between the too | |
| 18:55:46 | sean-k-mooney | i also found where the ceph logs are we have them for both jobs so no regression there | |
| 18:56:14 | dansmith | they're starting tempest now | |