| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 18:33:03 | dansmith | so we're not passing when instances are totally dead because their disk didn't come up | |
| 18:33:16 | sean-k-mooney | so i tough tempest had a way to also add validation to test that did not request them | |
| 18:33:22 | dansmith | they're on by default in devstack too, AFAICT | |
| 18:33:36 | dansmith | all my local testing has them on and I realized late these jobs opted out of them | |
| 18:34:13 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| 18:35:34 | dansmith | I'll please ask everyone here to cross all crossable appendages | |
| 18:45:14 | sean-k-mooney | so the nova job with package that had the failing test had 4G of swap and its memory low point was 130MB with all swapp used | |
| 18:45:48 | dansmith | ah I hadn't started looking yet, but yeah, interesting | |
| 18:45:50 | dansmith | did it oom? | |
| 18:45:58 | sean-k-mooney | checkign now | |
| 18:46:10 | sean-k-mooney | yep | |
| 18:46:14 | dansmith | yeah | |
| 18:46:21 | sean-k-mooney | Apr 26 17:18:11 np0033857622 kernel: /usr/bin/python invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0 | |
| 18:46:23 | dansmith | also ceph-osd | |
| 18:46:30 | dansmith | so man, let's hope there's a memory leak that is fixed now | |
| 18:46:51 | dansmith | 1.5GiB resident at OOM time | |
| 18:46:54 | sean-k-mooney | it was right at the end too | |
| 18:47:35 | sean-k-mooney | actuly it OOM'd a python proce the the osd | |
| 18:48:08 | dansmith | ah at basically the same time | |
| 18:48:16 | sean-k-mooney | one thing that si true of both josb is i dont think we collect any ceph logs | |
| 18:48:22 | dansmith | wait, no, | |
| 18:48:35 | dansmith | the python process invoked the killer but it didn't kill the python thing first right? | |
| 18:49:05 | sean-k-mooney | i need to look again but maybe | |
| 18:49:33 | sean-k-mooney | ya so python triggered it | |
| 18:49:53 | sean-k-mooney | and they ya it killed the osd | |
| 18:50:04 | sean-k-mooney | because it presumabel had a higher omm score | |
| 18:50:09 | dansmith | right | |
| 18:50:48 | sean-k-mooney | there are caches that can be tunned in cpeh to reduce the osd memroy usage just an fyi | |
| 18:51:35 | sean-k-mooney | i think bluestore is the (only?) backend format now but it has caches that can be reduced to reduce the osd memory usage if i remmeber coreectly so we can also try that if need | |
| 18:51:45 | sean-k-mooney | i wonder if cephadm uses diffent defualt the ubuntu | |
| 18:52:25 | dansmith | yeah, could be | |
| 18:53:05 | sean-k-mooney | memory_tracker low_point: 341952 so the cephadm job used less memroy over all | |
| 18:53:31 | sean-k-mooney | about 200mb less at its low point | |
| 18:53:42 | dansmith | that could also be affected by performance of the node too.. if one is much faster cpu than the other, it could have meant more activity at a single point or something | |
| 18:53:47 | sean-k-mooney | oh and it had swap free too | |
| 18:53:55 | dansmith | but I agree, that's a strong indicator that it's doing better | |
| 18:54:59 | sean-k-mooney | the ceph.confs are diffent between the too | |
| 18:55:46 | sean-k-mooney | i also found where the ceph logs are we have them for both jobs so no regression there | |
| 18:56:14 | dansmith | they're starting tempest now | |
| 18:56:35 | dansmith | well, one is | |
| 18:58:52 | dansmith | it's also interesting that the nova job has some services disabled that are not disabled in the base jobs AFAIK, which should reduce not only the test load but also the static footprint | |
| 18:59:43 | sean-k-mooney | we turn of swift for one i belive | |
| 18:59:52 | sean-k-mooney | and afew other thigns to manage memory | |
| 18:59:56 | sean-k-mooney | like heat | |
| 19:00:00 | dansmith | and cinder-backup | |
| 19:00:06 | dansmith | which uses a lot of memory for somer eason | |
| 19:03:38 | sean-k-mooney | just looking at the providres while we wait | |
| 19:03:49 | sean-k-mooney | both josb ran on ovh-bhs1 | |
| 19:03:58 | sean-k-mooney | so they hopfully had similar hardware | |
| 19:05:57 | sean-k-mooney | on a side note my laptop refhes has shipped which is proably a good thing since my fans are spinnig up trying to look at thses loogs | |
| 19:07:34 | sean-k-mooney | Elapsed time: 956 sec so just over 15 mins that ok for devstack on a vm | |
| 19:31:28 | dansmith | got one ssh timeout failure, | |
| 19:31:45 | dansmith | but it looks like a normal one, not even specifically volume-related | |
| 19:31:58 | sean-k-mooney | ack so we can proably ignore it | |
| 19:32:15 | dansmith | yeah, hope so | |
| 19:32:18 | opendevreview | Jay Faulkner proposed openstack/nova-specs master: Re-Propose "Ironic Shards" for Bobcat/2023.2 https://review.opendev.org/c/openstack/nova-specs/+/881643 | |
| 19:32:23 | dansmith | oh, but... | |
| 19:32:41 | dansmith | it's been three minutes since the last test finished, which might mean it... | |
| 19:32:47 | dansmith | oh yep, just exploded | |
| 19:32:48 | dansmith | dammit | |
| 19:33:26 | dansmith | looks like everything is failing now, so maybe it just OOMed | |
| 19:34:42 | sean-k-mooney | if so then i would suggestg kicking the swap to 8G for now and we can evaluate other options if that is not enough | |
| 19:35:18 | dansmith | yeah, I can never remember how to do that.. do you have a pointer to a job I can copy? | |
| 19:37:18 | sean-k-mooney | sure ill get it | |
| 19:38:03 | sean-k-mooney | configure_swap_size: 8192 | |
| 19:38:05 | sean-k-mooney | https://github.com/openstack/devstack/blob/master/.zuul.yaml#L569 | |
| 19:38:21 | dansmith | ah, right outside of devstack_vars | |
| 19:38:29 | dansmith | thanks.. we'll see what the logs say | |
| 19:38:33 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/881585/4/.zuul.yaml#603 | |
| 19:38:48 | sean-k-mooney | ya so its set to 4G now jsut bump that or do it in the base job | |
| 19:39:15 | sean-k-mooney | *g* | |
| 19:39:28 | sean-k-mooney | ... 8G you got the point | |
| 19:39:37 | dansmith | yeah it's hard failing now, so something must have gone boom | |
| 19:39:55 | dansmith | I guess that's better than just random fails because it's something we have _some_ control over | |
| 19:39:57 | sean-k-mooney | am im goig to go eat so ill check back later o/ | |
| 19:40:02 | dansmith | o/ | |
| 19:58:05 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:00:58 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:14:10 | dansmith | gouthamr: second successful run in a row on the cephadm job | |
| 20:14:23 | gouthamr | \o/ | |
| 20:14:27 | dansmith | the nova one is more complicated and based on it and it seems to have OOMed or some other major failure | |
| 20:14:39 | dansmith | I'll up the swap to 8g when it finishes and we'll get another data point | |
| 20:17:34 | gouthamr | that's great dansmith; i wanted to check - you're trying to leave "devstack-plugin-ceph-tempest-py3" job alone.. any reason not to switch that to cephadm and delete the special "cephadm" job? | |
| 20:18:39 | dansmith | only just so I could continue to have the comparisons, since at every point we're trying to get a grasp on what helps and hurts | |
| 20:18:55 | dansmith | but yeah, if you want me to just fold them in I guess I can | |
| 20:19:31 | dansmith | I don't have the same feeling that the distro packages are necessarily worse than the upstream ones (although I'm happy if they are and that's a benefit) | |
| 20:19:41 | dansmith | so I'm not in a big hurry to abandon that I guess :) | |
| 20:20:00 | gouthamr | you're being conservative; but there's no bandwidth to maintain both imho :) | |
| 20:20:27 | dansmith | the distro-based job OOMed again, so I guess I want to see if upping the swap makes that work or if it just grows further and OOMs there as well | |
| 20:20:29 | gouthamr | and we're in this situation because we tried to split attention, it was tempting for me at least to not touch what was working | |
| 20:21:03 | dansmith | gouthamr: ack, well, I've already marked it as non-voting which doesn't really hurt anything in the short term, but whatever | |
| 20:22:02 | gouthamr | ack; we can get you unblocked first and make that call, democratically, on the ML? | |
| 20:22:40 | dansmith | gouthamr: I have this all ready to go as soon as the nova job finishes to capture logs, so let me push it up as it is (with 8G) and then we can swap things around after that so I can see what the distro job does with more | |
| 20:22:57 | gouthamr | ack dansmith | |
| 20:23:26 | dansmith | gouthamr: yes of course.. I'm certainly not arguing to keep it in such that we need a vote or anything, so if you're actively hoping to drop that support from devstack or something I certainly won't argue against it | |
| 20:23:35 | dansmith | just want one job that works, is all :) | |
| 20:23:43 | gouthamr | ++ | |
| 20:41:38 | dansmith | wow, no oom on the nova job, but just constant fail until it timed out | |
| 20:43:55 | dansmith | and like no errors in n-cpu log | |
| 20:47:03 | dansmith | seems like all cinder fails | |