Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
18:46:14 dansmith yeah
18:46:21 sean-k-mooney Apr 26 17:18:11 np0033857622 kernel: /usr/bin/python invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
18:46:23 dansmith also ceph-osd
18:46:30 dansmith so man, let's hope there's a memory leak that is fixed now
18:46:51 dansmith 1.5GiB resident at OOM time
18:46:54 sean-k-mooney it was right at the end too
18:47:35 sean-k-mooney actuly it OOM'd a python proce the the osd
18:48:08 dansmith ah at basically the same time
18:48:16 sean-k-mooney one thing that si true of both josb is i dont think we collect any ceph logs
18:48:22 dansmith wait, no,
18:48:35 dansmith the python process invoked the killer but it didn't kill the python thing first right?
18:49:05 sean-k-mooney i need to look again but maybe
18:49:33 sean-k-mooney ya so python triggered it
18:49:53 sean-k-mooney and they ya it killed the osd
18:50:04 sean-k-mooney because it presumabel had a higher omm score
18:50:09 dansmith right
18:50:48 sean-k-mooney there are caches that can be tunned in cpeh to reduce the osd memroy usage just an fyi
18:51:35 sean-k-mooney i think bluestore is the (only?) backend format now but it has caches that can be reduced to reduce the osd memory usage if i remmeber coreectly so we can also try that if need
18:51:45 sean-k-mooney i wonder if cephadm uses diffent defualt the ubuntu
18:52:25 dansmith yeah, could be
18:53:05 sean-k-mooney memory_tracker low_point: 341952 so the cephadm job used less memroy over all
18:53:31 sean-k-mooney about 200mb less at its low point
18:53:42 dansmith that could also be affected by performance of the node too.. if one is much faster cpu than the other, it could have meant more activity at a single point or something
18:53:47 sean-k-mooney oh and it had swap free too
18:53:55 dansmith but I agree, that's a strong indicator that it's doing better
18:54:59 sean-k-mooney the ceph.confs are diffent between the too
18:55:46 sean-k-mooney i also found where the ceph logs are we have them for both jobs so no regression there
18:56:14 dansmith they're starting tempest now
18:56:35 dansmith well, one is
18:58:52 dansmith it's also interesting that the nova job has some services disabled that are not disabled in the base jobs AFAIK, which should reduce not only the test load but also the static footprint
18:59:43 sean-k-mooney we turn of swift for one i belive
18:59:52 sean-k-mooney and afew other thigns to manage memory
18:59:56 sean-k-mooney like heat
19:00:00 dansmith and cinder-backup
19:00:06 dansmith which uses a lot of memory for somer eason
19:03:38 sean-k-mooney just looking at the providres while we wait
19:03:49 sean-k-mooney both josb ran on ovh-bhs1
19:03:58 sean-k-mooney so they hopfully had similar hardware
19:05:57 sean-k-mooney on a side note my laptop refhes has shipped which is proably a good thing since my fans are spinnig up trying to look at thses loogs
19:07:34 sean-k-mooney Elapsed time: 956 sec so just over 15 mins that ok for devstack on a vm
19:31:28 dansmith got one ssh timeout failure,
19:31:45 dansmith but it looks like a normal one, not even specifically volume-related
19:31:58 sean-k-mooney ack so we can proably ignore it
19:32:15 dansmith yeah, hope so
19:32:18 opendevreview Jay Faulkner proposed openstack/nova-specs master: Re-Propose "Ironic Shards" for Bobcat/2023.2 https://review.opendev.org/c/openstack/nova-specs/+/881643
19:32:23 dansmith oh, but...
19:32:41 dansmith it's been three minutes since the last test finished, which might mean it...
19:32:47 dansmith oh yep, just exploded
19:32:48 dansmith dammit
19:33:26 dansmith looks like everything is failing now, so maybe it just OOMed
19:34:42 sean-k-mooney if so then i would suggestg kicking the swap to 8G for now and we can evaluate other options if that is not enough
19:35:18 dansmith yeah, I can never remember how to do that.. do you have a pointer to a job I can copy?
19:37:18 sean-k-mooney sure ill get it
19:38:03 sean-k-mooney configure_swap_size: 8192
19:38:05 sean-k-mooney https://github.com/openstack/devstack/blob/master/.zuul.yaml#L569
19:38:21 dansmith ah, right outside of devstack_vars
19:38:29 dansmith thanks.. we'll see what the logs say
19:38:33 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/881585/4/.zuul.yaml#603
19:38:48 sean-k-mooney ya so its set to 4G now jsut bump that or do it in the base job
19:39:15 sean-k-mooney *g*
19:39:28 sean-k-mooney ... 8G you got the point
19:39:37 dansmith yeah it's hard failing now, so something must have gone boom
19:39:55 dansmith I guess that's better than just random fails because it's something we have _some_ control over
19:39:57 sean-k-mooney am im goig to go eat so ill check back later o/
19:40:02 dansmith o/
19:58:05 opendevreview Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644
20:00:58 opendevreview Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644
20:14:10 dansmith gouthamr: second successful run in a row on the cephadm job
20:14:23 gouthamr \o/
20:14:27 dansmith the nova one is more complicated and based on it and it seems to have OOMed or some other major failure
20:14:39 dansmith I'll up the swap to 8g when it finishes and we'll get another data point
20:17:34 gouthamr that's great dansmith; i wanted to check - you're trying to leave "devstack-plugin-ceph-tempest-py3" job alone.. any reason not to switch that to cephadm and delete the special "cephadm" job?
20:18:39 dansmith only just so I could continue to have the comparisons, since at every point we're trying to get a grasp on what helps and hurts
20:18:55 dansmith but yeah, if you want me to just fold them in I guess I can
20:19:31 dansmith I don't have the same feeling that the distro packages are necessarily worse than the upstream ones (although I'm happy if they are and that's a benefit)
20:19:41 dansmith so I'm not in a big hurry to abandon that I guess :)
20:20:00 gouthamr you're being conservative; but there's no bandwidth to maintain both imho :)
20:20:27 dansmith the distro-based job OOMed again, so I guess I want to see if upping the swap makes that work or if it just grows further and OOMs there as well
20:20:29 gouthamr and we're in this situation because we tried to split attention, it was tempting for me at least to not touch what was working
20:21:03 dansmith gouthamr: ack, well, I've already marked it as non-voting which doesn't really hurt anything in the short term, but whatever
20:22:02 gouthamr ack; we can get you unblocked first and make that call, democratically, on the ML?
20:22:40 dansmith gouthamr: I have this all ready to go as soon as the nova job finishes to capture logs, so let me push it up as it is (with 8G) and then we can swap things around after that so I can see what the distro job does with more
20:22:57 gouthamr ack dansmith
20:23:26 dansmith gouthamr: yes of course.. I'm certainly not arguing to keep it in such that we need a vote or anything, so if you're actively hoping to drop that support from devstack or something I certainly won't argue against it
20:23:35 dansmith just want one job that works, is all :)
20:23:43 gouthamr ++
20:41:38 dansmith wow, no oom on the nova job, but just constant fail until it timed out
20:43:55 dansmith and like no errors in n-cpu log
20:47:03 dansmith seems like all cinder fails
22:01:34 dansmith nova job failed again
22:34:02 dansmith gouthamr: here's the report from the nova job: https://2db12bf686954cafec50-571baf10d8fb9f8a9cb5a5e38315002b.ssl.cf5.rackcdn.com/881585/4/check/nova-ceph-multistore/4694eec/testr_results.html
22:34:07 dansmith not nearly as much fail as before
22:34:15 dansmith the second job is a failure waiting for cinder to detach
22:34:36 dansmith the first one is a nova detach, I need to check that test to see if it's doing an ssh wait, which might help
22:35:11 dansmith oh actually that first one is in the cinder tempest plugin, not one of ours
22:36:56 dansmith looks like it's probably not
22:41:37 dansmith gmann: around?
22:43:08 gmann dansmith: hi
22:43:40 dansmith gmann: looks like the scenario.manager create_server in tempest is not waiting for sshableness
22:43:57 dansmith gmann: and the cinder tempest plugin has a test that uses that and may be part of the volume detach issue(s)

Earlier   Later