| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 18:48:16 | sean-k-mooney | one thing that si true of both josb is i dont think we collect any ceph logs | |
| 18:48:22 | dansmith | wait, no, | |
| 18:48:35 | dansmith | the python process invoked the killer but it didn't kill the python thing first right? | |
| 18:49:05 | sean-k-mooney | i need to look again but maybe | |
| 18:49:33 | sean-k-mooney | ya so python triggered it | |
| 18:49:53 | sean-k-mooney | and they ya it killed the osd | |
| 18:50:04 | sean-k-mooney | because it presumabel had a higher omm score | |
| 18:50:09 | dansmith | right | |
| 18:50:48 | sean-k-mooney | there are caches that can be tunned in cpeh to reduce the osd memroy usage just an fyi | |
| 18:51:35 | sean-k-mooney | i think bluestore is the (only?) backend format now but it has caches that can be reduced to reduce the osd memory usage if i remmeber coreectly so we can also try that if need | |
| 18:51:45 | sean-k-mooney | i wonder if cephadm uses diffent defualt the ubuntu | |
| 18:52:25 | dansmith | yeah, could be | |
| 18:53:05 | sean-k-mooney | memory_tracker low_point: 341952 so the cephadm job used less memroy over all | |
| 18:53:31 | sean-k-mooney | about 200mb less at its low point | |
| 18:53:42 | dansmith | that could also be affected by performance of the node too.. if one is much faster cpu than the other, it could have meant more activity at a single point or something | |
| 18:53:47 | sean-k-mooney | oh and it had swap free too | |
| 18:53:55 | dansmith | but I agree, that's a strong indicator that it's doing better | |
| 18:54:59 | sean-k-mooney | the ceph.confs are diffent between the too | |
| 18:55:46 | sean-k-mooney | i also found where the ceph logs are we have them for both jobs so no regression there | |
| 18:56:14 | dansmith | they're starting tempest now | |
| 18:56:35 | dansmith | well, one is | |
| 18:58:52 | dansmith | it's also interesting that the nova job has some services disabled that are not disabled in the base jobs AFAIK, which should reduce not only the test load but also the static footprint | |
| 18:59:43 | sean-k-mooney | we turn of swift for one i belive | |
| 18:59:52 | sean-k-mooney | and afew other thigns to manage memory | |
| 18:59:56 | sean-k-mooney | like heat | |
| 19:00:00 | dansmith | and cinder-backup | |
| 19:00:06 | dansmith | which uses a lot of memory for somer eason | |
| 19:03:38 | sean-k-mooney | just looking at the providres while we wait | |
| 19:03:49 | sean-k-mooney | both josb ran on ovh-bhs1 | |
| 19:03:58 | sean-k-mooney | so they hopfully had similar hardware | |
| 19:05:57 | sean-k-mooney | on a side note my laptop refhes has shipped which is proably a good thing since my fans are spinnig up trying to look at thses loogs | |
| 19:07:34 | sean-k-mooney | Elapsed time: 956 sec so just over 15 mins that ok for devstack on a vm | |
| 19:31:28 | dansmith | got one ssh timeout failure, | |
| 19:31:45 | dansmith | but it looks like a normal one, not even specifically volume-related | |
| 19:31:58 | sean-k-mooney | ack so we can proably ignore it | |
| 19:32:15 | dansmith | yeah, hope so | |
| 19:32:18 | opendevreview | Jay Faulkner proposed openstack/nova-specs master: Re-Propose "Ironic Shards" for Bobcat/2023.2 https://review.opendev.org/c/openstack/nova-specs/+/881643 | |
| 19:32:23 | dansmith | oh, but... | |
| 19:32:41 | dansmith | it's been three minutes since the last test finished, which might mean it... | |
| 19:32:47 | dansmith | oh yep, just exploded | |
| 19:32:48 | dansmith | dammit | |
| 19:33:26 | dansmith | looks like everything is failing now, so maybe it just OOMed | |
| 19:34:42 | sean-k-mooney | if so then i would suggestg kicking the swap to 8G for now and we can evaluate other options if that is not enough | |
| 19:35:18 | dansmith | yeah, I can never remember how to do that.. do you have a pointer to a job I can copy? | |
| 19:37:18 | sean-k-mooney | sure ill get it | |
| 19:38:03 | sean-k-mooney | configure_swap_size: 8192 | |
| 19:38:05 | sean-k-mooney | https://github.com/openstack/devstack/blob/master/.zuul.yaml#L569 | |
| 19:38:21 | dansmith | ah, right outside of devstack_vars | |
| 19:38:29 | dansmith | thanks.. we'll see what the logs say | |
| 19:38:33 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/881585/4/.zuul.yaml#603 | |
| 19:38:48 | sean-k-mooney | ya so its set to 4G now jsut bump that or do it in the base job | |
| 19:39:15 | sean-k-mooney | *g* | |
| 19:39:28 | sean-k-mooney | ... 8G you got the point | |
| 19:39:37 | dansmith | yeah it's hard failing now, so something must have gone boom | |
| 19:39:55 | dansmith | I guess that's better than just random fails because it's something we have _some_ control over | |
| 19:39:57 | sean-k-mooney | am im goig to go eat so ill check back later o/ | |
| 19:40:02 | dansmith | o/ | |
| 19:58:05 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:00:58 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:14:10 | dansmith | gouthamr: second successful run in a row on the cephadm job | |
| 20:14:23 | gouthamr | \o/ | |
| 20:14:27 | dansmith | the nova one is more complicated and based on it and it seems to have OOMed or some other major failure | |
| 20:14:39 | dansmith | I'll up the swap to 8g when it finishes and we'll get another data point | |
| 20:17:34 | gouthamr | that's great dansmith; i wanted to check - you're trying to leave "devstack-plugin-ceph-tempest-py3" job alone.. any reason not to switch that to cephadm and delete the special "cephadm" job? | |
| 20:18:39 | dansmith | only just so I could continue to have the comparisons, since at every point we're trying to get a grasp on what helps and hurts | |
| 20:18:55 | dansmith | but yeah, if you want me to just fold them in I guess I can | |
| 20:19:31 | dansmith | I don't have the same feeling that the distro packages are necessarily worse than the upstream ones (although I'm happy if they are and that's a benefit) | |
| 20:19:41 | dansmith | so I'm not in a big hurry to abandon that I guess :) | |
| 20:20:00 | gouthamr | you're being conservative; but there's no bandwidth to maintain both imho :) | |
| 20:20:27 | dansmith | the distro-based job OOMed again, so I guess I want to see if upping the swap makes that work or if it just grows further and OOMs there as well | |
| 20:20:29 | gouthamr | and we're in this situation because we tried to split attention, it was tempting for me at least to not touch what was working | |
| 20:21:03 | dansmith | gouthamr: ack, well, I've already marked it as non-voting which doesn't really hurt anything in the short term, but whatever | |
| 20:22:02 | gouthamr | ack; we can get you unblocked first and make that call, democratically, on the ML? | |
| 20:22:40 | dansmith | gouthamr: I have this all ready to go as soon as the nova job finishes to capture logs, so let me push it up as it is (with 8G) and then we can swap things around after that so I can see what the distro job does with more | |
| 20:22:57 | gouthamr | ack dansmith | |
| 20:23:26 | dansmith | gouthamr: yes of course.. I'm certainly not arguing to keep it in such that we need a vote or anything, so if you're actively hoping to drop that support from devstack or something I certainly won't argue against it | |
| 20:23:35 | dansmith | just want one job that works, is all :) | |
| 20:23:43 | gouthamr | ++ | |
| 20:41:38 | dansmith | wow, no oom on the nova job, but just constant fail until it timed out | |
| 20:43:55 | dansmith | and like no errors in n-cpu log | |
| 20:47:03 | dansmith | seems like all cinder fails | |
| 22:01:34 | dansmith | nova job failed again | |
| 22:34:02 | dansmith | gouthamr: here's the report from the nova job: https://2db12bf686954cafec50-571baf10d8fb9f8a9cb5a5e38315002b.ssl.cf5.rackcdn.com/881585/4/check/nova-ceph-multistore/4694eec/testr_results.html | |
| 22:34:07 | dansmith | not nearly as much fail as before | |
| 22:34:15 | dansmith | the second job is a failure waiting for cinder to detach | |
| 22:34:36 | dansmith | the first one is a nova detach, I need to check that test to see if it's doing an ssh wait, which might help | |
| 22:35:11 | dansmith | oh actually that first one is in the cinder tempest plugin, not one of ours | |
| 22:36:56 | dansmith | looks like it's probably not | |
| 22:41:37 | dansmith | gmann: around? | |
| 22:43:08 | gmann | dansmith: hi | |
| 22:43:40 | dansmith | gmann: looks like the scenario.manager create_server in tempest is not waiting for sshableness | |
| 22:43:57 | dansmith | gmann: and the cinder tempest plugin has a test that uses that and may be part of the volume detach issue(s) | |
| 22:44:11 | dansmith | gmann: is it possible to make that create_server() always wait? | |
| 22:44:59 | gmann | this one right? https://github.com/openstack/cinder-tempest-plugin/blob/d6989d3c1a31f2bebf94f2a7a8dac1d9eb788b1b/cinder_tempest_plugin/scenario/test_volume_encrypted.py#L86 | |
| 22:45:24 | dansmith | gmann: yeah, but I'm wondering if create_server() itself can wait so anything that uses it will wait | |
| 22:45:45 | gmann | dansmith: I think it make sense to wait for SSH by default in scenario tests | |
| 22:45:48 | dansmith | gmann: the current thinking is that a lot of our volume detach problems come from trying to attach a volume before an instance is booted | |
| 22:45:58 | dansmith | gmann: ack, okay | |
| 22:46:33 | gmann | yeah, and with nature of most of scenario tests, booting server fully is always good before trying things on that | |
| 22:47:20 | dansmith | yeah, okay lemme try to do that | |