| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 19:03:38 | sean-k-mooney | just looking at the providres while we wait | |
| 19:03:49 | sean-k-mooney | both josb ran on ovh-bhs1 | |
| 19:03:58 | sean-k-mooney | so they hopfully had similar hardware | |
| 19:05:57 | sean-k-mooney | on a side note my laptop refhes has shipped which is proably a good thing since my fans are spinnig up trying to look at thses loogs | |
| 19:07:34 | sean-k-mooney | Elapsed time: 956 sec so just over 15 mins that ok for devstack on a vm | |
| 19:31:28 | dansmith | got one ssh timeout failure, | |
| 19:31:45 | dansmith | but it looks like a normal one, not even specifically volume-related | |
| 19:31:58 | sean-k-mooney | ack so we can proably ignore it | |
| 19:32:15 | dansmith | yeah, hope so | |
| 19:32:18 | opendevreview | Jay Faulkner proposed openstack/nova-specs master: Re-Propose "Ironic Shards" for Bobcat/2023.2 https://review.opendev.org/c/openstack/nova-specs/+/881643 | |
| 19:32:23 | dansmith | oh, but... | |
| 19:32:41 | dansmith | it's been three minutes since the last test finished, which might mean it... | |
| 19:32:47 | dansmith | oh yep, just exploded | |
| 19:32:48 | dansmith | dammit | |
| 19:33:26 | dansmith | looks like everything is failing now, so maybe it just OOMed | |
| 19:34:42 | sean-k-mooney | if so then i would suggestg kicking the swap to 8G for now and we can evaluate other options if that is not enough | |
| 19:35:18 | dansmith | yeah, I can never remember how to do that.. do you have a pointer to a job I can copy? | |
| 19:37:18 | sean-k-mooney | sure ill get it | |
| 19:38:03 | sean-k-mooney | configure_swap_size: 8192 | |
| 19:38:05 | sean-k-mooney | https://github.com/openstack/devstack/blob/master/.zuul.yaml#L569 | |
| 19:38:21 | dansmith | ah, right outside of devstack_vars | |
| 19:38:29 | dansmith | thanks.. we'll see what the logs say | |
| 19:38:33 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/881585/4/.zuul.yaml#603 | |
| 19:38:48 | sean-k-mooney | ya so its set to 4G now jsut bump that or do it in the base job | |
| 19:39:15 | sean-k-mooney | *g* | |
| 19:39:28 | sean-k-mooney | ... 8G you got the point | |
| 19:39:37 | dansmith | yeah it's hard failing now, so something must have gone boom | |
| 19:39:55 | dansmith | I guess that's better than just random fails because it's something we have _some_ control over | |
| 19:39:57 | sean-k-mooney | am im goig to go eat so ill check back later o/ | |
| 19:40:02 | dansmith | o/ | |
| 19:58:05 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:00:58 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 20:14:10 | dansmith | gouthamr: second successful run in a row on the cephadm job | |
| 20:14:23 | gouthamr | \o/ | |
| 20:14:27 | dansmith | the nova one is more complicated and based on it and it seems to have OOMed or some other major failure | |
| 20:14:39 | dansmith | I'll up the swap to 8g when it finishes and we'll get another data point | |
| 20:17:34 | gouthamr | that's great dansmith; i wanted to check - you're trying to leave "devstack-plugin-ceph-tempest-py3" job alone.. any reason not to switch that to cephadm and delete the special "cephadm" job? | |
| 20:18:39 | dansmith | only just so I could continue to have the comparisons, since at every point we're trying to get a grasp on what helps and hurts | |
| 20:18:55 | dansmith | but yeah, if you want me to just fold them in I guess I can | |
| 20:19:31 | dansmith | I don't have the same feeling that the distro packages are necessarily worse than the upstream ones (although I'm happy if they are and that's a benefit) | |
| 20:19:41 | dansmith | so I'm not in a big hurry to abandon that I guess :) | |
| 20:20:00 | gouthamr | you're being conservative; but there's no bandwidth to maintain both imho :) | |
| 20:20:27 | dansmith | the distro-based job OOMed again, so I guess I want to see if upping the swap makes that work or if it just grows further and OOMs there as well | |
| 20:20:29 | gouthamr | and we're in this situation because we tried to split attention, it was tempting for me at least to not touch what was working | |
| 20:21:03 | dansmith | gouthamr: ack, well, I've already marked it as non-voting which doesn't really hurt anything in the short term, but whatever | |
| 20:22:02 | gouthamr | ack; we can get you unblocked first and make that call, democratically, on the ML? | |
| 20:22:40 | dansmith | gouthamr: I have this all ready to go as soon as the nova job finishes to capture logs, so let me push it up as it is (with 8G) and then we can swap things around after that so I can see what the distro job does with more | |
| 20:22:57 | gouthamr | ack dansmith | |
| 20:23:26 | dansmith | gouthamr: yes of course.. I'm certainly not arguing to keep it in such that we need a vote or anything, so if you're actively hoping to drop that support from devstack or something I certainly won't argue against it | |
| 20:23:35 | dansmith | just want one job that works, is all :) | |
| 20:23:43 | gouthamr | ++ | |
| 20:41:38 | dansmith | wow, no oom on the nova job, but just constant fail until it timed out | |
| 20:43:55 | dansmith | and like no errors in n-cpu log | |
| 20:47:03 | dansmith | seems like all cinder fails | |
| 22:01:34 | dansmith | nova job failed again | |
| 22:34:02 | dansmith | gouthamr: here's the report from the nova job: https://2db12bf686954cafec50-571baf10d8fb9f8a9cb5a5e38315002b.ssl.cf5.rackcdn.com/881585/4/check/nova-ceph-multistore/4694eec/testr_results.html | |
| 22:34:07 | dansmith | not nearly as much fail as before | |
| 22:34:15 | dansmith | the second job is a failure waiting for cinder to detach | |
| 22:34:36 | dansmith | the first one is a nova detach, I need to check that test to see if it's doing an ssh wait, which might help | |
| 22:35:11 | dansmith | oh actually that first one is in the cinder tempest plugin, not one of ours | |
| 22:36:56 | dansmith | looks like it's probably not | |
| 22:41:37 | dansmith | gmann: around? | |
| 22:43:08 | gmann | dansmith: hi | |
| 22:43:40 | dansmith | gmann: looks like the scenario.manager create_server in tempest is not waiting for sshableness | |
| 22:43:57 | dansmith | gmann: and the cinder tempest plugin has a test that uses that and may be part of the volume detach issue(s) | |
| 22:44:11 | dansmith | gmann: is it possible to make that create_server() always wait? | |
| 22:44:59 | gmann | this one right? https://github.com/openstack/cinder-tempest-plugin/blob/d6989d3c1a31f2bebf94f2a7a8dac1d9eb788b1b/cinder_tempest_plugin/scenario/test_volume_encrypted.py#L86 | |
| 22:45:24 | dansmith | gmann: yeah, but I'm wondering if create_server() itself can wait so anything that uses it will wait | |
| 22:45:45 | gmann | dansmith: I think it make sense to wait for SSH by default in scenario tests | |
| 22:45:48 | dansmith | gmann: the current thinking is that a lot of our volume detach problems come from trying to attach a volume before an instance is booted | |
| 22:45:58 | dansmith | gmann: ack, okay | |
| 22:46:33 | gmann | yeah, and with nature of most of scenario tests, booting server fully is always good before trying things on that | |
| 22:47:20 | dansmith | yeah, okay lemme try to do that | |
| 22:47:52 | gmann | dansmith: ok. we need to create/pass the validation resources from scenario manager which is needed for SSH/ping | |
| 22:48:46 | gmann | or may be via get_remote_client | |
| 22:49:40 | dansmith | yeah | |
| 23:48:35 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-27 | |||
| 01:29:53 | dansmith | ugh, failed 102 on that last one | |
| 01:36:19 | dansmith | er, no, I suck at numbers.. 9, and all scenario, so must be related to my tempest change | |
| 03:59:16 | opendevreview | yatin proposed openstack/nova master: [DNM] Check detatch issue with lower tb size https://review.opendev.org/c/openstack/nova/+/881690 | |
| 08:17:21 | sean-k-mooney | gibi: bauzas: fun issue https://bugs.launchpad.net/os-vif/+bug/2017868 | |
| 08:17:41 | bauzas | ack | |
| 08:20:11 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 08:20:49 | sean-k-mooney | im still considering -2ing ^ | |
| 10:02:46 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement is_luks_inside_qcow2 funtion https://review.opendev.org/c/openstack/nova/+/854030 | |
| 10:03:18 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement encryption on backingStore https://review.opendev.org/c/openstack/nova/+/870012 | |
| 14:30:55 | opendevreview | ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| 15:57:12 | gouthamr | o/ dansmith - i found a "rbd image is busy" error in a different change parented on your devstack-plugin-change: https://zuul.opendev.org/t/openstack/build/278d411484764c55829c26cec2860bee/log/controller/logs/screen-n-cpu.txt#10742 | |
| 15:58:05 | gouthamr | im still looking, but, the cephfs job runs some scenario tests where it boots up nova VMs, attaches manila shares to them... these tests are failing to delete the VMs after | |
| 15:59:05 | gouthamr | the DELETE /servers/{id} call returns with 204, we wait in a loop, the server doesn't go away within 300 seconds.. and n-cpu logs don't have an error.. | |
| 16:01:29 | gouthamr | i don't know if the "rbd.ImageBusy" exception has anything to do with this silent failure | |
| 16:02:46 | gouthamr | it might | |
| 16:08:53 | dansmith | gouthamr: yeah I think eharney mentioned that it's been a pervasive problem, but I usually see it like once or twice and not a huge raft of them in one job when there's not something else going wrong | |
| 16:09:50 | gouthamr | i think in this job it's been consistently failing.. | |
| 16:10:16 | gouthamr | maybe i am missing a config opt here that you've set elsewhere | |
| 17:20:05 | opendevreview | sean mooney proposed openstack/os-vif master: [WIP] set default qos policy https://review.opendev.org/c/openstack/os-vif/+/881751 | |
| 22:52:45 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-28 | |||
| 07:47:09 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 12:09:24 | opendevreview | Stephen Finucane proposed openstack/nova master: libvirt: Remove unnecessary arg https://review.opendev.org/c/openstack/nova/+/881817 | |
| 12:14:33 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |