| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-06-10 | |||
| 09:07:16 | bauzas | 20% is a high rate of flakiness | |
| 09:07:18 | stephenfin | then we achieve nothing with this | |
| 09:07:45 | gibi | for me moving from a full red due to pep8 (mypy) failure to a falky red due to unstable jobs is still progress | |
| 09:07:47 | bauzas | except we go from 100% of failures to a random 20% from what I understando | |
| 09:07:52 | bauzas | this | |
| 09:08:40 | bauzas | either way, I proposed but I don't have opinions | |
| 09:08:59 | bauzas | the last call, tho, is that I need to get my kids in a min | |
| 09:09:59 | bauzas | https://zuul.opendev.org/t/openstack/status#795744 | |
| 09:10:14 | bauzas | we don't run the flaky jobs | |
| 09:10:52 | bauzas | 30 mins roughly sorry | |
| 09:11:00 | stephenfin | idk, if the failure rate is that high that we can't land a simple reqs patch, that would suggest we need to be working on the other jobs. I see lyarwood has already been looking. I can start now too | |
| 09:11:39 | bauzas | stephenfin: yup, we need to parallelize efforts, I don't disagree | |
| 09:12:01 | bauzas | my patch is just a hack | |
| 09:12:17 | bauzas | and we need to consider themigrate and next jobs as the top prio | |
| 09:12:22 | bauzas | I absolutely don't disagree | |
| 09:12:43 | bauzas | and I also absolutely agree we need to land the types-paramiko changes | |
| 09:13:18 | bauzas | it's just, again, a way to unblock the gate even if flakey with the other jobs | |
| 09:13:21 | opendevreview | Victor Coutellier proposed openstack/nova master: Allow configuration of direct-snapshot feature https://review.opendev.org/c/openstack/nova/+/794837 | |
| 09:50:24 | gibi | the requirement patch failed again. I'm looking at the grenade failure... | |
| 09:50:36 | gibi | https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html | |
| 09:55:04 | elodilles | gibi: sorry, meanwhile i've commented 'recheck' on it :X | |
| 09:55:11 | gibi | elodilles: no worries | |
| 09:55:15 | gibi | the failure is unrelated | |
| 09:55:32 | gibi | but based on this morining discussion above I feel we need to stabilize our gate | |
| 09:55:54 | bauzas | I'm back | |
| 09:56:52 | bauzas | gibi: stephenfin disagrees with the quick fix approach of removing mypy so I don't want to opiniate here | |
| 09:57:25 | bauzas | what i agree tho is that the gate failures are a PITA that need other pair of eyes | |
| 09:57:46 | bauzas | do we have a bug for tracking the nova-next and other jobs issues ? | |
| 09:59:00 | bauzas | we probably need other kinda workaround change for those jobs I guess | |
| 09:59:27 | gibi | bauzas: lyarwood promised to file a bug for the libvirt.libvirtError: unable to connect to server at 'ubuntu-focal-rax-dfw-0025050041:49152': Connection refused | |
| 09:59:30 | gibi | case | |
| 09:59:40 | gibi | I'm looking at the last grenade failure | |
| 09:59:50 | gibi | which is a cinder volume detach timeout | |
| 10:00:00 | gibi | at least seem so far | |
| 10:00:06 | bauzas | humpf, yet another focal detach saga ? | |
| 10:01:00 | admin1 | hi all .. checking if anyone knows this .. what is the percentage latency/performance difference between underlying filesystem vs nova ephemeral disks on top of it .. ( qcow2) and if there is a way to speed up iops performance by changing the format to raw for the vms .. and if such way exist in openstack ? | |
| 10:01:37 | elodilles | this volume failure seems quite frequent (however, mostly with volume stuck in 'in-use' state) | |
| 10:01:40 | elodilles | http://logstash.openstack.org/#/dashboard/file/logstash.json?query=message:%5C%22failed%20to%20reach%20available%20status%20(current%5C%22 | |
| 10:02:33 | gibi | elodilles: both in-use and detaching state happens | |
| 10:02:34 | bauzas | elodilles: gibi: which jobs are we talking about ? | |
| 10:02:55 | bauzas | I see nova-live-migrate and nova-ceph-multistore, right? | |
| 10:02:56 | gibi | bauzas: I'm looking at the last failure in https://review.opendev.org/c/openstack/nova/+/795533 which is in nova-grenade-multinode | |
| 10:03:04 | bauzas | holy snap | |
| 10:03:47 | bauzas | I was about to propose to make some failing jobs non-voting until we identify a proper fix, but if that's an issue occurring on a large set of jobs, neverming this proposal | |
| 10:04:18 | bauzas | we're in a fscking bad situation :/ | |
| 10:06:18 | bauzas | elodilles: the in-use issue seems to not happen for a while | |
| 10:09:59 | gibi | I see both | |
| 10:10:10 | gibi | in http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22failed%20to%20reach%20available%20status%5C%22 | |
| 10:14:05 | elodilles | well, i haven't realized but logstash shows only failures from 4th June | |
| 10:16:13 | elodilles | i mean, only that day | |
| 10:16:36 | bauzas | yep | |
| 10:16:43 | bauzas | some transient issue I think | |
| 10:16:44 | gibi | elodilles: I think that is a shortcoming of logstash | |
| 10:16:52 | bauzas | ... or this ? | |
| 10:16:53 | gibi | I see the same recently | |
| 10:17:16 | bauzas | anyway, logstash is sunsetting unfortunately | |
| 10:17:18 | elodilles | i also think it's some indexing issue, or something like that... | |
| 10:17:33 | bauzas | so we can no longer count on it :( | |
| 10:17:41 | bauzas | (even if i told about it) | |
| 10:33:09 | gibi | I cannot figure out this failure | |
| 10:33:14 | gibi | I lost in cinder | |
| 10:38:34 | gibi | I wait for lyarwood to look at it but other than that I can only file a bug on cinder | |
| 10:53:35 | gibi | lyarwood: o/ | |
| 10:53:38 | lyarwood | gibi: Can you throw me a pointer to some logs? | |
| 10:53:59 | gibi | sure | |
| 10:55:05 | gibi | https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html | |
| 10:55:17 | gibi | this is the recent grenade job failure on https://review.opendev.org/c/openstack/nova/+/795533 | |
| 10:55:57 | gibi | but I think this 'failed to reach available status (current detaching)' and 'failed to reach available status (current in-use)' are pretty common failures | |
| 10:56:21 | lyarwood | right it's still the detach logic | |
| 10:56:59 | lyarwood | https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/controller/logs/screen-n-cpu.txt#25665 | |
| 10:57:16 | lyarwood | I had a few hits of this last week and wanted to dump the instance console on failure | |
| 10:57:31 | lyarwood | led me down a rabbit hole that I've not had time to revisit this week | |
| 10:57:48 | lyarwood | https://review.opendev.org/c/openstack/tempest/+/794757 was my initial attempt | |
| 10:58:01 | gibi | ohh | |
| 10:58:13 | gibi | thanks for the info | |
| 10:58:23 | gibi | I went to the cinder side | |
| 10:58:26 | gibi | and lost | |
| 11:00:24 | gibi | lyarwood: what I saw that the tempest sent a volume detachment https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/job-output.txt#58720 and that led to the volume being in detaching state | |
| 11:02:05 | gibi | and that succeeded in nova https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/controller/logs/screen-n-cpu.txt#23456 | |
| 11:03:37 | gibi | ohh it does not | |
| 11:03:48 | gibi | it oly succeded from the persistent domain | |
| 11:04:10 | lyarwood | right req-6f7d27cf-3d82-4e8d-ad25-7f53249a05a0 fails to detach the volume from the live domain | |
| 11:04:17 | gibi | yeah, now I see | |
| 11:04:54 | lyarwood | I traced this all through previously and couldn't see any issues with n-cpu, libvirt or even QEMU tbh | |
| 11:05:04 | lyarwood | so I wanted to see what the state of the guest OS was | |
| 11:05:14 | lyarwood | before I reported a bug to the QEMU folks | |
| 11:05:18 | lyarwood | and/or cirros | |
| 11:07:33 | lyarwood | https://storage.gra.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_ec1/795533/5/check/nova-live-migration/ec1e40c/testr_results.html - the latest run has failed in nova-live-migration | |
| 11:07:58 | lyarwood | with a timeout during the initial POST, odd. | |
| 11:08:57 | opendevmeet | Launchpad bug 1929446 in OpenStack Compute (nova) "check_can_live_migrate_source taking > 60 seconds in CI" [Medium,Triaged] | |
| 11:08:57 | lyarwood | oh I wonder if it's https://bugs.launchpad.net/nova/+bug/1929446 | |
| 11:09:14 | lyarwood | sean-k-mooney: ^ did you get anywhere with that? | |
| 11:09:21 | bauzas | fwiw, the mypy removal is green from the CI https://review.opendev.org/c/openstack/nova/+/795744 | |
| 11:11:11 | lyarwood | It's still likely to fail in the actual gate with all of these failures however right? | |
| 11:12:45 | sean-k-mooney | i think its stil ovsdbapp but i did not see a way to stop the polling in the lib | |
| 11:13:32 | sean-k-mooney | i can take another look im wondering if we need to move the polling to the prive sep deamon or into a real pthread | |
| 11:16:16 | lyarwood | sean-k-mooney: I still don't understand the timing here | |
| 11:16:42 | lyarwood | sean-k-mooney: is there a long running os-vif thing in the background or is it related to check_can_live_migrate_source? | |
| 11:17:13 | sean-k-mooney | the first | |
| 11:17:32 | lyarwood | kk the short term workaround is just to bump the rpc timeout I guess | |