Earlier  
Posted Nick Remark
#openstack-nova - 2021-06-10
09:07:16 bauzas 20% is a high rate of flakiness
09:07:18 stephenfin then we achieve nothing with this
09:07:45 gibi for me moving from a full red due to pep8 (mypy) failure to a falky red due to unstable jobs is still progress
09:07:47 bauzas except we go from 100% of failures to a random 20% from what I understando
09:07:52 bauzas this
09:08:40 bauzas either way, I proposed but I don't have opinions
09:08:59 bauzas the last call, tho, is that I need to get my kids in a min
09:09:59 bauzas https://zuul.opendev.org/t/openstack/status#795744
09:10:14 bauzas we don't run the flaky jobs
09:10:52 bauzas 30 mins roughly sorry
09:11:00 stephenfin idk, if the failure rate is that high that we can't land a simple reqs patch, that would suggest we need to be working on the other jobs. I see lyarwood has already been looking. I can start now too
09:11:39 bauzas stephenfin: yup, we need to parallelize efforts, I don't disagree
09:12:01 bauzas my patch is just a hack
09:12:17 bauzas and we need to consider themigrate and next jobs as the top prio
09:12:22 bauzas I absolutely don't disagree
09:12:43 bauzas and I also absolutely agree we need to land the types-paramiko changes
09:13:18 bauzas it's just, again, a way to unblock the gate even if flakey with the other jobs
09:13:21 opendevreview Victor Coutellier proposed openstack/nova master: Allow configuration of direct-snapshot feature https://review.opendev.org/c/openstack/nova/+/794837
09:50:24 gibi the requirement patch failed again. I'm looking at the grenade failure...
09:50:36 gibi https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html
09:55:04 elodilles gibi: sorry, meanwhile i've commented 'recheck' on it :X
09:55:11 gibi elodilles: no worries
09:55:15 gibi the failure is unrelated
09:55:32 gibi but based on this morining discussion above I feel we need to stabilize our gate
09:55:54 bauzas I'm back
09:56:52 bauzas gibi: stephenfin disagrees with the quick fix approach of removing mypy so I don't want to opiniate here
09:57:25 bauzas what i agree tho is that the gate failures are a PITA that need other pair of eyes
09:57:46 bauzas do we have a bug for tracking the nova-next and other jobs issues ?
09:59:00 bauzas we probably need other kinda workaround change for those jobs I guess
09:59:27 gibi bauzas: lyarwood promised to file a bug for the libvirt.libvirtError: unable to connect to server at 'ubuntu-focal-rax-dfw-0025050041:49152': Connection refused
09:59:30 gibi case
09:59:40 gibi I'm looking at the last grenade failure
09:59:50 gibi which is a cinder volume detach timeout
10:00:00 gibi at least seem so far
10:00:06 bauzas humpf, yet another focal detach saga ?
10:01:00 admin1 hi all .. checking if anyone knows this .. what is the percentage latency/performance difference between underlying filesystem vs nova ephemeral disks on top of it .. ( qcow2) and if there is a way to speed up iops performance by changing the format to raw for the vms .. and if such way exist in openstack ?
10:01:37 elodilles this volume failure seems quite frequent (however, mostly with volume stuck in 'in-use' state)
10:01:40 elodilles http://logstash.openstack.org/#/dashboard/file/logstash.json?query=message:%5C%22failed%20to%20reach%20available%20status%20(current%5C%22
10:02:33 gibi elodilles: both in-use and detaching state happens
10:02:34 bauzas elodilles: gibi: which jobs are we talking about ?
10:02:55 bauzas I see nova-live-migrate and nova-ceph-multistore, right?
10:02:56 gibi bauzas: I'm looking at the last failure in https://review.opendev.org/c/openstack/nova/+/795533 which is in nova-grenade-multinode
10:03:04 bauzas holy snap
10:03:47 bauzas I was about to propose to make some failing jobs non-voting until we identify a proper fix, but if that's an issue occurring on a large set of jobs, neverming this proposal
10:04:18 bauzas we're in a fscking bad situation :/
10:06:18 bauzas elodilles: the in-use issue seems to not happen for a while
10:09:59 gibi I see both
10:10:10 gibi in http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22failed%20to%20reach%20available%20status%5C%22
10:14:05 elodilles well, i haven't realized but logstash shows only failures from 4th June
10:16:13 elodilles i mean, only that day
10:16:36 bauzas yep
10:16:43 bauzas some transient issue I think
10:16:44 gibi elodilles: I think that is a shortcoming of logstash
10:16:52 bauzas ... or this ?
10:16:53 gibi I see the same recently
10:17:16 bauzas anyway, logstash is sunsetting unfortunately
10:17:18 elodilles i also think it's some indexing issue, or something like that...
10:17:33 bauzas so we can no longer count on it :(
10:17:41 bauzas (even if i told about it)
10:33:09 gibi I cannot figure out this failure
10:33:14 gibi I lost in cinder
10:38:34 gibi I wait for lyarwood to look at it but other than that I can only file a bug on cinder
10:53:35 gibi lyarwood: o/
10:53:38 lyarwood gibi: Can you throw me a pointer to some logs?
10:53:59 gibi sure
10:55:05 gibi https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html
10:55:17 gibi this is the recent grenade job failure on https://review.opendev.org/c/openstack/nova/+/795533
10:55:57 gibi but I think this 'failed to reach available status (current detaching)' and 'failed to reach available status (current in-use)' are pretty common failures
10:56:21 lyarwood right it's still the detach logic
10:56:59 lyarwood https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/controller/logs/screen-n-cpu.txt#25665
10:57:16 lyarwood I had a few hits of this last week and wanted to dump the instance console on failure
10:57:31 lyarwood led me down a rabbit hole that I've not had time to revisit this week
10:57:48 lyarwood https://review.opendev.org/c/openstack/tempest/+/794757 was my initial attempt
10:58:01 gibi ohh
10:58:13 gibi thanks for the info
10:58:23 gibi I went to the cinder side
10:58:26 gibi and lost
11:00:24 gibi lyarwood: what I saw that the tempest sent a volume detachment https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/job-output.txt#58720 and that led to the volume being in detaching state
11:02:05 gibi and that succeeded in nova https://zuul.opendev.org/t/openstack/build/24485b8c450740c5946c39dcf310a746/log/controller/logs/screen-n-cpu.txt#23456
11:03:37 gibi ohh it does not
11:03:48 gibi it oly succeded from the persistent domain
11:04:10 lyarwood right req-6f7d27cf-3d82-4e8d-ad25-7f53249a05a0 fails to detach the volume from the live domain
11:04:17 gibi yeah, now I see
11:04:54 lyarwood I traced this all through previously and couldn't see any issues with n-cpu, libvirt or even QEMU tbh
11:05:04 lyarwood so I wanted to see what the state of the guest OS was
11:05:14 lyarwood before I reported a bug to the QEMU folks
11:05:18 lyarwood and/or cirros
11:07:33 lyarwood https://storage.gra.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_ec1/795533/5/check/nova-live-migration/ec1e40c/testr_results.html - the latest run has failed in nova-live-migration
11:07:58 lyarwood with a timeout during the initial POST, odd.
11:08:57 opendevmeet Launchpad bug 1929446 in OpenStack Compute (nova) "check_can_live_migrate_source taking > 60 seconds in CI" [Medium,Triaged]
11:08:57 lyarwood oh I wonder if it's https://bugs.launchpad.net/nova/+bug/1929446
11:09:14 lyarwood sean-k-mooney: ^ did you get anywhere with that?
11:09:21 bauzas fwiw, the mypy removal is green from the CI https://review.opendev.org/c/openstack/nova/+/795744
11:11:11 lyarwood It's still likely to fail in the actual gate with all of these failures however right?
11:12:45 sean-k-mooney i think its stil ovsdbapp but i did not see a way to stop the polling in the lib
11:13:32 sean-k-mooney i can take another look im wondering if we need to move the polling to the prive sep deamon or into a real pthread
11:16:16 lyarwood sean-k-mooney: I still don't understand the timing here
11:16:42 lyarwood sean-k-mooney: is there a long running os-vif thing in the background or is it related to check_can_live_migrate_source?
11:17:13 sean-k-mooney the first
11:17:32 lyarwood kk the short term workaround is just to bump the rpc timeout I guess

Earlier   Later