Earlier  
Posted Nick Remark
#openstack-nova - 2021-06-10
09:02:08 gibi stephenfin: if it lands then we are done
09:02:21 gibi if bauzas's patch lands first, then we revert that when yours land
09:02:33 bauzas gibi: stephenfin: patch is up for just removing mypy run until we fix the requirements
09:02:39 stephenfin that makes no sense to me though
09:02:46 stephenfin the mypy run is causing the other failures
09:02:48 stephenfin *isn't
09:03:01 bauzas stephenfin: sure
09:03:08 stephenfin so they have an equal chance of failing randomly
09:03:23 bauzas I'm just proposing this one because I guess nova-next WONT be running on my patch
09:03:35 bauzas it's just a tox.ini change
09:03:39 stephenfin but what does this achieve?
09:03:39 gibi stephenfin: not equal chance iff the tox.ini change does not trigger the unstable jobs
09:03:56 gibi stephenfin: if it trigger the same job, then I agree that it mypy removal patch is pointless
09:04:08 stephenfin even if it doesn't, so what?
09:04:23 bauzas gibi: stephenfin: yup, indeed, if this runs the same jobs, then nevermind my one
09:04:53 bauzas stephenfin: I was off yesterday but from what I've seen our gate is blocked
09:05:03 bauzas I just want to unblock it asap
09:05:10 stephenfin right, but it's blocked because of the flaky tests
09:05:16 stephenfin not because of the mypy thing
09:05:16 bauzas the clean fix requires a requirement bump
09:05:22 stephenfin we have a fix for the mypy thing
09:05:29 bauzas stephenfin: yup, I saw it
09:05:43 bauzas stephenfin: I'm just trying to see whether we can land things easier
09:06:00 bauzas again, unblocking the gate seems to me the most important
09:06:11 bauzas we could revert stuff later
09:06:16 stephenfin but you still won't be able to land anything that causes nova-next or nova-grenade-mutlinode to run?
09:06:20 stephenfin so it's not unblocked
09:06:38 bauzas stephenfin: that's a classic chicken-and-egg issue
09:06:47 bauzas you have 2 unrelated gate issues
09:06:48 stephenfin I mean, if they're flaky then they're flaky for everything, surely?
09:07:07 bauzas stephenfin: I don't disagree
09:07:16 bauzas 20% is a high rate of flakiness
09:07:18 stephenfin then we achieve nothing with this
09:07:45 gibi for me moving from a full red due to pep8 (mypy) failure to a falky red due to unstable jobs is still progress
09:07:47 bauzas except we go from 100% of failures to a random 20% from what I understando
09:07:52 bauzas this
09:08:40 bauzas either way, I proposed but I don't have opinions
09:08:59 bauzas the last call, tho, is that I need to get my kids in a min
09:09:59 bauzas https://zuul.opendev.org/t/openstack/status#795744
09:10:14 bauzas we don't run the flaky jobs
09:10:52 bauzas 30 mins roughly sorry
09:11:00 stephenfin idk, if the failure rate is that high that we can't land a simple reqs patch, that would suggest we need to be working on the other jobs. I see lyarwood has already been looking. I can start now too
09:11:39 bauzas stephenfin: yup, we need to parallelize efforts, I don't disagree
09:12:01 bauzas my patch is just a hack
09:12:17 bauzas and we need to consider themigrate and next jobs as the top prio
09:12:22 bauzas I absolutely don't disagree
09:12:43 bauzas and I also absolutely agree we need to land the types-paramiko changes
09:13:18 bauzas it's just, again, a way to unblock the gate even if flakey with the other jobs
09:13:21 opendevreview Victor Coutellier proposed openstack/nova master: Allow configuration of direct-snapshot feature https://review.opendev.org/c/openstack/nova/+/794837
09:50:24 gibi the requirement patch failed again. I'm looking at the grenade failure...
09:50:36 gibi https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html
09:55:04 elodilles gibi: sorry, meanwhile i've commented 'recheck' on it :X
09:55:11 gibi elodilles: no worries
09:55:15 gibi the failure is unrelated
09:55:32 gibi but based on this morining discussion above I feel we need to stabilize our gate
09:55:54 bauzas I'm back
09:56:52 bauzas gibi: stephenfin disagrees with the quick fix approach of removing mypy so I don't want to opiniate here
09:57:25 bauzas what i agree tho is that the gate failures are a PITA that need other pair of eyes
09:57:46 bauzas do we have a bug for tracking the nova-next and other jobs issues ?
09:59:00 bauzas we probably need other kinda workaround change for those jobs I guess
09:59:27 gibi bauzas: lyarwood promised to file a bug for the libvirt.libvirtError: unable to connect to server at 'ubuntu-focal-rax-dfw-0025050041:49152': Connection refused
09:59:30 gibi case
09:59:40 gibi I'm looking at the last grenade failure
09:59:50 gibi which is a cinder volume detach timeout
10:00:00 gibi at least seem so far
10:00:06 bauzas humpf, yet another focal detach saga ?
10:01:00 admin1 hi all .. checking if anyone knows this .. what is the percentage latency/performance difference between underlying filesystem vs nova ephemeral disks on top of it .. ( qcow2) and if there is a way to speed up iops performance by changing the format to raw for the vms .. and if such way exist in openstack ?
10:01:37 elodilles this volume failure seems quite frequent (however, mostly with volume stuck in 'in-use' state)
10:01:40 elodilles http://logstash.openstack.org/#/dashboard/file/logstash.json?query=message:%5C%22failed%20to%20reach%20available%20status%20(current%5C%22
10:02:33 gibi elodilles: both in-use and detaching state happens
10:02:34 bauzas elodilles: gibi: which jobs are we talking about ?
10:02:55 bauzas I see nova-live-migrate and nova-ceph-multistore, right?
10:02:56 gibi bauzas: I'm looking at the last failure in https://review.opendev.org/c/openstack/nova/+/795533 which is in nova-grenade-multinode
10:03:04 bauzas holy snap
10:03:47 bauzas I was about to propose to make some failing jobs non-voting until we identify a proper fix, but if that's an issue occurring on a large set of jobs, neverming this proposal
10:04:18 bauzas we're in a fscking bad situation :/
10:06:18 bauzas elodilles: the in-use issue seems to not happen for a while
10:09:59 gibi I see both
10:10:10 gibi in http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22failed%20to%20reach%20available%20status%5C%22
10:14:05 elodilles well, i haven't realized but logstash shows only failures from 4th June
10:16:13 elodilles i mean, only that day
10:16:36 bauzas yep
10:16:43 bauzas some transient issue I think
10:16:44 gibi elodilles: I think that is a shortcoming of logstash
10:16:52 bauzas ... or this ?
10:16:53 gibi I see the same recently
10:17:16 bauzas anyway, logstash is sunsetting unfortunately
10:17:18 elodilles i also think it's some indexing issue, or something like that...
10:17:33 bauzas so we can no longer count on it :(
10:17:41 bauzas (even if i told about it)
10:33:09 gibi I cannot figure out this failure
10:33:14 gibi I lost in cinder
10:38:34 gibi I wait for lyarwood to look at it but other than that I can only file a bug on cinder
10:53:35 gibi lyarwood: o/
10:53:38 lyarwood gibi: Can you throw me a pointer to some logs?
10:53:59 gibi sure
10:55:05 gibi https://f141fb01d9c1d07df646-94aaf0771088c81abb9a09d47e91a608.ssl.cf1.rackcdn.com/795533/5/check/nova-grenade-multinode/24485b8/testr_results.html
10:55:17 gibi this is the recent grenade job failure on https://review.opendev.org/c/openstack/nova/+/795533
10:55:57 gibi but I think this 'failed to reach available status (current detaching)' and 'failed to reach available status (current in-use)' are pretty common failures
10:56:21 lyarwood right it's still the detach logic

Earlier   Later