Earlier  
Posted Nick Remark
#openstack-nova - 2022-10-11
16:00:45 bauzas heyho
16:00:50 elodilles o/
16:01:09 bauzas who's around ?
16:02:08 bauzas we can start if needed
16:02:20 bauzas hopefully, should be a quick one, as we have the PTG next week
16:02:43 elodilles ++
16:02:50 bauzas #topic Bugs (stuck/critical)
16:02:56 bauzas #info No Critical bug
16:03:01 bauzas #link https://bugs.launchpad.net/nova/+bugs?search=Search&field.status=New 4 new untriaged bugs (-1 since the last meeting)
16:03:01 gibi o/
16:03:05 dansmith o/
16:03:10 bauzas #link https://storyboard.openstack.org/#!/project/openstack/placement 26 open stories (+0 since the last meeting) in Storyboard for Placement
16:03:20 bauzas elodilles: thanks for having looked at the bugs
16:03:24 elodilles np
16:03:30 bauzas anything you would want to discuss ?
16:03:36 bauzas #info Add yourself in the team bug roster if you want to help https://etherpad.opendev.org/p/nova-bug-triage-roster
16:03:48 elodilles maybe one bug
16:03:52 elodilles (or two)
16:04:14 Uggla */
16:04:19 elodilles related to the well-know volume timeout failure
16:04:33 elodilles a new general bug was open: https://bugs.launchpad.net/nova/+bug/1992328
16:04:52 elodilles i know that there are a ton of similar bugs already
16:05:04 elodilles some are more specific and some are general
16:05:18 bauzas elodilles: do you want to close it as a duplicate ?
16:05:34 elodilles if we have the exact same somewhere, then we could
16:05:44 elodilles otherwise we can keep it open...
16:06:11 elodilles also saw that maybe somewhat coupled with this 30 days old bug: https://bugs.launchpad.net/nova/+bug/1989232
16:06:30 elodilles i mean this is also opened around the same issue
16:07:50 elodilles otherwise i didn't have time to go deep into them
16:08:05 bauzas well, I dunno
16:08:09 elodilles so any idea about what to do with these bugs are welcome
16:08:38 gibi these need to be trobule shooted to find the root cause
16:08:47 gibi but I don't have time for that
16:09:05 gibi timeout is a the visible fault but there is deeper reasons why that happnes
16:09:28 bauzas ok, so let's continue to have this bug report then
16:09:43 bauzas and if someone wants to fix it, he/she could duplicate it if needed
16:09:57 elodilles ack
16:10:59 bauzas ok, any other bug report to look at ?
16:11:11 elodilles nothing else from me
16:12:52 bauzas ok, continuing then
16:13:09 bauzas gibi: can you be the bug baton for next week ?
16:13:15 gibi lets see
16:13:17 bauzas even if next week it will be the PTG ?
16:13:30 gibi I will take it
16:13:32 bauzas ok
16:13:34 bauzas thanks
16:13:41 bauzas #info bug baton is being passed to gibi
16:13:58 bauzas (sorry, passing it to you, as I was having the baton for 2 weeks :) )
16:14:13 bauzas moving on
16:14:15 bauzas #topic Gate status
16:14:20 bauzas #link https://bugs.launchpad.net/nova/+bugs?field.tag=gate-failure Nova gate bugs
16:14:26 bauzas #link https://zuul.openstack.org/builds?project=openstack%2Fnova&project=openstack%2Fplacement&pipeline=periodic-weekly Nova&Placement periodic jobs status
16:14:28 bauzas heh :)
16:14:44 bauzas as you see, we have a timeout with the centos9-fips job
16:15:20 bauzas but I dunno for how long it was having this timeout
16:16:20 elodilles https://zuul.openstack.org/builds?job_name=tempest-centos9-stream-fips&project=openstack%2Fnova&project=openstack%2Fplacement&pipeline=periodic-weekly&skip=0
16:16:28 elodilles it seems always?
16:16:33 bauzas yes
16:17:53 bauzas I guess the owner of the job was adalee, right?.
16:18:40 bauzas this is bizarre, I don't see where the job run is timing out
16:20:01 sean-k-mooney well the default timeout i think is 2 hours
16:20:14 sean-k-mooney so it might just need a little longer sicne the fips job does a reboot
16:20:24 bauzas yeah but the job seems to have done
16:20:30 bauzas be* done
16:20:38 bauzas anyway, nothing urgent
16:20:53 bauzas this is just the fact we can't get the result
16:21:00 bauzas even if it looks tempest works fine
16:21:11 sean-k-mooney its posible it has addtional test in a post playbook
16:21:24 bauzas I don't know who wrote the job
16:21:34 bauzas but I'll try to find it
16:21:46 bauzas should be easy to find
16:22:24 clarkb https://zuul.openstack.org/build/5f6e6a2f65ee4a5e90754d07142fab9f/log/job-output.txt#23945-23946 shows where it is timing out
16:23:24 bauzas found who added it https://review.opendev.org/c/openstack/nova/+/831844
16:23:54 bauzas thanks clarkb, will look
16:24:31 sean-k-mooney ok its tempest full
16:24:47 sean-k-mooney so its two tempest runs
16:25:05 bauzas | RUN END RESULT_TIMED_OUT: [untrusted : opendev.org/openstack/tempest/playbooks/devstack-tempest.yaml@master]
16:25:09 sean-k-mooney the first one completed but the second that runs the slow test and senarior tests timed out
16:25:10 bauzas hmpf
16:25:24 bauzas yup
16:25:45 sean-k-mooney so ya it just need anther say 30 mins added to the tiemout to be safe
16:26:05 sean-k-mooney it also could have just been a slow node
16:26:10 clarkb and maybe double check logs to see that something didn't get stuck there due to fips
16:26:19 clarkb there is a bit of a time delta between the timeout and tempest reporting anything
16:26:38 clarkb oh its only 15 seconds nevermind
16:26:42 sean-k-mooney yep
16:26:47 sean-k-mooney there isnt really a break in the logs
16:26:53 clarkb I mathed wrong the first time
16:27:08 bauzas yes, this is just a slow test
16:27:42 bauzas but should we modify the timeout for that job to be larger ?
16:27:46 sean-k-mooney i dont know how much time the reboot adds for fips but 2 hours is proably borderlien
16:28:01 bauzas we can DNM this at least
16:28:16 bauzas and verify if adding more timeout time helps
16:28:17 sean-k-mooney i would add 30mins and monitor it and see roughly how long it takes over the next few weeks
16:28:23 bauzas right
16:28:46 bauzas I can propose a patch
16:28:51 sean-k-mooney ack
16:29:21 bauzas #action bauzas to add a 30min more timeout for the centos9-fips periodic job so we will see whether it fixes the timeout
16:29:27 bauzas moving on
16:29:38 bauzas #info Please look at the gate failures and file a bug report with the gate-failure tag.

Earlier   Later