Earlier  
Posted Nick Remark
#openstack-nova - 2022-11-29
16:28:15 sean-k-mooney ah ok will check after meeting
16:28:19 bauzas sean-k-mooney: you'll get the link to the gerrit change
16:28:20 sean-k-mooney nothing else form me
16:28:29 auniyal hand-raise: zuul frequent timeout issue/fails - this seems to be resource issue, is it possible zuull resource can be increased ?
16:29:09 bauzas sean-k-mooney: tl;dr: yes we skipped the related tests but maybe they are actually not needed as you said
16:29:16 sean-k-mooney auniyal: not really timeout are not that common in our jobs
16:29:24 bauzas auniyal: see what I said above, we had problems with the gate very recently
16:29:26 sean-k-mooney auniyal: do you have an example
16:29:31 auniyal in morning when there are less number of jobs running if we run same, job gets passed
16:29:40 auniyal like less then 20
16:29:47 auniyal right now 60 jobs are running
16:29:52 sean-k-mooney that should not really be a thing
16:30:03 sean-k-mooney unless we have issues with our ci providers
16:30:17 bauzas auniyal: if you speak about job results telling timeouts, agreed with sean-k-mooney, you should tell which ones so we could investigate
16:30:24 sean-k-mooney we ocationally have issues with slow providers but its not normally coralated with the number of runnign jobs
16:30:29 bauzas yup
16:30:35 auniyal ack
16:30:38 bauzas timeouts are generally an infra issue
16:30:43 bauzas from a ci provider
16:30:50 bauzas but "generally"
16:31:04 bauzas which means sometimes we may have a larger problem
16:31:07 sean-k-mooney auniyal: do you have a gerrit link to a change where it happend
16:31:12 dansmith are they fips jobs?
16:31:31 clarkb bauzas: I'm not sure I agree with that statement
16:31:31 sean-k-mooney oh ya it could be that did we add the extra 30 mins ot the job yet
16:31:38 clarkb we have significant amounts of very inefficient test payload
16:31:55 clarkb yes slow providers make that worse, but we have lots of ability to improve things in the jobs just about every time I look
16:32:15 sean-k-mooney clarkb: we dont often see timeouts in the jobs that run on the nova gate
16:32:29 sean-k-mooney we tent to be well within the job timeout interval
16:32:45 sean-k-mooney that is not nessialy the same for other projects
16:32:46 clarkb (it is common for tempets jobs to dig into swap which slows everything down, devstack uses osc which is super slow because it gets a new token for every request and has python spin up time, ansible loops are costly with large numbers of entries and so on)
16:32:58 auniyal sean, I am trying to find a link but its time taking
16:33:04 clarkb sean-k-mooney: yes swap is a common cause for the difference in behaviors and that isn't an infra issue
16:33:15 clarkb sean-k-mooney: and devstack runtime could be ~halved if we stopped using osc
16:33:25 clarkb or improved osc's startup and token acquisition time
16:33:31 sean-k-mooney clarkb: ack
16:33:36 clarkb I just want to avoid the idea its an infra issue so ignore it
16:33:39 sean-k-mooney ya the osc thing is a long runing known issue
16:33:49 clarkb this assertion gets made often then I go looking and there is plenty of job payload that is just slow
16:33:55 sean-k-mooney the parallel improments dansmith did helped indirectly
16:33:55 bauzas clarkb: sorry, I was unclear
16:33:59 auniyal although, I have experinced this alot, if my zuul, is not passing at night time (IST), even after recheck I ran them in morning, then pass
16:34:10 bauzas clarkb: I wasn't advocating about someone else's fault
16:34:52 bauzas clarkb: I was just explaining to some new nova contributor that given the current situation, we only have timeouts with nova jobs due to some ci provider issue
16:35:04 clarkb bauzas: right I disagree with that
16:35:13 bauzas clarkb: but I agree with you on some jobs that are wasting resources
16:35:13 clarkb jobs timeout due to an accumulation of slow steps
16:35:21 clarkb some of those may be due to a slow provider or slow instance
16:35:30 clarkb but, it is extremely rare that this is the only problem
16:35:34 sean-k-mooney clarkb: we tend to be seeing an avgerate runtime at about 75% or less of the job timeout in my experince
16:35:42 clarkb and I know nova tempest jobs have a large number of other slowness problems
16:35:55 clarkb sean-k-mooney: yes, but if a job digs deeply into swap its all downhill from there
16:35:56 bauzas clarkb: that's a fair point
16:35:56 sean-k-mooney we have 2 hour timeouts on our tempest jobs and we rarly go above about 90 mins
16:36:05 clarkb suddenly your 75% typical runtime can balloon to 200%
16:36:08 bauzas except the fips one
16:36:14 sean-k-mooney clarkb: sure but i dont think we are
16:36:25 sean-k-mooney but its somethign we can look at
16:36:40 sean-k-mooney auniyal: the best thing you can do is provide us an example and we can look into it
16:36:46 clarkb ++ to looking at it
16:36:47 sean-k-mooney and then see if there is a trend
16:37:10 auniyal ack
16:38:36 bauzas I actually wonder how we can track the trend
16:38:50 sean-k-mooney https://zuul.openstack.org/builds?project=openstack%2Fnova&result=TIMED_OUT&skip=0
16:38:56 sean-k-mooney that but its currently loading
16:39:17 sean-k-mooney we have a couple every few days
16:39:18 bauzas sure, but you don't have the time a SUCCESS job runs
16:39:28 bauzas which is what we should track
16:39:29 clarkb you can show both success and timeouts in a listing
16:39:35 clarkb (and failures, etc)
16:39:37 sean-k-mooney well we can change the result to filter both
16:39:51 bauzas the duration field, shit, missed it
16:40:01 sean-k-mooney we also hav ento fixed the fips job
16:40:09 sean-k-mooney ill create a patch for that i think
16:40:16 bauzas sean-k-mooney: I said I should do it
16:40:27 sean-k-mooney bauzas: ok please do
16:40:33 bauzas sean-k-mooney: that's simple to do and that's like 4 weeks I promised it
16:41:13 bauzas sean-k-mooney: you know what ? I'll end this meeting by now so everyone can do what they want, including me writing a zuul patch :)
16:41:25 sean-k-mooney :)
16:41:35 bauzas having said it,
16:41:39 bauzas thanks folks
16:41:43 bauzas #endmeeting
16:41:43 opendevmeet Meeting ended Tue Nov 29 16:41:43 2022 UTC. Information about MeetBot at http://wiki.debian.org/MeetBot . (v 0.1.4)
16:41:43 opendevmeet Minutes: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.html
16:41:43 opendevmeet Minutes (text): https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.txt
16:41:43 opendevmeet Log: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.log.html
16:42:03 gibi o/
16:42:22 chateaulav o/
16:42:29 elodilles o/
16:44:07 bauzas like, https://zuul.openstack.org/job/tempest-centos9-stream-fips
16:48:28 clarkb sean-k-mooney: one thing I've found is fairly consistent for our systemic job timeouts is that we've got a number of steps to each job: base setup (configuring mirrors, configuring ssh keys, configuring git repos), test env setup (tox/devstack/whatever), actual testing, and log collection. Each of these tends to have unnecessary slow bits that add up over the course of a job.
16:48:30 clarkb When you then run into something like a slower node or swapping or slowness fetching an external resource it is very easy to tip over the timeout
16:49:00 clarkb sean-k-mooney: imo it would be helpful for us to try and whittle away at that accumulated slowness tech debt to make us less susceptible when we run into an overall slower situation.
16:49:27 clarkb I've worked on that a bit in the base jobs and log collection side of things as that has broad impact. But the downsides here are that it has broad impact so I have to be extremely careful to maintain backward compatibility
16:49:44 clarkb but the same approaches can be taken to improve things like devstack (did you know it installs tempest 3 times!)
16:50:26 clarkb I think improving memory consumption would also help avoid slowness caused by swapping. privsep is a fairly outsized offender here
16:53:16 bauzas clarkb: I'm curious about privsep being memory greedy
16:53:30 bauzas and I wonder why
16:53:43 clarkb I suspect because it grows buffers to handle all the input and output sent through it
16:54:05 bauzas our internal customers haven't reported such problem, but I guess because of lack of evidence rather than not having a problem

Earlier   Later