Earlier  
Posted Nick Remark
#openstack-nova - 2022-11-29
16:21:11 bauzas #agreed Dec-14th will be a spec review day and Jan-10th will be an implementation review day, mark your calendars
16:21:41 bauzas #action bauzas to send an email about it
16:22:16 bauzas #agreed Some nova-cores can review some features changes around Dec 15th, you now know about it
16:22:27 gibi :)
16:22:28 bauzas OK, that's it
16:22:43 bauzas moving on
16:22:50 bauzas (sorry, that was a long discussion)
16:22:54 bauzas #topic Review priorities
16:23:00 bauzas #link https://review.opendev.org/q/status:open+(project:openstack/nova+OR+project:openstack/placement+OR+project:openstack/os-traits+OR+project:openstack/os-resource-classes+OR+project:openstack/os-vif+OR+project:openstack/python-novaclient+OR+project:openstack/osc-placement)+(label:Review-Priority%252B1+OR+label:Review-Priority%252B2)
16:23:05 bauzas #info As a reminder, cores eager to review changes can +1 to indicate their interest, +2 for committing to the review
16:23:30 bauzas I'm happy to see people using it
16:23:56 bauzas that's it for that topic
16:24:00 bauzas next one
16:24:07 bauzas #topic Stable Branches
16:24:13 bauzas elodilles: your turn
16:24:16 elodilles ack
16:24:20 elodilles this will be short
16:24:23 elodilles #info stable branches seem to be unblocked / OK
16:24:27 elodilles #info stable branch status / gate failures tracking etherpad: https://etherpad.opendev.org/p/nova-stable-branch-ci
16:24:30 elodilles that's it
16:25:58 gibi nice
16:26:14 bauzas was quick and awesome
16:26:36 bauzas last topic but not the least in theory,
16:26:45 bauzas #topic Open discussion
16:26:55 bauzas nothing in the wikipage
16:26:58 bauzas so
16:27:04 bauzas anything to discuss here by now ?
16:27:08 gibi -
16:27:17 sean-k-mooney did you merge skipign the failing nova-lvm tests yet
16:27:26 sean-k-mooney or is the master gate still explodingon that
16:27:30 bauzas I think yesterday we said we could discuss during this meeting about the test skips
16:27:46 bauzas but given we merged gmann's patch, the ship has sailed
16:27:56 sean-k-mooney ack
16:28:02 sean-k-mooney so they are disabeled currently
16:28:05 bauzas sean-k-mooney: see my ML thread above ^
16:28:06 sean-k-mooney the failing detach tests
16:28:15 sean-k-mooney ah ok will check after meeting
16:28:19 bauzas sean-k-mooney: you'll get the link to the gerrit change
16:28:20 sean-k-mooney nothing else form me
16:28:29 auniyal hand-raise: zuul frequent timeout issue/fails - this seems to be resource issue, is it possible zuull resource can be increased ?
16:29:09 bauzas sean-k-mooney: tl;dr: yes we skipped the related tests but maybe they are actually not needed as you said
16:29:16 sean-k-mooney auniyal: not really timeout are not that common in our jobs
16:29:24 bauzas auniyal: see what I said above, we had problems with the gate very recently
16:29:26 sean-k-mooney auniyal: do you have an example
16:29:31 auniyal in morning when there are less number of jobs running if we run same, job gets passed
16:29:40 auniyal like less then 20
16:29:47 auniyal right now 60 jobs are running
16:29:52 sean-k-mooney that should not really be a thing
16:30:03 sean-k-mooney unless we have issues with our ci providers
16:30:17 bauzas auniyal: if you speak about job results telling timeouts, agreed with sean-k-mooney, you should tell which ones so we could investigate
16:30:24 sean-k-mooney we ocationally have issues with slow providers but its not normally coralated with the number of runnign jobs
16:30:29 bauzas yup
16:30:35 auniyal ack
16:30:38 bauzas timeouts are generally an infra issue
16:30:43 bauzas from a ci provider
16:30:50 bauzas but "generally"
16:31:04 bauzas which means sometimes we may have a larger problem
16:31:07 sean-k-mooney auniyal: do you have a gerrit link to a change where it happend
16:31:12 dansmith are they fips jobs?
16:31:31 sean-k-mooney oh ya it could be that did we add the extra 30 mins ot the job yet
16:31:31 clarkb bauzas: I'm not sure I agree with that statement
16:31:38 clarkb we have significant amounts of very inefficient test payload
16:31:55 clarkb yes slow providers make that worse, but we have lots of ability to improve things in the jobs just about every time I look
16:32:15 sean-k-mooney clarkb: we dont often see timeouts in the jobs that run on the nova gate
16:32:29 sean-k-mooney we tent to be well within the job timeout interval
16:32:45 sean-k-mooney that is not nessialy the same for other projects
16:32:46 clarkb (it is common for tempets jobs to dig into swap which slows everything down, devstack uses osc which is super slow because it gets a new token for every request and has python spin up time, ansible loops are costly with large numbers of entries and so on)
16:32:58 auniyal sean, I am trying to find a link but its time taking
16:33:04 clarkb sean-k-mooney: yes swap is a common cause for the difference in behaviors and that isn't an infra issue
16:33:15 clarkb sean-k-mooney: and devstack runtime could be ~halved if we stopped using osc
16:33:25 clarkb or improved osc's startup and token acquisition time
16:33:31 sean-k-mooney clarkb: ack
16:33:36 clarkb I just want to avoid the idea its an infra issue so ignore it
16:33:39 sean-k-mooney ya the osc thing is a long runing known issue
16:33:49 clarkb this assertion gets made often then I go looking and there is plenty of job payload that is just slow
16:33:55 bauzas clarkb: sorry, I was unclear
16:33:55 sean-k-mooney the parallel improments dansmith did helped indirectly
16:33:59 auniyal although, I have experinced this alot, if my zuul, is not passing at night time (IST), even after recheck I ran them in morning, then pass
16:34:10 bauzas clarkb: I wasn't advocating about someone else's fault
16:34:52 bauzas clarkb: I was just explaining to some new nova contributor that given the current situation, we only have timeouts with nova jobs due to some ci provider issue
16:35:04 clarkb bauzas: right I disagree with that
16:35:13 clarkb jobs timeout due to an accumulation of slow steps
16:35:13 bauzas clarkb: but I agree with you on some jobs that are wasting resources
16:35:21 clarkb some of those may be due to a slow provider or slow instance
16:35:30 clarkb but, it is extremely rare that this is the only problem
16:35:34 sean-k-mooney clarkb: we tend to be seeing an avgerate runtime at about 75% or less of the job timeout in my experince
16:35:42 clarkb and I know nova tempest jobs have a large number of other slowness problems
16:35:55 clarkb sean-k-mooney: yes, but if a job digs deeply into swap its all downhill from there
16:35:56 sean-k-mooney we have 2 hour timeouts on our tempest jobs and we rarly go above about 90 mins
16:35:56 bauzas clarkb: that's a fair point
16:36:05 clarkb suddenly your 75% typical runtime can balloon to 200%
16:36:08 bauzas except the fips one
16:36:14 sean-k-mooney clarkb: sure but i dont think we are
16:36:25 sean-k-mooney but its somethign we can look at
16:36:40 sean-k-mooney auniyal: the best thing you can do is provide us an example and we can look into it
16:36:46 clarkb ++ to looking at it
16:36:47 sean-k-mooney and then see if there is a trend
16:37:10 auniyal ack
16:38:36 bauzas I actually wonder how we can track the trend
16:38:50 sean-k-mooney https://zuul.openstack.org/builds?project=openstack%2Fnova&result=TIMED_OUT&skip=0

Earlier   Later