| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-29 | |||
| 16:30:38 | bauzas | timeouts are generally an infra issue | |
| 16:30:43 | bauzas | from a ci provider | |
| 16:30:50 | bauzas | but "generally" | |
| 16:31:04 | bauzas | which means sometimes we may have a larger problem | |
| 16:31:07 | sean-k-mooney | auniyal: do you have a gerrit link to a change where it happend | |
| 16:31:12 | dansmith | are they fips jobs? | |
| 16:31:31 | sean-k-mooney | oh ya it could be that did we add the extra 30 mins ot the job yet | |
| 16:31:31 | clarkb | bauzas: I'm not sure I agree with that statement | |
| 16:31:38 | clarkb | we have significant amounts of very inefficient test payload | |
| 16:31:55 | clarkb | yes slow providers make that worse, but we have lots of ability to improve things in the jobs just about every time I look | |
| 16:32:15 | sean-k-mooney | clarkb: we dont often see timeouts in the jobs that run on the nova gate | |
| 16:32:29 | sean-k-mooney | we tent to be well within the job timeout interval | |
| 16:32:45 | sean-k-mooney | that is not nessialy the same for other projects | |
| 16:32:46 | clarkb | (it is common for tempets jobs to dig into swap which slows everything down, devstack uses osc which is super slow because it gets a new token for every request and has python spin up time, ansible loops are costly with large numbers of entries and so on) | |
| 16:32:58 | auniyal | sean, I am trying to find a link but its time taking | |
| 16:33:04 | clarkb | sean-k-mooney: yes swap is a common cause for the difference in behaviors and that isn't an infra issue | |
| 16:33:15 | clarkb | sean-k-mooney: and devstack runtime could be ~halved if we stopped using osc | |
| 16:33:25 | clarkb | or improved osc's startup and token acquisition time | |
| 16:33:31 | sean-k-mooney | clarkb: ack | |
| 16:33:36 | clarkb | I just want to avoid the idea its an infra issue so ignore it | |
| 16:33:39 | sean-k-mooney | ya the osc thing is a long runing known issue | |
| 16:33:49 | clarkb | this assertion gets made often then I go looking and there is plenty of job payload that is just slow | |
| 16:33:55 | bauzas | clarkb: sorry, I was unclear | |
| 16:33:55 | sean-k-mooney | the parallel improments dansmith did helped indirectly | |
| 16:33:59 | auniyal | although, I have experinced this alot, if my zuul, is not passing at night time (IST), even after recheck I ran them in morning, then pass | |
| 16:34:10 | bauzas | clarkb: I wasn't advocating about someone else's fault | |
| 16:34:52 | bauzas | clarkb: I was just explaining to some new nova contributor that given the current situation, we only have timeouts with nova jobs due to some ci provider issue | |
| 16:35:04 | clarkb | bauzas: right I disagree with that | |
| 16:35:13 | clarkb | jobs timeout due to an accumulation of slow steps | |
| 16:35:13 | bauzas | clarkb: but I agree with you on some jobs that are wasting resources | |
| 16:35:21 | clarkb | some of those may be due to a slow provider or slow instance | |
| 16:35:30 | clarkb | but, it is extremely rare that this is the only problem | |
| 16:35:34 | sean-k-mooney | clarkb: we tend to be seeing an avgerate runtime at about 75% or less of the job timeout in my experince | |
| 16:35:42 | clarkb | and I know nova tempest jobs have a large number of other slowness problems | |
| 16:35:55 | clarkb | sean-k-mooney: yes, but if a job digs deeply into swap its all downhill from there | |
| 16:35:56 | sean-k-mooney | we have 2 hour timeouts on our tempest jobs and we rarly go above about 90 mins | |
| 16:35:56 | bauzas | clarkb: that's a fair point | |
| 16:36:05 | clarkb | suddenly your 75% typical runtime can balloon to 200% | |
| 16:36:08 | bauzas | except the fips one | |
| 16:36:14 | sean-k-mooney | clarkb: sure but i dont think we are | |
| 16:36:25 | sean-k-mooney | but its somethign we can look at | |
| 16:36:40 | sean-k-mooney | auniyal: the best thing you can do is provide us an example and we can look into it | |
| 16:36:46 | clarkb | ++ to looking at it | |
| 16:36:47 | sean-k-mooney | and then see if there is a trend | |
| 16:37:10 | auniyal | ack | |
| 16:38:36 | bauzas | I actually wonder how we can track the trend | |
| 16:38:50 | sean-k-mooney | https://zuul.openstack.org/builds?project=openstack%2Fnova&result=TIMED_OUT&skip=0 | |
| 16:38:56 | sean-k-mooney | that but its currently loading | |
| 16:39:17 | sean-k-mooney | we have a couple every few days | |
| 16:39:18 | bauzas | sure, but you don't have the time a SUCCESS job runs | |
| 16:39:28 | bauzas | which is what we should track | |
| 16:39:29 | clarkb | you can show both success and timeouts in a listing | |
| 16:39:35 | clarkb | (and failures, etc) | |
| 16:39:37 | sean-k-mooney | well we can change the result to filter both | |
| 16:39:51 | bauzas | the duration field, shit, missed it | |
| 16:40:01 | sean-k-mooney | we also hav ento fixed the fips job | |
| 16:40:09 | sean-k-mooney | ill create a patch for that i think | |
| 16:40:16 | bauzas | sean-k-mooney: I said I should do it | |
| 16:40:27 | sean-k-mooney | bauzas: ok please do | |
| 16:40:33 | bauzas | sean-k-mooney: that's simple to do and that's like 4 weeks I promised it | |
| 16:41:13 | bauzas | sean-k-mooney: you know what ? I'll end this meeting by now so everyone can do what they want, including me writing a zuul patch :) | |
| 16:41:25 | sean-k-mooney | :) | |
| 16:41:35 | bauzas | having said it, | |
| 16:41:39 | bauzas | thanks folks | |
| 16:41:43 | opendevmeet | Log: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.log.html | |
| 16:41:43 | opendevmeet | Minutes (text): https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.txt | |
| 16:41:43 | opendevmeet | Minutes: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.html | |
| 16:41:43 | opendevmeet | Meeting ended Tue Nov 29 16:41:43 2022 UTC. Information about MeetBot at http://wiki.debian.org/MeetBot . (v 0.1.4) | |
| 16:41:43 | bauzas | #endmeeting | |
| 16:42:03 | gibi | o/ | |
| 16:42:22 | chateaulav | o/ | |
| 16:42:29 | elodilles | o/ | |
| 16:44:07 | bauzas | like, https://zuul.openstack.org/job/tempest-centos9-stream-fips | |
| 16:48:28 | clarkb | sean-k-mooney: one thing I've found is fairly consistent for our systemic job timeouts is that we've got a number of steps to each job: base setup (configuring mirrors, configuring ssh keys, configuring git repos), test env setup (tox/devstack/whatever), actual testing, and log collection. Each of these tends to have unnecessary slow bits that add up over the course of a job. | |
| 16:48:30 | clarkb | When you then run into something like a slower node or swapping or slowness fetching an external resource it is very easy to tip over the timeout | |
| 16:49:00 | clarkb | sean-k-mooney: imo it would be helpful for us to try and whittle away at that accumulated slowness tech debt to make us less susceptible when we run into an overall slower situation. | |
| 16:49:27 | clarkb | I've worked on that a bit in the base jobs and log collection side of things as that has broad impact. But the downsides here are that it has broad impact so I have to be extremely careful to maintain backward compatibility | |
| 16:49:44 | clarkb | but the same approaches can be taken to improve things like devstack (did you know it installs tempest 3 times!) | |
| 16:50:26 | clarkb | I think improving memory consumption would also help avoid slowness caused by swapping. privsep is a fairly outsized offender here | |
| 16:53:16 | bauzas | clarkb: I'm curious about privsep being memory greedy | |
| 16:53:30 | bauzas | and I wonder why | |
| 16:53:43 | clarkb | I suspect because it grows buffers to handle all the input and output sent through it | |
| 16:54:05 | clarkb | one way to maybe improve things is to stop running a different privsep for each service whcih creates a bunch of large buffers. We might be able to get away with one large buffer | |
| 16:54:05 | bauzas | our internal customers haven't reported such problem, but I guess because of lack of evidence rather than not having a problem | |
| 16:54:29 | clarkb | or buffer things with intentionally smaller buffers | |
| 16:54:40 | bauzas | agreed | |
| 16:54:48 | bauzas | a stream is costly | |
| 16:57:01 | sean-k-mooney | clarkb: yep although we have enough fo a buffer in the nova project that we get a time out failure only 1 or twice a week | |
| 16:57:18 | clarkb | sean-k-mooney: yes, but you've also set your timeout to two hours | |
| 16:57:36 | clarkb | (one hour was the goal once upon a time) | |
| 16:57:50 | sean-k-mooney | ack | |
| 16:58:03 | sean-k-mooney | shareing privesep is a security issue | |
| 16:58:09 | sean-k-mooney | so i dont think we can ever do that | |
| 16:58:22 | sean-k-mooney | nova will have more privespe deamons eventurlaly | |
| 16:59:16 | clarkb | does privsep prevent random processes from connecting to it? If not this is equivalent. If so it could also apply restrictions on what a specific process can do (granted this is maybe a larger attack surface than refusing to talk at all) | |
| 16:59:42 | sean-k-mooney | you are ment to use file permissiosn to limit access at the file system level | |
| 16:59:53 | sean-k-mooney | but no | |
| 17:00:00 | sean-k-mooney | not as part of privsep itself | |
| 17:00:25 | sean-k-mooney | clarkb: you need one privsep deamon per privsep context currently | |
| 17:00:42 | sean-k-mooney | for it to proplry provide the correct permission enforcement/escalation | |