| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-17 | |||
| 16:13:40 | gibi | yep, that would be my goal of disabling it temporary to see if the OOM just moves to another test case | |
| 16:13:44 | gibi | and to see which test case | |
| 16:13:48 | gibi | to find a pattern | |
| 16:13:48 | bauzas | (I also verified that nothing changed on the tempest side since 1 year for this test) | |
| 16:14:05 | dansmith | gibi: you could also rename it I think and change the sort ordering | |
| 16:14:10 | bauzas | yeah | |
| 16:14:11 | dansmith | afaik, we run tests sorted per worker | |
| 16:14:27 | bauzas | I was wondering, maybe this was a problem due to another test | |
| 16:14:28 | gibi | dansmith: good idea | |
| 16:14:45 | bauzas | dansmith: I think we can ask stestr to modifyh the sort | |
| 16:14:55 | dansmith | oh? | |
| 16:14:59 | bauzas | but I need to remember how to do it | |
| 16:15:05 | gibi | bauzas: on that I extracted all the test cases form the killed worker from multiple runs and the only test case overlap was this tc | |
| 16:15:35 | gibi | so if other test causing the issue then it is not a single test but a set of tests | |
| 16:15:49 | gibi | otherwise I would see an overlap | |
| 16:15:51 | bauzas | gibi: well, yeah, but that maybe means that the previous tests were adding more memory before so that's only with this test that OOMkiller wants to kill | |
| 16:16:14 | bauzas | as you see, this is a very simple test | |
| 16:16:22 | gibi | that is my point above, if a single test adds the extra memory usage then that woudl show up as an overlap between runs | |
| 16:16:30 | gibi | but it doesn't | |
| 16:16:45 | bauzas | gibi: that's why I'll try to see how to ask stestr to modify the sort | |
| 16:17:12 | gibi | yeah, moving this tc to the end can help to see if there is a set of tests that trigger this behavior | |
| 16:18:07 | gibi | anyhow I think we can move on | |
| 16:19:03 | bauzas | cool | |
| 16:19:19 | bauzas | #link https://bugs.launchpad.net/nova/+bugs?search=Search&field.status=New 27 new untriaged bugs (+0 since the last meeting) | |
| 16:19:45 | bauzas | I triaged a few bugs todaty | |
| 16:19:57 | bauzas | #link https://etherpad.opendev.org/p/nova-bug-triage-20230110 | |
| 16:20:16 | bauzas | nothing to report here by now | |
| 16:20:21 | bauzas | #info Add yourself in the team bug roster if you want to help https://etherpad.opendev.org/p/nova-bug-triage-roster | |
| 16:20:28 | bauzas | gibi: wants to get the bug baton this week ? | |
| 16:21:23 | gibi | bauzas: sure I can | |
| 16:21:29 | bauzas | thanks alot | |
| 16:21:47 | bauzas | #info bug baton is being passed to gibi | |
| 16:21:53 | bauzas | #topic Gate status | |
| 16:21:58 | bauzas | #link https://bugs.launchpad.net/nova/+bugs?field.tag=gate-failure Nova gate bugs | |
| 16:22:14 | bauzas | we already discussed about the main one, wanting to discuss other CI bugs ? | |
| 16:22:38 | gibi | just a sort summary | |
| 16:22:38 | bauzas | looks not | |
| 16:22:41 | bauzas | ah | |
| 16:22:50 | bauzas | we're listening to you | |
| 16:22:53 | gibi | I see failures in our functional tests | |
| 16:23:13 | gibi | one is about missing db tables so it is probably interference between test cases | |
| 16:23:21 | gibi | we saw that before | |
| 16:23:30 | gibi | fixed it but not we had a non 100% fix | |
| 16:23:39 | bauzas | :/ | |
| 16:24:08 | gibi | and there is a failure with db cursor need a reset | |
| 16:24:14 | gibi | it might be related to the above | |
| 16:24:17 | gibi | not sure yet | |
| 16:24:47 | bauzas | lovely | |
| 16:24:55 | gibi | these two I wanted to mention | |
| 16:25:08 | gibi | but there are other open bugs that appear in the gate time to time | |
| 16:25:09 | bauzas | flipping strest worker runs would help to trigger the races | |
| 16:25:20 | gibi | so it is fairly hard to land things overall | |
| 16:25:29 | bauzas | I could try to reproduce those functests locally | |
| 16:25:45 | bauzas | this would exhaust my laptop, but worth trying | |
| 16:26:18 | bauzas | gibi: let's then discuss this tomorrow as well | |
| 16:26:23 | gibi | sure | |
| 16:26:48 | bauzas | I mean, I have my power mgmt series to work on, but if we can't land things, nothing will merge either way. | |
| 16:27:03 | sean-k-mooney | the gate is not totally blocked | |
| 16:27:12 | sean-k-mooney | but its flaky enough that its hard | |
| 16:27:15 | bauzas | yeah, but rechecking is not a great option | |
| 16:27:16 | gibi | yepp | |
| 16:27:22 | sean-k-mooney | ya its not | |
| 16:27:29 | bauzas | agreed, I'm not sending the signal our gate is busted | |
| 16:27:35 | bauzas | but we know this is hard | |
| 16:27:43 | sean-k-mooney | one thing i have noticed is the py3.10 functional job seams more stable then py38 | |
| 16:27:47 | bauzas | and let me go to the next topic and you'll understand why | |
| 16:28:01 | sean-k-mooney | for the db issues | |
| 16:28:09 | sean-k-mooney | but that could be just the ones i happend to look at | |
| 16:28:14 | bauzas | ok | |
| 16:28:22 | clarkb | sean-k-mooney: 3.10 introduced a much more deterministic thread scheduler. Also its quite a bit quicker in some projects which helps generally | |
| 16:28:32 | bauzas | ah, gdk | |
| 16:28:54 | bauzas | we probably have tests not correctly cleaning up data | |
| 16:28:58 | bauzas | so we need to bisect them | |
| 16:29:01 | sean-k-mooney | ya so im wondifing if we are blocked we might want to make the 3.8 one non voting while we try to fix this | |
| 16:29:14 | sean-k-mooney | but there are other issues so i dont think that will help much | |
| 16:29:19 | sean-k-mooney | just somethign to keep in mind | |
| 16:29:20 | bauzas | sean-k-mooney: before going that road, lemme try to bisect the faulty tests | |
| 16:29:25 | dansmith | I've seen it both waysm | |
| 16:29:26 | sean-k-mooney | yep | |
| 16:29:32 | dansmith | 3.10 passing with 3.8 failing and the other way | |
| 16:29:39 | dansmith | so I don't think disabling one gets us much | |
| 16:29:41 | sean-k-mooney | ok then its jsut flaky | |
| 16:29:50 | bauzas | lovely | |
| 16:29:53 | bauzas | moving on | |
| 16:29:59 | bauzas | we have some agenda today | |
| 16:30:09 | bauzas | #link https://zuul.openstack.org/builds?project=openstack%2Fnova&project=openstack%2Fplacement&pipeline=periodic-weekly Nova&Placement periodic jobs status | |
| 16:30:14 | bauzas | that's fun | |
| 16:30:34 | bauzas | despite https://review.opendev.org/c/openstack/tempest/+/866049 was merged, we still have the centos9-fips job timeouting | |
| 16:30:40 | bauzas | so I looked at the job def | |
| 16:30:57 | bauzas | and looks to me it no longer depends on the job I added extra timeout :) | |
| 16:31:22 | bauzas | so basically the patch that took 2 months to get landed is basically useless for our pipeline | |
| 16:31:26 | bauzas | funny, as I said | |
| 16:31:41 | gmann | I think we had progress on running fips testing on ubuntu but need to check if we have job ready. that can replace c9-fips jobs | |
| 16:31:48 | bauzas | so I'll just add the extra timeout on our local job definition | |
| 16:32:00 | opendevreview | Dan Smith proposed openstack/nova master: WIP: Detect host renames and abort startup https://review.opendev.org/c/openstack/nova/+/863920 | |
| 16:32:10 | bauzas | gmann: that's good to hear | |
| 16:32:29 | gmann | not merged yet #link https://review.opendev.org/c/openstack/project-config/+/867112 | |
| 16:32:33 | bauzas | gmann: we could put fips in check pipeline then | |
| 16:32:47 | gmann | yeah that is plan once we have ubuntu based job | |
| 16:32:55 | bauzas | gmann: as a reminder, given centos9s, fips is on periodic pipeline | |