| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-17 | |||
| 16:04:38 | gibi | sure | |
| 16:04:57 | gibi | I updated the bug | |
| 16:05:01 | bauzas | ok, so, gibi (mostly) and I looked at this one today | |
| 16:05:05 | gibi | I think it is tempest.api.compute.admin.test_volume.AttachSCSIVolumeTestJSON.test_attach_scsi_disk_with_config_drive test case that tirggers the OOM | |
| 16:05:11 | bauzas | yeah | |
| 16:05:17 | bauzas | and like I said, I tried to find wherer | |
| 16:06:17 | bauzas | but I wasn't able to see | |
| 16:06:34 | bauzas | context : https://github.com/openstack/tempest/blob/7c8b49becef78a257e2515970a552c84982f59cd/tempest/api/compute/admin/test_volume.py#L84-L120 | |
| 16:06:46 | bauzas | we try to create an image | |
| 16:06:52 | bauzas | then we create an instance | |
| 16:07:06 | bauzas | and then a volume which we attach to the instance | |
| 16:07:44 | gibi | I haven't had time to look into the actual tc yet | |
| 16:07:45 | sean-k-mooney | p/ | |
| 16:08:08 | gibi | also it would be nice to see how the python interpreter rss size grows during the test execution | |
| 16:08:11 | dansmith | yeah surely seems like a benign test case | |
| 16:08:38 | sean-k-mooney | we unfortuently dont have the memtacker stuff form devstack | |
| 16:08:50 | sean-k-mooney | btu it would be nice if we coudl get that and also dmsg in the tox based tests | |
| 16:08:52 | bauzas | I tried to grep the testname in n-api | |
| 16:09:01 | bauzas | but I wasn't finding it | |
| 16:09:07 | bauzas | so, either we no longer use it | |
| 16:09:17 | bauzas | or we were not yet calling the nova-api | |
| 16:09:34 | bauzas | which means we have the kill before creating the instance | |
| 16:09:55 | bauzas | but I could be wrong | |
| 16:10:53 | bauzas | anyway, folks are ok if we modify the bug to High ? | |
| 16:10:59 | bauzas | bug report* | |
| 16:11:32 | gibi | tomorrow I will continue looking but we can also tentatively try to disable this single test to see if that removes the OOM problem | |
| 16:12:18 | gibi | bauzas: I'm not against having this as High | |
| 16:12:26 | bauzas | ok | |
| 16:12:35 | bauzas | then let's look again tomorrow and we'll see what to do | |
| 16:13:08 | bauzas | this time I'm just afraid to remove this test because we don't know why we have a OOMkill | |
| 16:13:18 | bauzas | this could arrive to another test then | |
| 16:13:40 | gibi | yep, that would be my goal of disabling it temporary to see if the OOM just moves to another test case | |
| 16:13:44 | gibi | and to see which test case | |
| 16:13:48 | gibi | to find a pattern | |
| 16:13:48 | bauzas | (I also verified that nothing changed on the tempest side since 1 year for this test) | |
| 16:14:05 | dansmith | gibi: you could also rename it I think and change the sort ordering | |
| 16:14:10 | bauzas | yeah | |
| 16:14:11 | dansmith | afaik, we run tests sorted per worker | |
| 16:14:27 | bauzas | I was wondering, maybe this was a problem due to another test | |
| 16:14:28 | gibi | dansmith: good idea | |
| 16:14:45 | bauzas | dansmith: I think we can ask stestr to modifyh the sort | |
| 16:14:55 | dansmith | oh? | |
| 16:14:59 | bauzas | but I need to remember how to do it | |
| 16:15:05 | gibi | bauzas: on that I extracted all the test cases form the killed worker from multiple runs and the only test case overlap was this tc | |
| 16:15:35 | gibi | so if other test causing the issue then it is not a single test but a set of tests | |
| 16:15:49 | gibi | otherwise I would see an overlap | |
| 16:15:51 | bauzas | gibi: well, yeah, but that maybe means that the previous tests were adding more memory before so that's only with this test that OOMkiller wants to kill | |
| 16:16:14 | bauzas | as you see, this is a very simple test | |
| 16:16:22 | gibi | that is my point above, if a single test adds the extra memory usage then that woudl show up as an overlap between runs | |
| 16:16:30 | gibi | but it doesn't | |
| 16:16:45 | bauzas | gibi: that's why I'll try to see how to ask stestr to modify the sort | |
| 16:17:12 | gibi | yeah, moving this tc to the end can help to see if there is a set of tests that trigger this behavior | |
| 16:18:07 | gibi | anyhow I think we can move on | |
| 16:19:03 | bauzas | cool | |
| 16:19:19 | bauzas | #link https://bugs.launchpad.net/nova/+bugs?search=Search&field.status=New 27 new untriaged bugs (+0 since the last meeting) | |
| 16:19:45 | bauzas | I triaged a few bugs todaty | |
| 16:19:57 | bauzas | #link https://etherpad.opendev.org/p/nova-bug-triage-20230110 | |
| 16:20:16 | bauzas | nothing to report here by now | |
| 16:20:21 | bauzas | #info Add yourself in the team bug roster if you want to help https://etherpad.opendev.org/p/nova-bug-triage-roster | |
| 16:20:28 | bauzas | gibi: wants to get the bug baton this week ? | |
| 16:21:23 | gibi | bauzas: sure I can | |
| 16:21:29 | bauzas | thanks alot | |
| 16:21:47 | bauzas | #info bug baton is being passed to gibi | |
| 16:21:53 | bauzas | #topic Gate status | |
| 16:21:58 | bauzas | #link https://bugs.launchpad.net/nova/+bugs?field.tag=gate-failure Nova gate bugs | |
| 16:22:14 | bauzas | we already discussed about the main one, wanting to discuss other CI bugs ? | |
| 16:22:38 | gibi | just a sort summary | |
| 16:22:38 | bauzas | looks not | |
| 16:22:41 | bauzas | ah | |
| 16:22:50 | bauzas | we're listening to you | |
| 16:22:53 | gibi | I see failures in our functional tests | |
| 16:23:13 | gibi | one is about missing db tables so it is probably interference between test cases | |
| 16:23:21 | gibi | we saw that before | |
| 16:23:30 | gibi | fixed it but not we had a non 100% fix | |
| 16:23:39 | bauzas | :/ | |
| 16:24:08 | gibi | and there is a failure with db cursor need a reset | |
| 16:24:14 | gibi | it might be related to the above | |
| 16:24:17 | gibi | not sure yet | |
| 16:24:47 | bauzas | lovely | |
| 16:24:55 | gibi | these two I wanted to mention | |
| 16:25:08 | gibi | but there are other open bugs that appear in the gate time to time | |
| 16:25:09 | bauzas | flipping strest worker runs would help to trigger the races | |
| 16:25:20 | gibi | so it is fairly hard to land things overall | |
| 16:25:29 | bauzas | I could try to reproduce those functests locally | |
| 16:25:45 | bauzas | this would exhaust my laptop, but worth trying | |
| 16:26:18 | bauzas | gibi: let's then discuss this tomorrow as well | |
| 16:26:23 | gibi | sure | |
| 16:26:48 | bauzas | I mean, I have my power mgmt series to work on, but if we can't land things, nothing will merge either way. | |
| 16:27:03 | sean-k-mooney | the gate is not totally blocked | |
| 16:27:12 | sean-k-mooney | but its flaky enough that its hard | |
| 16:27:15 | bauzas | yeah, but rechecking is not a great option | |
| 16:27:16 | gibi | yepp | |
| 16:27:22 | sean-k-mooney | ya its not | |
| 16:27:29 | bauzas | agreed, I'm not sending the signal our gate is busted | |
| 16:27:35 | bauzas | but we know this is hard | |
| 16:27:43 | sean-k-mooney | one thing i have noticed is the py3.10 functional job seams more stable then py38 | |
| 16:27:47 | bauzas | and let me go to the next topic and you'll understand why | |
| 16:28:01 | sean-k-mooney | for the db issues | |
| 16:28:09 | sean-k-mooney | but that could be just the ones i happend to look at | |
| 16:28:14 | bauzas | ok | |