| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-07 | |||
| 08:13:25 | opendevreview | melanie witt proposed openstack/nova master: libvirt: Introduce support for qcow2 with LUKS https://review.opendev.org/c/openstack/nova/+/772273 | |
| 09:04:17 | bauzas | tobias-urdin: excellent catch on the RPC pin alias, many thanks | |
| 09:04:31 | bauzas | it would have broken seriously our users if we were having it released | |
| 09:27:36 | bauzas | yet another strange DB issue https://storage.gra.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_d00/868237/9/check/nova-tox-functional-py38/d00d1ff/testr_results.html | |
| 09:27:50 | bauzas | gibi: I'm writing the etherpad | |
| 09:28:03 | bauzas | and I'm reporting a new gate failure | |
| 09:29:15 | gibi | bauzas: the failure you linked is probably https://bugs.launchpad.net/nova/+bug/1946339 | |
| 09:29:51 | gibi | so no need for a new bug report | |
| 09:30:45 | bauzas | gibi: seems unrelated | |
| 09:30:56 | bauzas | I get a DB issue due to a missing table | |
| 09:31:09 | bauzas | oh | |
| 09:31:29 | gibi | sqlite3.OperationalError: no such table: instance_faults | |
| 09:31:29 | gibi | both fails with | |
| 09:31:41 | gibi | after a message timeout | |
| 09:31:51 | bauzas | yeah, not the same table but quite the same behaviour | |
| 09:32:03 | bauzas | well, the bug report title is too specific | |
| 09:33:06 | bauzas | gibi: agreed, seems the same | |
| 09:47:20 | bauzas | gibi: https://etherpad.opendev.org/p/nova-ci-failures | |
| 09:49:49 | bauzas | damn shit, opensearch urls aren't shareable | |
| 09:52:40 | gibi | bauzas: feel free to update the title of https://bugs.launchpad.net/nova/+bug/1946339 | |
| 09:52:58 | bauzas | gibi: I'm not good at naming | |
| 09:53:10 | bauzas | and I'm not sure of the rootcause | |
| 09:53:26 | bauzas | I guess this is due to long-running threads | |
| 09:53:40 | bauzas | so, due to the length it takes, then we delete the DB | |
| 09:54:06 | gibi | bauzas: changed the title then | |
| 09:54:08 | bauzas | and so, when the thread stops, then we try to lookup the table which no longer exists | |
| 09:54:22 | bauzas | amirite on the root cause ? | |
| 09:54:52 | gibi | bauzas: one of the root cause is described in the commit message here https://review.opendev.org/c/openstack/nova/+/814036 | |
| 09:56:24 | gibi | what I we need to figure out what is the new sequence of events that leads to a similar leak | |
| 09:56:25 | bauzas | yeah ok, so we agree | |
| 09:56:35 | bauzas | we have spawned greenlets | |
| 09:56:42 | bauzas | that finish after the db is cleaned up | |
| 09:56:59 | bauzas | we should hold the test until those greenlets finish | |
| 09:58:16 | gibi | 1) we need to figure out which eventlet leaks 2) then we need to figure out how that leaked eventlet affects the later test case, i.e. where it the global variable that supports the leak 3) then we can figure out how to fix it | |
| 09:59:12 | tobias-urdin | bauzas: no worries, funny coincidence that i was digging around in the RPC layer and saw it :) | |
| 09:59:53 | gibi | in the fixed case it was RPC handling eventlet was leaked that after 60 sec waked up do to timeout and used nova.rpc.get_versioned_notifier() to get a fresh notifier to the actually running test case | |
| 10:00:21 | opendevreview | Jorge San Emeterio proposed openstack/nova master: Dividing global privsep profile https://review.opendev.org/c/openstack/nova/+/871729 | |
| 10:00:51 | gibi | the fix there was to ignore notifications if it is comming from a test case that is different from the currently running one | |
| 10:03:53 | bauzas | gibi: grabbing a coffee, this is a hard issue to dig into | |
| 10:04:11 | gibi | bauzas: ack I will add a logsearch query for it as soon as I have one | |
| 10:04:27 | bauzas | gibi: thanks, I was just doing it but I'm not expert of the tool | |
| 10:04:51 | bauzas | I'm very disappointed we can't longer share our ex-logstack urls | |
| 10:04:54 | bauzas | logstash* | |
| 10:05:17 | gibi | the SQL cursor one https://bugs.launchpad.net/nova/+bug/2002782 is not that frequent, we hit it 9 times in 20 days so I would ignore it for now | |
| 10:05:19 | bauzas | not very handy for finding how much we're doomed | |
| 10:05:29 | bauzas | gibi: yup, I was about to tell you | |
| 10:05:47 | bauzas | gibi: that's why I wanted to add the opensearch one | |
| 10:09:18 | gibi | that 1525 hits in 7 days seem way to much for https://bugs.launchpad.net/nova/+bug/1946339 I suspect some false positives there | |
| 10:10:00 | bauzas | probably | |
| 10:11:29 | bauzas | gibi: fancy to share your logsearch config files ? | |
| 10:12:05 | gibi | https://github.com/gibizer/zuul-log-search-config | |
| 10:14:54 | gibi | so we have sort of false positives. For example this https://zuul.opendev.org/t/openstack/build/31e9de9e4a574df6a2f45546927954fe/log/job-output.txt#23436 partially reproduced https://bugs.launchpad.net/nova/+bug/1946339 but in this case the test case did not fail due the the missing DB table just the stack trace was logged. | |
| 10:15:27 | gibi | so we probably have a lot of hits that has the stack trace but no failed tests | |
| 10:16:04 | bauzas | gibi: I can amend the opensearch query to ensure the outcome is FAILURE | |
| 10:16:51 | gibi | yepp, this is an example of a passed tox run with the missing DB table test case https://zuul.opendev.org/t/openstack/build/f389ba6c4d9f4a04a5b6f09e253d864b/log/job-output.txt#23501 | |
| 10:17:05 | bauzas | 191 hits :) | |
| 10:17:10 | gibi | s/missing DB table test case/missing DB table stack trace/ | |
| 10:17:11 | bauzas | in the last 7 days | |
| 10:18:20 | bauzas | gibi: sorry for this dumb question but I wanna rush | |
| 10:18:34 | bauzas | gibi: how to use the configs from the other repo into the main one ? | |
| 10:18:43 | bauzas | just all the files ? | |
| 10:18:50 | bauzas | or just checkouting some of them ? | |
| 10:19:53 | gibi | bauzas: what I do is: 1) create a venv and install the tool with pip install git+http://github.com/gibizer/zuul-log-search 2) next to the .venv clone the config repo | |
| 10:20:01 | bauzas | did 1) | |
| 10:20:13 | bauzas | did 2) | |
| 10:20:31 | gibi | It uses .logsearch.conf.d/ in the current directory if exists. Otherwise, uses $XDG_CONFIG_HOME/logsearch/ if XDG_CONFIG_HOME is defined. Otherwise, uses ~/.config/logsearch/. | |
| 10:20:44 | bauzas | yeah so mv the whole dir ? | |
| 10:20:51 | bauzas | that was my question | |
| 10:21:20 | bauzas | I see a config subdir in the zuul-log-search | |
| 10:21:24 | bauzas | but it seems unused | |
| 10:21:28 | gibi | https://paste.opendev.org/show/bTV1aUuajhb3uPYNb8Mp/ | |
| 10:22:06 | gibi | sorry, so create the .vevn in the clone config repo. | |
| 10:22:54 | bauzas | I see | |
| 10:23:17 | bauzas | or ln -sf this config dir | |
| 10:23:36 | bauzas | which is what I'll be using | |
| 10:24:09 | gibi | ack | |
| 10:24:58 | bauzas | yay, that works | |
| 10:28:40 | gibi | I will go and collect other frequent gate failures based on the query logsearch build --project openstack/nova --voting --pipeline gate --result FAILURE --branch master --days 7 | |
| 10:29:45 | bauzas | gibi: iiuc, Builds with matching logs 160/162 means that over 162 job runs with FAILURE, 160 of them were having the query I asked ? | |
| 10:29:52 | bauzas | so, 98% of them | |
| 10:30:05 | gibi | we have some gate runs which are TIMED_OUT too logsearch build --project openstack/nova --voting --pipeline gate --result TIMED_OUT --branch master --days 7 I think this is what dansmith mentioned yesterday | |
| 10:30:14 | gibi | bauzas: yes | |
| 10:30:31 | bauzas | gibi: maybe the query I make is too large, as you mentioned | |
| 10:30:46 | bauzas | request was 'sqlite3.OperationalError: no such table: instance_faults" | |
| 10:32:07 | gibi | bauzas: that will pick up the cases when you see the stack trace without that killing the test but job failed for other reason | |
| 10:32:11 | gibi | but we need to live with it | |
| 10:32:30 | bauzas | yup | |
| 10:32:51 | bauzas | I think we now have enough to work with | |
| 10:33:04 | bauzas | I'll try to do this digging thing | |
| 10:33:38 | gibi | ack | |
| 10:34:24 | opendevreview | Maxim Monin proposed openstack/nova master: Server Rescue leads to Server ERROR state if base image is deleted https://review.opendev.org/c/openstack/nova/+/872385 | |
| 10:34:31 | bauzas | I tried to ask opensearch to give me the occurrences down to 30 days | |
| 10:34:47 | bauzas | I'll try to see whether it started to reappear at some point in time | |
| 10:35:29 | bauzas | mmm, interesting | |
| 10:35:38 | bauzas | grabbing occurrences for the last 2 months | |
| 10:36:38 | bauzas | gibi: https://imgur.com/a/wOeLcUf | |
| 10:37:02 | gibi | bauzas: we have limited log storage | |
| 10:37:03 | bauzas | it started recently | |
| 10:37:17 | bauzas | less than one month of storage ? | |
| 10:37:24 | bauzas | or more ? | |