| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-07 | |||
| 09:52:58 | bauzas | gibi: I'm not good at naming | |
| 09:53:10 | bauzas | and I'm not sure of the rootcause | |
| 09:53:26 | bauzas | I guess this is due to long-running threads | |
| 09:53:40 | bauzas | so, due to the length it takes, then we delete the DB | |
| 09:54:06 | gibi | bauzas: changed the title then | |
| 09:54:08 | bauzas | and so, when the thread stops, then we try to lookup the table which no longer exists | |
| 09:54:22 | bauzas | amirite on the root cause ? | |
| 09:54:52 | gibi | bauzas: one of the root cause is described in the commit message here https://review.opendev.org/c/openstack/nova/+/814036 | |
| 09:56:24 | gibi | what I we need to figure out what is the new sequence of events that leads to a similar leak | |
| 09:56:25 | bauzas | yeah ok, so we agree | |
| 09:56:35 | bauzas | we have spawned greenlets | |
| 09:56:42 | bauzas | that finish after the db is cleaned up | |
| 09:56:59 | bauzas | we should hold the test until those greenlets finish | |
| 09:58:16 | gibi | 1) we need to figure out which eventlet leaks 2) then we need to figure out how that leaked eventlet affects the later test case, i.e. where it the global variable that supports the leak 3) then we can figure out how to fix it | |
| 09:59:12 | tobias-urdin | bauzas: no worries, funny coincidence that i was digging around in the RPC layer and saw it :) | |
| 09:59:53 | gibi | in the fixed case it was RPC handling eventlet was leaked that after 60 sec waked up do to timeout and used nova.rpc.get_versioned_notifier() to get a fresh notifier to the actually running test case | |
| 10:00:21 | opendevreview | Jorge San Emeterio proposed openstack/nova master: Dividing global privsep profile https://review.opendev.org/c/openstack/nova/+/871729 | |
| 10:00:51 | gibi | the fix there was to ignore notifications if it is comming from a test case that is different from the currently running one | |
| 10:03:53 | bauzas | gibi: grabbing a coffee, this is a hard issue to dig into | |
| 10:04:11 | gibi | bauzas: ack I will add a logsearch query for it as soon as I have one | |
| 10:04:27 | bauzas | gibi: thanks, I was just doing it but I'm not expert of the tool | |
| 10:04:51 | bauzas | I'm very disappointed we can't longer share our ex-logstack urls | |
| 10:04:54 | bauzas | logstash* | |
| 10:05:17 | gibi | the SQL cursor one https://bugs.launchpad.net/nova/+bug/2002782 is not that frequent, we hit it 9 times in 20 days so I would ignore it for now | |
| 10:05:19 | bauzas | not very handy for finding how much we're doomed | |
| 10:05:29 | bauzas | gibi: yup, I was about to tell you | |
| 10:05:47 | bauzas | gibi: that's why I wanted to add the opensearch one | |
| 10:09:18 | gibi | that 1525 hits in 7 days seem way to much for https://bugs.launchpad.net/nova/+bug/1946339 I suspect some false positives there | |
| 10:10:00 | bauzas | probably | |
| 10:11:29 | bauzas | gibi: fancy to share your logsearch config files ? | |
| 10:12:05 | gibi | https://github.com/gibizer/zuul-log-search-config | |
| 10:14:54 | gibi | so we have sort of false positives. For example this https://zuul.opendev.org/t/openstack/build/31e9de9e4a574df6a2f45546927954fe/log/job-output.txt#23436 partially reproduced https://bugs.launchpad.net/nova/+bug/1946339 but in this case the test case did not fail due the the missing DB table just the stack trace was logged. | |
| 10:15:27 | gibi | so we probably have a lot of hits that has the stack trace but no failed tests | |
| 10:16:04 | bauzas | gibi: I can amend the opensearch query to ensure the outcome is FAILURE | |
| 10:16:51 | gibi | yepp, this is an example of a passed tox run with the missing DB table test case https://zuul.opendev.org/t/openstack/build/f389ba6c4d9f4a04a5b6f09e253d864b/log/job-output.txt#23501 | |
| 10:17:05 | bauzas | 191 hits :) | |
| 10:17:10 | gibi | s/missing DB table test case/missing DB table stack trace/ | |
| 10:17:11 | bauzas | in the last 7 days | |
| 10:18:20 | bauzas | gibi: sorry for this dumb question but I wanna rush | |
| 10:18:34 | bauzas | gibi: how to use the configs from the other repo into the main one ? | |
| 10:18:43 | bauzas | just all the files ? | |
| 10:18:50 | bauzas | or just checkouting some of them ? | |
| 10:19:53 | gibi | bauzas: what I do is: 1) create a venv and install the tool with pip install git+http://github.com/gibizer/zuul-log-search 2) next to the .venv clone the config repo | |
| 10:20:01 | bauzas | did 1) | |
| 10:20:13 | bauzas | did 2) | |
| 10:20:31 | gibi | It uses .logsearch.conf.d/ in the current directory if exists. Otherwise, uses $XDG_CONFIG_HOME/logsearch/ if XDG_CONFIG_HOME is defined. Otherwise, uses ~/.config/logsearch/. | |
| 10:20:44 | bauzas | yeah so mv the whole dir ? | |
| 10:20:51 | bauzas | that was my question | |
| 10:21:20 | bauzas | I see a config subdir in the zuul-log-search | |
| 10:21:24 | bauzas | but it seems unused | |
| 10:21:28 | gibi | https://paste.opendev.org/show/bTV1aUuajhb3uPYNb8Mp/ | |
| 10:22:06 | gibi | sorry, so create the .vevn in the clone config repo. | |
| 10:22:54 | bauzas | I see | |
| 10:23:17 | bauzas | or ln -sf this config dir | |
| 10:23:36 | bauzas | which is what I'll be using | |
| 10:24:09 | gibi | ack | |
| 10:24:58 | bauzas | yay, that works | |
| 10:28:40 | gibi | I will go and collect other frequent gate failures based on the query logsearch build --project openstack/nova --voting --pipeline gate --result FAILURE --branch master --days 7 | |
| 10:29:45 | bauzas | gibi: iiuc, Builds with matching logs 160/162 means that over 162 job runs with FAILURE, 160 of them were having the query I asked ? | |
| 10:29:52 | bauzas | so, 98% of them | |
| 10:30:05 | gibi | we have some gate runs which are TIMED_OUT too logsearch build --project openstack/nova --voting --pipeline gate --result TIMED_OUT --branch master --days 7 I think this is what dansmith mentioned yesterday | |
| 10:30:14 | gibi | bauzas: yes | |
| 10:30:31 | bauzas | gibi: maybe the query I make is too large, as you mentioned | |
| 10:30:46 | bauzas | request was 'sqlite3.OperationalError: no such table: instance_faults" | |
| 10:32:07 | gibi | bauzas: that will pick up the cases when you see the stack trace without that killing the test but job failed for other reason | |
| 10:32:11 | gibi | but we need to live with it | |
| 10:32:30 | bauzas | yup | |
| 10:32:51 | bauzas | I think we now have enough to work with | |
| 10:33:04 | bauzas | I'll try to do this digging thing | |
| 10:33:38 | gibi | ack | |
| 10:34:24 | opendevreview | Maxim Monin proposed openstack/nova master: Server Rescue leads to Server ERROR state if base image is deleted https://review.opendev.org/c/openstack/nova/+/872385 | |
| 10:34:31 | bauzas | I tried to ask opensearch to give me the occurrences down to 30 days | |
| 10:34:47 | bauzas | I'll try to see whether it started to reappear at some point in time | |
| 10:35:29 | bauzas | mmm, interesting | |
| 10:35:38 | bauzas | grabbing occurrences for the last 2 months | |
| 10:36:38 | bauzas | gibi: https://imgur.com/a/wOeLcUf | |
| 10:37:02 | gibi | bauzas: we have limited log storage | |
| 10:37:03 | bauzas | it started recently | |
| 10:37:17 | bauzas | less than one month of storage ? | |
| 10:37:24 | bauzas | or more ? | |
| 10:37:30 | gibi | with the old logstash it was about a month | |
| 10:37:39 | gibi | I don't know about the new one | |
| 10:37:52 | bauzas | gibi: with old logstash, I was sure it was a month | |
| 10:38:21 | bauzas | anyway, if so, let's start to find the regression by other way | |
| 10:38:46 | bauzas | and I suspect this can't be reproduced locally | |
| 10:39:04 | bauzas | or I would need to speed down my laptop | |
| 10:50:08 | dvo-plv | Hello, <sean-k-mooney> | |
| 10:50:28 | dvo-plv | I would like to continue our coversation, which we had at friday | |
| 10:50:39 | dvo-plv | We talked about packed_ring option | |
| 10:51:04 | gibi | bauzas: yeah one thing you can try is to slow down things and increase the frequency of the test case that was failed by duplicating in many times | |
| 10:51:04 | dvo-plv | I would like to discuss schedulet. | |
| 10:51:36 | dvo-plv | The situation when user did not ask about COMPUTE_NET_VIRTIO_PACKED trait, but we need to handle migration in some way. I found that scheduler has ALL_REQUEST_FILTERS array with different filters. My eye falls on the accelerators_filter. I suggest implement packed_ring filtering in the same way as in this method. Also this give us ability to avoid situation when user want to start VM on the node where this feature is not unavailable | |
| 10:53:01 | sean-k-mooney | dvo-plv: good thinking but that would be the legacy approch | |
| 10:54:24 | sean-k-mooney | dvo-plv: my counter propsal is this. when a vm is spwaned on a host if it support COMPUTE_NET_VIRTIO_PACKED set a flag in the instance_system_metadta to record that. then instead of adding a post placement filter add a pre placement filter here https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py | |
| 10:55:17 | sean-k-mooney | dvo-plv: unless this is the accelerators filter you ment https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py#L260-L273 | |
| 10:55:29 | sean-k-mooney | if so then yes it would be very similar to that | |
| 10:55:50 | sean-k-mooney | we would either check for a extra_spec and add the trait in an identical way | |
| 10:56:07 | sean-k-mooney | or check the instnace_system_metadta for the flag. | |
| 10:56:41 | sean-k-mooney | the former would take effect when booting a vm that explictly request this the latter for any vm that was spwaned on a host with this capablity | |
| 10:58:05 | sean-k-mooney | one of the main probalem with the instnace_system_metadata approch is im not sure the request_spec has that field | |