Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-07
10:04:11 gibi bauzas: ack I will add a logsearch query for it as soon as I have one
10:04:27 bauzas gibi: thanks, I was just doing it but I'm not expert of the tool
10:04:51 bauzas I'm very disappointed we can't longer share our ex-logstack urls
10:04:54 bauzas logstash*
10:05:17 gibi the SQL cursor one https://bugs.launchpad.net/nova/+bug/2002782 is not that frequent, we hit it 9 times in 20 days so I would ignore it for now
10:05:19 bauzas not very handy for finding how much we're doomed
10:05:29 bauzas gibi: yup, I was about to tell you
10:05:47 bauzas gibi: that's why I wanted to add the opensearch one
10:09:18 gibi that 1525 hits in 7 days seem way to much for https://bugs.launchpad.net/nova/+bug/1946339 I suspect some false positives there
10:10:00 bauzas probably
10:11:29 bauzas gibi: fancy to share your logsearch config files ?
10:12:05 gibi https://github.com/gibizer/zuul-log-search-config
10:14:54 gibi so we have sort of false positives. For example this https://zuul.opendev.org/t/openstack/build/31e9de9e4a574df6a2f45546927954fe/log/job-output.txt#23436 partially reproduced https://bugs.launchpad.net/nova/+bug/1946339 but in this case the test case did not fail due the the missing DB table just the stack trace was logged.
10:15:27 gibi so we probably have a lot of hits that has the stack trace but no failed tests
10:16:04 bauzas gibi: I can amend the opensearch query to ensure the outcome is FAILURE
10:16:51 gibi yepp, this is an example of a passed tox run with the missing DB table test case https://zuul.opendev.org/t/openstack/build/f389ba6c4d9f4a04a5b6f09e253d864b/log/job-output.txt#23501
10:17:05 bauzas 191 hits :)
10:17:10 gibi s/missing DB table test case/missing DB table stack trace/
10:17:11 bauzas in the last 7 days
10:18:20 bauzas gibi: sorry for this dumb question but I wanna rush
10:18:34 bauzas gibi: how to use the configs from the other repo into the main one ?
10:18:43 bauzas just all the files ?
10:18:50 bauzas or just checkouting some of them ?
10:19:53 gibi bauzas: what I do is: 1) create a venv and install the tool with pip install git+http://github.com/gibizer/zuul-log-search 2) next to the .venv clone the config repo
10:20:01 bauzas did 1)
10:20:13 bauzas did 2)
10:20:31 gibi It uses .logsearch.conf.d/ in the current directory if exists. Otherwise, uses $XDG_CONFIG_HOME/logsearch/ if XDG_CONFIG_HOME is defined. Otherwise, uses ~/.config/logsearch/.
10:20:44 bauzas yeah so mv the whole dir ?
10:20:51 bauzas that was my question
10:21:20 bauzas I see a config subdir in the zuul-log-search
10:21:24 bauzas but it seems unused
10:21:28 gibi https://paste.opendev.org/show/bTV1aUuajhb3uPYNb8Mp/
10:22:06 gibi sorry, so create the .vevn in the clone config repo.
10:22:54 bauzas I see
10:23:17 bauzas or ln -sf this config dir
10:23:36 bauzas which is what I'll be using
10:24:09 gibi ack
10:24:58 bauzas yay, that works
10:28:40 gibi I will go and collect other frequent gate failures based on the query logsearch build --project openstack/nova --voting --pipeline gate --result FAILURE --branch master --days 7
10:29:45 bauzas gibi: iiuc, Builds with matching logs 160/162 means that over 162 job runs with FAILURE, 160 of them were having the query I asked ?
10:29:52 bauzas so, 98% of them
10:30:05 gibi we have some gate runs which are TIMED_OUT too logsearch build --project openstack/nova --voting --pipeline gate --result TIMED_OUT --branch master --days 7 I think this is what dansmith mentioned yesterday
10:30:14 gibi bauzas: yes
10:30:31 bauzas gibi: maybe the query I make is too large, as you mentioned
10:30:46 bauzas request was 'sqlite3.OperationalError: no such table: instance_faults"
10:32:07 gibi bauzas: that will pick up the cases when you see the stack trace without that killing the test but job failed for other reason
10:32:11 gibi but we need to live with it
10:32:30 bauzas yup
10:32:51 bauzas I think we now have enough to work with
10:33:04 bauzas I'll try to do this digging thing
10:33:38 gibi ack
10:34:24 opendevreview Maxim Monin proposed openstack/nova master: Server Rescue leads to Server ERROR state if base image is deleted https://review.opendev.org/c/openstack/nova/+/872385
10:34:31 bauzas I tried to ask opensearch to give me the occurrences down to 30 days
10:34:47 bauzas I'll try to see whether it started to reappear at some point in time
10:35:29 bauzas mmm, interesting
10:35:38 bauzas grabbing occurrences for the last 2 months
10:36:38 bauzas gibi: https://imgur.com/a/wOeLcUf
10:37:02 gibi bauzas: we have limited log storage
10:37:03 bauzas it started recently
10:37:17 bauzas less than one month of storage ?
10:37:24 bauzas or more ?
10:37:30 gibi with the old logstash it was about a month
10:37:39 gibi I don't know about the new one
10:37:52 bauzas gibi: with old logstash, I was sure it was a month
10:38:21 bauzas anyway, if so, let's start to find the regression by other way
10:38:46 bauzas and I suspect this can't be reproduced locally
10:39:04 bauzas or I would need to speed down my laptop
10:50:08 dvo-plv Hello, <sean-k-mooney>
10:50:28 dvo-plv I would like to continue our coversation, which we had at friday
10:50:39 dvo-plv We talked about packed_ring option
10:51:04 dvo-plv I would like to discuss schedulet.
10:51:04 gibi bauzas: yeah one thing you can try is to slow down things and increase the frequency of the test case that was failed by duplicating in many times
10:51:36 dvo-plv The situation when user did not ask about COMPUTE_NET_VIRTIO_PACKED trait, but we need to handle migration in some way. I found that scheduler has ALL_REQUEST_FILTERS array with different filters. My eye falls on the accelerators_filter. I suggest implement packed_ring filtering in the same way as in this method. Also this give us ability to avoid situation when user want to start VM on the node where this feature is not unavailable
10:53:01 sean-k-mooney dvo-plv: good thinking but that would be the legacy approch
10:54:24 sean-k-mooney dvo-plv: my counter propsal is this. when a vm is spwaned on a host if it support COMPUTE_NET_VIRTIO_PACKED set a flag in the instance_system_metadta to record that. then instead of adding a post placement filter add a pre placement filter here https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py
10:55:17 sean-k-mooney dvo-plv: unless this is the accelerators filter you ment https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py#L260-L273
10:55:29 sean-k-mooney if so then yes it would be very similar to that
10:55:50 sean-k-mooney we would either check for a extra_spec and add the trait in an identical way
10:56:07 sean-k-mooney or check the instnace_system_metadta for the flag.
10:56:41 sean-k-mooney the former would take effect when booting a vm that explictly request this the latter for any vm that was spwaned on a host with this capablity
10:58:05 sean-k-mooney one of the main probalem with the instnace_system_metadata approch is im not sure the request_spec has that field
10:58:43 sean-k-mooney the approch we take really comes down to one choice. does the packed ring format need to be opt-in or automatic
10:59:02 sean-k-mooney if its opt in via a flavor/image property then the prefilter is trivial
10:59:56 sean-k-mooney the request spec has both the image properties and flavor extra spec so you can just check them and add the required trait as the accleror filter does
11:01:07 sean-k-mooney looking at https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py the instance_system_metadata is not part of the request spec currently
11:01:43 sean-k-mooney so to that ahat approch we woul d have ot modify the request spec which im not sure is the right thing to do here.
11:02:19 sean-k-mooney the request spec and instance_system_metadata live in different DBs (api vs cell db)
11:03:54 bauzas gibi: I chose the stestr approach of --until-failure
11:04:01 dvo-plv I check instance_system_metadata table and for me it looks like we will mix different OpenStack's layers ( instance and host) , because there is no type of data for instance like that, we have there some image, project, and user info. And I did not found how this table link with request_spec what we create at the scheduler
11:04:04 sean-k-mooney dvo-plv: so based on that i would suggest we take the opt in/out approch and use flavor/image properties
11:05:05 gibi bauzas: that is independent from increasing the chance of catching it by increasing the number of test case to run that we know can fail due to the issue. you can do both
11:05:10 sean-k-mooney dvo-plv: instance_system_metadata is a generic key value store for storing internal information about the instnace. such as the embeded image properites
11:05:25 sean-k-mooney and its not accesable to the schduler genreally
11:06:04 bauzas gibi: we know that the issue is not on a single test
11:06:31 bauzas so while the testrunner runs, I'm looking at every single failure to see the stacktrace and find a pattern
11:06:41 gibi bauzas: yes, but we can grab a list of test cases run by a failed test worker. That list contains both the test case that leaked and the test case that failed due to the leak
11:06:48 gibi bauzas: we know what is the latter
11:06:48 opendevreview Merged openstack/nova master: Fix 6.2 compute RPC version alias https://review.opendev.org/c/openstack/nova/+/872804
11:07:11 sean-k-mooney dvo-plv: what you would actully need to do is have the conductor populate a filed on the request spec when you do a live migration. i feel like that approch is more complex then requried
11:07:13 gibi bauzas: so we can run the same testcase list

Earlier   Later