| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-07 | |||
| 10:20:31 | gibi | It uses .logsearch.conf.d/ in the current directory if exists. Otherwise, uses $XDG_CONFIG_HOME/logsearch/ if XDG_CONFIG_HOME is defined. Otherwise, uses ~/.config/logsearch/. | |
| 10:20:44 | bauzas | yeah so mv the whole dir ? | |
| 10:20:51 | bauzas | that was my question | |
| 10:21:20 | bauzas | I see a config subdir in the zuul-log-search | |
| 10:21:24 | bauzas | but it seems unused | |
| 10:21:28 | gibi | https://paste.opendev.org/show/bTV1aUuajhb3uPYNb8Mp/ | |
| 10:22:06 | gibi | sorry, so create the .vevn in the clone config repo. | |
| 10:22:54 | bauzas | I see | |
| 10:23:17 | bauzas | or ln -sf this config dir | |
| 10:23:36 | bauzas | which is what I'll be using | |
| 10:24:09 | gibi | ack | |
| 10:24:58 | bauzas | yay, that works | |
| 10:28:40 | gibi | I will go and collect other frequent gate failures based on the query logsearch build --project openstack/nova --voting --pipeline gate --result FAILURE --branch master --days 7 | |
| 10:29:45 | bauzas | gibi: iiuc, Builds with matching logs 160/162 means that over 162 job runs with FAILURE, 160 of them were having the query I asked ? | |
| 10:29:52 | bauzas | so, 98% of them | |
| 10:30:05 | gibi | we have some gate runs which are TIMED_OUT too logsearch build --project openstack/nova --voting --pipeline gate --result TIMED_OUT --branch master --days 7 I think this is what dansmith mentioned yesterday | |
| 10:30:14 | gibi | bauzas: yes | |
| 10:30:31 | bauzas | gibi: maybe the query I make is too large, as you mentioned | |
| 10:30:46 | bauzas | request was 'sqlite3.OperationalError: no such table: instance_faults" | |
| 10:32:07 | gibi | bauzas: that will pick up the cases when you see the stack trace without that killing the test but job failed for other reason | |
| 10:32:11 | gibi | but we need to live with it | |
| 10:32:30 | bauzas | yup | |
| 10:32:51 | bauzas | I think we now have enough to work with | |
| 10:33:04 | bauzas | I'll try to do this digging thing | |
| 10:33:38 | gibi | ack | |
| 10:34:24 | opendevreview | Maxim Monin proposed openstack/nova master: Server Rescue leads to Server ERROR state if base image is deleted https://review.opendev.org/c/openstack/nova/+/872385 | |
| 10:34:31 | bauzas | I tried to ask opensearch to give me the occurrences down to 30 days | |
| 10:34:47 | bauzas | I'll try to see whether it started to reappear at some point in time | |
| 10:35:29 | bauzas | mmm, interesting | |
| 10:35:38 | bauzas | grabbing occurrences for the last 2 months | |
| 10:36:38 | bauzas | gibi: https://imgur.com/a/wOeLcUf | |
| 10:37:02 | gibi | bauzas: we have limited log storage | |
| 10:37:03 | bauzas | it started recently | |
| 10:37:17 | bauzas | less than one month of storage ? | |
| 10:37:24 | bauzas | or more ? | |
| 10:37:30 | gibi | with the old logstash it was about a month | |
| 10:37:39 | gibi | I don't know about the new one | |
| 10:37:52 | bauzas | gibi: with old logstash, I was sure it was a month | |
| 10:38:21 | bauzas | anyway, if so, let's start to find the regression by other way | |
| 10:38:46 | bauzas | and I suspect this can't be reproduced locally | |
| 10:39:04 | bauzas | or I would need to speed down my laptop | |
| 10:50:08 | dvo-plv | Hello, <sean-k-mooney> | |
| 10:50:28 | dvo-plv | I would like to continue our coversation, which we had at friday | |
| 10:50:39 | dvo-plv | We talked about packed_ring option | |
| 10:51:04 | gibi | bauzas: yeah one thing you can try is to slow down things and increase the frequency of the test case that was failed by duplicating in many times | |
| 10:51:04 | dvo-plv | I would like to discuss schedulet. | |
| 10:51:36 | dvo-plv | The situation when user did not ask about COMPUTE_NET_VIRTIO_PACKED trait, but we need to handle migration in some way. I found that scheduler has ALL_REQUEST_FILTERS array with different filters. My eye falls on the accelerators_filter. I suggest implement packed_ring filtering in the same way as in this method. Also this give us ability to avoid situation when user want to start VM on the node where this feature is not unavailable | |
| 10:53:01 | sean-k-mooney | dvo-plv: good thinking but that would be the legacy approch | |
| 10:54:24 | sean-k-mooney | dvo-plv: my counter propsal is this. when a vm is spwaned on a host if it support COMPUTE_NET_VIRTIO_PACKED set a flag in the instance_system_metadta to record that. then instead of adding a post placement filter add a pre placement filter here https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py | |
| 10:55:17 | sean-k-mooney | dvo-plv: unless this is the accelerators filter you ment https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py#L260-L273 | |
| 10:55:29 | sean-k-mooney | if so then yes it would be very similar to that | |
| 10:55:50 | sean-k-mooney | we would either check for a extra_spec and add the trait in an identical way | |
| 10:56:07 | sean-k-mooney | or check the instnace_system_metadta for the flag. | |
| 10:56:41 | sean-k-mooney | the former would take effect when booting a vm that explictly request this the latter for any vm that was spwaned on a host with this capablity | |
| 10:58:05 | sean-k-mooney | one of the main probalem with the instnace_system_metadata approch is im not sure the request_spec has that field | |
| 10:58:43 | sean-k-mooney | the approch we take really comes down to one choice. does the packed ring format need to be opt-in or automatic | |
| 10:59:02 | sean-k-mooney | if its opt in via a flavor/image property then the prefilter is trivial | |
| 10:59:56 | sean-k-mooney | the request spec has both the image properties and flavor extra spec so you can just check them and add the required trait as the accleror filter does | |
| 11:01:07 | sean-k-mooney | looking at https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py the instance_system_metadata is not part of the request spec currently | |
| 11:01:43 | sean-k-mooney | so to that ahat approch we woul d have ot modify the request spec which im not sure is the right thing to do here. | |
| 11:02:19 | sean-k-mooney | the request spec and instance_system_metadata live in different DBs (api vs cell db) | |
| 11:03:54 | bauzas | gibi: I chose the stestr approach of --until-failure | |
| 11:04:01 | dvo-plv | I check instance_system_metadata table and for me it looks like we will mix different OpenStack's layers ( instance and host) , because there is no type of data for instance like that, we have there some image, project, and user info. And I did not found how this table link with request_spec what we create at the scheduler | |
| 11:04:04 | sean-k-mooney | dvo-plv: so based on that i would suggest we take the opt in/out approch and use flavor/image properties | |
| 11:05:05 | gibi | bauzas: that is independent from increasing the chance of catching it by increasing the number of test case to run that we know can fail due to the issue. you can do both | |
| 11:05:10 | sean-k-mooney | dvo-plv: instance_system_metadata is a generic key value store for storing internal information about the instnace. such as the embeded image properites | |
| 11:05:25 | sean-k-mooney | and its not accesable to the schduler genreally | |
| 11:06:04 | bauzas | gibi: we know that the issue is not on a single test | |
| 11:06:31 | bauzas | so while the testrunner runs, I'm looking at every single failure to see the stacktrace and find a pattern | |
| 11:06:41 | gibi | bauzas: yes, but we can grab a list of test cases run by a failed test worker. That list contains both the test case that leaked and the test case that failed due to the leak | |
| 11:06:48 | opendevreview | Merged openstack/nova master: Fix 6.2 compute RPC version alias https://review.opendev.org/c/openstack/nova/+/872804 | |
| 11:06:48 | gibi | bauzas: we know what is the latter | |
| 11:07:11 | sean-k-mooney | dvo-plv: what you would actully need to do is have the conductor populate a filed on the request spec when you do a live migration. i feel like that approch is more complex then requried | |
| 11:07:13 | gibi | bauzas: so we can run the same testcase list | |
| 11:07:22 | gibi | bauzas: as we know it contains both | |
| 11:07:36 | bauzas | gibi: I see your proposal | |
| 11:07:41 | gibi | bauzas: then we can increase the chance by adding more test cases that is in the latter category | |
| 11:07:47 | bauzas | I have the subunits | |
| 11:07:54 | bauzas | so I can generate a list | |
| 11:07:58 | bauzas | of failing tests | |
| 11:08:03 | bauzas | and duplicate that list | |
| 11:09:39 | sean-k-mooney | dvo-plv: if we were to leverage the instance_system_metadta we would likely need to extend the Destination filed to have addtional trait requests or something like that https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L1093 | |
| 11:09:41 | gibi | originally (in 2021) this way I was able to reproduce https://bugs.launchpad.net/nova/+bug/1946339 but I tried this couple weeks ago again and was not able to reproduce the current occasion after couple hour of --unit-failure run | |
| 11:10:47 | sean-k-mooney | dvo-plv: the destination object is constucted here https://github.com/openstack/nova/blob/1c46c4e9e5ba4b84816f5cadad0674f3a773e739/nova/conductor/tasks/live_migrate.py#L64 | |
| 11:11:19 | sean-k-mooney | dvo-plv: but as i said this more complex approch is only relevent if we wanted to automatically enabel this functionality | |
| 11:12:31 | sean-k-mooney | well technially its created here https://github.com/openstack/nova/blob/1c46c4e9e5ba4b84816f5cadad0674f3a773e739/nova/conductor/manager.py#L470 | |
| 11:13:36 | sean-k-mooney | this only matters for the live migration case as the feature can be renegociated on cold migration or other move operations | |
| 11:14:18 | bauzas | (functional-py38) [sbauza@sbauza nova]$ stestr load /tmp/zuul-logs.Edb6Rp/testrepository.subunit --subunit | subunit-filter -F | subunit-ls | |
| 11:14:18 | bauzas | nova.tests.functional.libvirt.test_vtpm.VTPMServersTest.test_create_server | |
| 11:14:30 | bauzas | gibi: I'm able to get the failing test | |
| 11:14:57 | bauzas | so given I'm looking at all the fetched logsearch subunits, I could extrapolate a list of usual suspects | |
| 11:15:38 | dvo-plv | So if you think that this way is very complex and can make code not so easy and familiar, maybe we should better use existing approach ( creates a new filter like accelerators_filter and check if the user requested packed option) what already exists and is easy to scale. I already check this approach, this approach also forbids migrating VM to the host without packed ring support and also start VM on the host without packed ring support | |
| 11:16:51 | sean-k-mooney | dvo-plv: yep that is the simpelest approch. we could in a future release enable it by default and add a migration mechaniums too if desired. | |
| 11:17:19 | sean-k-mooney | either by turning it on once we raise our min QEMU/Libvirt versions to one that means it will alwasy be aviable | |
| 11:17:39 | sean-k-mooney | or buy automatically adding the image property if not provided | |
| 11:18:09 | sean-k-mooney | so taking the explict approch now does not prevent use making it implict in the future | |
| 11:18:35 | sean-k-mooney | making it automatic now front loads a bunch of complexity | |
| 11:21:54 | gibi | bauzas: I tried that too. I fetched multiple failed worker test case list and intersected them it resulted in an empty list. probably we have multiple test cases that leaks | |
| 11:22:20 | gibi | bauzas: but you can be lucky | |
| 11:23:54 | bauzas | gibi: I have the uuids from logseearch but I don't have the subunit streams | |