| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-16 | |||
| 17:39:04 | bauzas | damn, who knows the Launchpad nick of Kirill ? /me needs to paperwork the right ownership of https://blueprints.launchpad.net/nova/+spec/ironic-vnc-console | |
| 17:40:55 | bauzas | anyway, I can live with that | |
| 17:44:24 | bauzas | wow, the numbers of accepted blueprints for Antelope are identical to Yoga | |
| 17:44:37 | bauzas | disclaimer: this is gonna be a productive 5-week | |
| 17:53:13 | sean-k-mooney | ya we have more then we will likely land but we shal see how it goes | |
| 17:53:53 | sean-k-mooney | pci in palcemnt is technially complete we jsut have some cleanups and a bugfix still waiting ot merge | |
| 17:53:58 | sean-k-mooney | but the feature is fully merged | |
| 17:54:44 | sean-k-mooney | im hoping artom's fqdn change, shaids evacuate change and dansmits uuid change will merge in the next week | |
| 17:55:02 | sean-k-mooney | we will see i guess based on review bandwith | |
| 18:40:06 | opendevreview | Merged openstack/nova master: Follow up for the PCI in placement series https://review.opendev.org/c/openstack/nova/+/855654 | |
| 19:44:26 | opendevreview | Merged openstack/nova master: Rename _to_device_spec_conf to _to_list_of_json_str https://review.opendev.org/c/openstack/nova/+/855648 | |
| 23:49:45 | opendevreview | Merged openstack/nova master: Enable new defaults and scope checks by default https://review.opendev.org/c/openstack/nova/+/866218 | |
| 23:49:52 | opendevreview | Merged openstack/nova master: Remove use of removeprefix https://review.opendev.org/c/openstack/nova/+/867788 | |
| 23:56:46 | opendevreview | Merged openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358 | |
| 23:56:54 | opendevreview | Merged openstack/nova master: Reproduce PCI pool filtering bug https://review.opendev.org/c/openstack/nova/+/855649 | |
| #openstack-nova - 2023-01-17 | |||
| 00:06:20 | opendevreview | Merged openstack/nova master: Update Availability zone doc page https://review.opendev.org/c/openstack/nova/+/846463 | |
| 09:17:17 | bauzas | gibi: re: https://bugs.launchpad.net/nova/+bug/2002951 OOM | |
| 09:17:44 | bauzas | gibi: based on the example you gave, those are the tests that were run for the failing worker https://paste.opendev.org/show/bUSshY14qpkpDQ5jraEt/ | |
| 09:18:11 | gibi | nothing really jumps out from that list | |
| 09:18:15 | bauzas | me too | |
| 09:29:10 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 09:30:29 | bauzas | gibi: looks like the test was downloading the image when it stacktraced | |
| 09:32:58 | bauzas | wait, no | |
| 09:33:08 | bauzas | timings don't match | |
| 09:34:40 | gibi | I don't think OOM kill will cause a stack trace, the process will simply disappear | |
| 09:34:53 | bauzas | my bad | |
| 09:34:57 | bauzas | I meant when it was killed | |
| 09:35:17 | gibi | also as we discussed the point where the OOM hit might not be close to the point where the killed process used up the excessive memory | |
| 09:35:27 | bauzas | I'm trying to find where the test was when the worker got killed | |
| 09:41:28 | gibi | from this we can rule out that it is on a specific provider https://paste.opendev.org/show/b1CpIgnpVmLh4YCUOIar/ I see failures on ovh, rax, inmotion | |
| 09:42:04 | bauzas | gibi: TIL how to ask subunit from a CI log : | |
| 09:42:05 | bauzas | (venv) [sbauza@sbauza zuul-logs.9HEwdg]$ cat testrepository.subunit | subunit-filter -s --xfail --with-tag=worker-0 | subunit-ls | |
| 09:42:24 | bauzas | a grep does the same but not by the same manner :D | |
| 09:43:17 | bauzas | gibi: do you have any idea why I'm seeing a tempest call 30 mins before the run is run ? | |
| 09:43:24 | bauzas | before the *test is run ? | |
| 09:43:58 | gibi | TZ difference in log? | |
| 09:44:25 | bauzas | gibi: https://paste.opendev.org/show/bI0yvTNy52PzFSQUsGze/ | |
| 09:46:18 | gibi | maybe job-output.txt rendered after the job failed | |
| 09:46:22 | gibi | hm | |
| 09:47:08 | gibi | I would believe the tempest_log over the job-output.txt about the time steps | |
| 09:47:13 | bauzas | me too | |
| 09:47:26 | bauzas | but look, the image eventually was downloaded | |
| 09:47:30 | bauzas | we can see the log | |
| 09:47:44 | bauzas | which means the HTTP call was done | |
| 09:47:55 | gibi | the OOM hit at 22:31:13 based on syslog | |
| 09:48:01 | gibi | that matches the tempest_log timestamp | |
| 09:48:11 | bauzas | good point then | |
| 09:48:21 | bauzas | gibi: I briefly looked at glance logs | |
| 09:48:43 | bauzas | as I said, the image was apparently fully downloaded in 7-ish secs | |
| 09:51:24 | bauzas | oh wait | |
| 09:52:32 | bauzas | gibi: https://paste.opendev.org/show/bLFZGO2MZTYjdRRV3DCM/ | |
| 09:55:07 | bauzas | looks like we were caching the image | |
| 09:55:36 | bauzas | as we got the new path, and then nothing | |
| 09:56:08 | bauzas | and the timings match this time | |
| 10:16:18 | gibi | hm this is interesting, in all the 16 nova-ceph-multistore jobs that failed in the last 10 days the same test case got killed https://paste.opendev.org/show/bmEzF6rFgucUibd4CqTX/ | |
| 10:20:00 | bauzas | gibi: and I guess we'll see the same, which is we want to get the image | |
| 10:32:36 | bauzas | gibi: I'm curious btw., I've seen you using a logsearch tool | |
| 10:33:01 | bauzas | is that a CLI about https://opensearch.logs.openstack.org/ ? | |
| 10:33:24 | gibi | nope, it is https://github.com/gibizer/zuul-log-search | |
| 10:33:44 | gibi | my homebrew tool for grepping zuul logs | |
| 10:50:17 | gibi | I took 6 recent runs and generated the list of test cases run in the killed worker | |
| 10:50:35 | gibi | then I checked for intersection of the set of test cases | |
| 10:50:46 | gibi | and it is only tempest.api.compute.admin.test_volume.AttachSCSIVolumeTestJSON.test_attach_scsi_disk_with_config_drive | |
| 10:50:54 | gibi | the one that got killed | |
| 10:51:37 | gibi | so this points to that single test case as a cause | |
| 10:56:42 | bauzas | gibi: I'm just grabbing another change logs for looking whether the test was also killed while downloading the image | |
| 10:57:19 | gibi | ack | |
| 10:57:51 | bauzas | ah, your tool only downloads a specific file if I use --file | |
| 10:58:02 | bauzas | gibi: can I get all zuul logs from a specific change ? | |
| 10:58:57 | gibi | no, you need to use --file to get a log downloaded. I did it opt-in as I mostly use it to wide search and I wanted to limit the disk and bandwidth usage | |
| 10:59:15 | bauzas | k | |
| 10:59:15 | gibi | feel free to open an issue in the repo to add such option | |
| 10:59:29 | bauzas | I can workaround it for a sec | |
| 11:19:56 | opendevreview | Kashyap Chamarthy proposed openstack/nova master: libvirt: At start-up skip compareCPU() with a workaround https://review.opendev.org/c/openstack/nova/+/870794 | |
| 11:20:26 | kashyap | gibi: When you get a minute, can you have a quick look at the unit test? I know I messed it up slightly but how I'm unclear :/ | |
| 11:36:16 | bauzas | gibi: I tried to look at alot of failing jobs and all of them are indeed failing with the same test | |
| 11:36:35 | bauzas | I tried to find where in https://github.com/openstack/tempest/blob/master/tempest/api/compute/admin/test_volume.py#L76 we have the oomkiller | |
| 11:36:47 | bauzas | but as you said, maybe it's killed after a few seconds | |
| 13:01:15 | sean-k-mooney | bauzas: im going to add a specless bluepint the meeting adgenda and try and implement it before then. we we decided to defer it thats ok but if we agree its trivial enough i would liek to include it in A | |
| 13:03:21 | bauzas | sean-k-mooney: ack | |
| 13:43:21 | zigo | Is there some docs somewhere explaining how to implement an OpenStack wsgi API with keystone auth? | |
| 13:43:49 | zigo | FYI, I already got the db migration with Alembic done ... | |
| 13:43:59 | zigo | (plus oslo_config setup...) | |
| 13:45:45 | zigo | User docs are sometimes lacking info, dev docs are almost inexistant ... :( | |
| 13:55:10 | bauzas | zigo: you are deliberatly left with the choice you want | |
| 13:55:36 | bauzas | you just need to use keystonemiddleware lib | |
| 13:55:58 | bauzas | https://pypi.org/project/keystonemiddleware/ | |
| 13:56:25 | bauzas | https://docs.openstack.org/keystonemiddleware/latest/middlewarearchitecture.html describes the strategies you can choose for Auth'ing | |
| 13:57:01 | bauzas | a recommandation is to use paste for pipelining the WSGI middlewares | |
| 13:58:35 | zigo | Thanks. But there's no code example is shown in the keystonemiddleware's doc. | |
| 13:59:08 | zigo | Like many stuff, I'm stuck with a "look at other project, and attempt cut/past, then see what it does" strategy... | |
| 13:59:48 | bauzas | hah | |
| 13:59:50 | bauzas | that | |
| 14:00:07 | zigo | :) | |
| 14:00:15 | bauzas | yeah, in generall the overall workflow is prescribed, like in https://docs.openstack.org/project-team-guide/index.html | |
| 14:00:48 | bauzas | but beyond this, this is the project's team responsbility to decide how to implement what they want | |
| 14:00:55 | bauzas | like, the WSGI framework they prefer | |
| 14:01:09 | bauzas | or even the WSGI server they'd run with devstack | |
| 14:03:06 | bauzas | zigo: but honestly, the keystonemiddleware plugin isn't that hard to use | |
| 14:03:44 | zigo | I don't think that's the hardest part indeed. I just don't know where to start! :) | |