| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-16 | |||
| 12:37:59 | sahid | cool thank you | |
| 12:38:11 | sean-k-mooney | yep exactly | |
| 12:38:30 | sean-k-mooney | and we can set teh retires and interval to whatever we think is a good default | |
| 12:39:06 | sahid | ok fairenough :) | |
| 12:43:25 | sean-k-mooney | ok done i was suggesting 1 or 3 because 1 i sthe current behavior and 3 is what qemu defaults too when it sends them | |
| 13:37:58 | kashyap | sean-k-mooney: BTW a small data point on that "mpx" saga: if Nova doesn't break at the first CPU compare in check_cpu_compatibility(), then using "cpu_model_extra_flags=-mpx" works | |
| 13:38:26 | kashyap | (That gives a hint too that the first compare is wrong) | |
| 13:39:34 | sean-k-mooney | the way it should be working is we should be removing the mpx flag form all the modles listed in cpu_models and if any of them pass then we proceed as normal | |
| 13:40:09 | sean-k-mooney | so as long as any of the listed modeles work with the cpu_model_extra_flags option applied then we shoudl boot | |
| 13:40:19 | sean-k-mooney | /boot/start the agent/ | |
| 13:41:00 | sean-k-mooney | although really if any of them are invlied with that combination we shoudl reject it as an error | |
| 16:33:27 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 16:56:35 | opendevreview | Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812 | |
| 16:56:50 | artom | bauzas, ^^ | |
| 16:57:00 | bauzas | artom: ack, will look | |
| 17:39:04 | bauzas | damn, who knows the Launchpad nick of Kirill ? /me needs to paperwork the right ownership of https://blueprints.launchpad.net/nova/+spec/ironic-vnc-console | |
| 17:40:55 | bauzas | anyway, I can live with that | |
| 17:44:24 | bauzas | wow, the numbers of accepted blueprints for Antelope are identical to Yoga | |
| 17:44:37 | bauzas | disclaimer: this is gonna be a productive 5-week | |
| 17:53:13 | sean-k-mooney | ya we have more then we will likely land but we shal see how it goes | |
| 17:53:53 | sean-k-mooney | pci in palcemnt is technially complete we jsut have some cleanups and a bugfix still waiting ot merge | |
| 17:53:58 | sean-k-mooney | but the feature is fully merged | |
| 17:54:44 | sean-k-mooney | im hoping artom's fqdn change, shaids evacuate change and dansmits uuid change will merge in the next week | |
| 17:55:02 | sean-k-mooney | we will see i guess based on review bandwith | |
| 18:40:06 | opendevreview | Merged openstack/nova master: Follow up for the PCI in placement series https://review.opendev.org/c/openstack/nova/+/855654 | |
| 19:44:26 | opendevreview | Merged openstack/nova master: Rename _to_device_spec_conf to _to_list_of_json_str https://review.opendev.org/c/openstack/nova/+/855648 | |
| 23:49:45 | opendevreview | Merged openstack/nova master: Enable new defaults and scope checks by default https://review.opendev.org/c/openstack/nova/+/866218 | |
| 23:49:52 | opendevreview | Merged openstack/nova master: Remove use of removeprefix https://review.opendev.org/c/openstack/nova/+/867788 | |
| 23:56:46 | opendevreview | Merged openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358 | |
| 23:56:54 | opendevreview | Merged openstack/nova master: Reproduce PCI pool filtering bug https://review.opendev.org/c/openstack/nova/+/855649 | |
| #openstack-nova - 2023-01-17 | |||
| 00:06:20 | opendevreview | Merged openstack/nova master: Update Availability zone doc page https://review.opendev.org/c/openstack/nova/+/846463 | |
| 09:17:17 | bauzas | gibi: re: https://bugs.launchpad.net/nova/+bug/2002951 OOM | |
| 09:17:44 | bauzas | gibi: based on the example you gave, those are the tests that were run for the failing worker https://paste.opendev.org/show/bUSshY14qpkpDQ5jraEt/ | |
| 09:18:11 | gibi | nothing really jumps out from that list | |
| 09:18:15 | bauzas | me too | |
| 09:29:10 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 09:30:29 | bauzas | gibi: looks like the test was downloading the image when it stacktraced | |
| 09:32:58 | bauzas | wait, no | |
| 09:33:08 | bauzas | timings don't match | |
| 09:34:40 | gibi | I don't think OOM kill will cause a stack trace, the process will simply disappear | |
| 09:34:53 | bauzas | my bad | |
| 09:34:57 | bauzas | I meant when it was killed | |
| 09:35:17 | gibi | also as we discussed the point where the OOM hit might not be close to the point where the killed process used up the excessive memory | |
| 09:35:27 | bauzas | I'm trying to find where the test was when the worker got killed | |
| 09:41:28 | gibi | from this we can rule out that it is on a specific provider https://paste.opendev.org/show/b1CpIgnpVmLh4YCUOIar/ I see failures on ovh, rax, inmotion | |
| 09:42:04 | bauzas | gibi: TIL how to ask subunit from a CI log : | |
| 09:42:05 | bauzas | (venv) [sbauza@sbauza zuul-logs.9HEwdg]$ cat testrepository.subunit | subunit-filter -s --xfail --with-tag=worker-0 | subunit-ls | |
| 09:42:24 | bauzas | a grep does the same but not by the same manner :D | |
| 09:43:17 | bauzas | gibi: do you have any idea why I'm seeing a tempest call 30 mins before the run is run ? | |
| 09:43:24 | bauzas | before the *test is run ? | |
| 09:43:58 | gibi | TZ difference in log? | |
| 09:44:25 | bauzas | gibi: https://paste.opendev.org/show/bI0yvTNy52PzFSQUsGze/ | |
| 09:46:18 | gibi | maybe job-output.txt rendered after the job failed | |
| 09:46:22 | gibi | hm | |
| 09:47:08 | gibi | I would believe the tempest_log over the job-output.txt about the time steps | |
| 09:47:13 | bauzas | me too | |
| 09:47:26 | bauzas | but look, the image eventually was downloaded | |
| 09:47:30 | bauzas | we can see the log | |
| 09:47:44 | bauzas | which means the HTTP call was done | |
| 09:47:55 | gibi | the OOM hit at 22:31:13 based on syslog | |
| 09:48:01 | gibi | that matches the tempest_log timestamp | |
| 09:48:11 | bauzas | good point then | |
| 09:48:21 | bauzas | gibi: I briefly looked at glance logs | |
| 09:48:43 | bauzas | as I said, the image was apparently fully downloaded in 7-ish secs | |
| 09:51:24 | bauzas | oh wait | |
| 09:52:32 | bauzas | gibi: https://paste.opendev.org/show/bLFZGO2MZTYjdRRV3DCM/ | |
| 09:55:07 | bauzas | looks like we were caching the image | |
| 09:55:36 | bauzas | as we got the new path, and then nothing | |
| 09:56:08 | bauzas | and the timings match this time | |
| 10:16:18 | gibi | hm this is interesting, in all the 16 nova-ceph-multistore jobs that failed in the last 10 days the same test case got killed https://paste.opendev.org/show/bmEzF6rFgucUibd4CqTX/ | |
| 10:20:00 | bauzas | gibi: and I guess we'll see the same, which is we want to get the image | |
| 10:32:36 | bauzas | gibi: I'm curious btw., I've seen you using a logsearch tool | |
| 10:33:01 | bauzas | is that a CLI about https://opensearch.logs.openstack.org/ ? | |
| 10:33:24 | gibi | nope, it is https://github.com/gibizer/zuul-log-search | |
| 10:33:44 | gibi | my homebrew tool for grepping zuul logs | |
| 10:50:17 | gibi | I took 6 recent runs and generated the list of test cases run in the killed worker | |
| 10:50:35 | gibi | then I checked for intersection of the set of test cases | |
| 10:50:46 | gibi | and it is only tempest.api.compute.admin.test_volume.AttachSCSIVolumeTestJSON.test_attach_scsi_disk_with_config_drive | |
| 10:50:54 | gibi | the one that got killed | |
| 10:51:37 | gibi | so this points to that single test case as a cause | |
| 10:56:42 | bauzas | gibi: I'm just grabbing another change logs for looking whether the test was also killed while downloading the image | |
| 10:57:19 | gibi | ack | |
| 10:57:51 | bauzas | ah, your tool only downloads a specific file if I use --file | |
| 10:58:02 | bauzas | gibi: can I get all zuul logs from a specific change ? | |
| 10:58:57 | gibi | no, you need to use --file to get a log downloaded. I did it opt-in as I mostly use it to wide search and I wanted to limit the disk and bandwidth usage | |
| 10:59:15 | bauzas | k | |
| 10:59:15 | gibi | feel free to open an issue in the repo to add such option | |
| 10:59:29 | bauzas | I can workaround it for a sec | |
| 11:19:56 | opendevreview | Kashyap Chamarthy proposed openstack/nova master: libvirt: At start-up skip compareCPU() with a workaround https://review.opendev.org/c/openstack/nova/+/870794 | |
| 11:20:26 | kashyap | gibi: When you get a minute, can you have a quick look at the unit test? I know I messed it up slightly but how I'm unclear :/ | |
| 11:36:16 | bauzas | gibi: I tried to look at alot of failing jobs and all of them are indeed failing with the same test | |
| 11:36:35 | bauzas | I tried to find where in https://github.com/openstack/tempest/blob/master/tempest/api/compute/admin/test_volume.py#L76 we have the oomkiller | |
| 11:36:47 | bauzas | but as you said, maybe it's killed after a few seconds | |
| 13:01:15 | sean-k-mooney | bauzas: im going to add a specless bluepint the meeting adgenda and try and implement it before then. we we decided to defer it thats ok but if we agree its trivial enough i would liek to include it in A | |
| 13:03:21 | bauzas | sean-k-mooney: ack | |
| 13:43:21 | zigo | Is there some docs somewhere explaining how to implement an OpenStack wsgi API with keystone auth? | |
| 13:43:49 | zigo | FYI, I already got the db migration with Alembic done ... | |
| 13:43:59 | zigo | (plus oslo_config setup...) | |
| 13:45:45 | zigo | User docs are sometimes lacking info, dev docs are almost inexistant ... :( | |
| 13:55:10 | bauzas | zigo: you are deliberatly left with the choice you want | |