| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-16 | |||
| 12:20:44 | sean-k-mooney | supporting both formats is required becasue you are not allowd to requrie config change to upgrade | |
| 12:25:26 | sahid | yes that the point I don't want to break things. | |
| 12:25:58 | sahid | sean-k-mooney: I'm not sure to understand you mean we can update from bool to int transparently as this is using a string to int? | |
| 12:27:17 | sean-k-mooney | we cant do it transparently because its string to int | |
| 12:27:23 | sahid | oh.. but the pb in our case will be that, a True will not be converted to a int | |
| 12:27:28 | sean-k-mooney | in c it would be int to int | |
| 12:28:06 | sean-k-mooney | we still need to accpet yes/y/True ectra in the config | |
| 12:28:18 | sean-k-mooney | and that woudl have to be converted to 1 i guess | |
| 12:28:37 | sean-k-mooney | but you would also have to accpet 1 and any other values you are supproting | |
| 12:29:32 | sean-k-mooney | so ya after your change you still need to be able to handel a config with True in it as valid if it was to be transparent | |
| 12:29:39 | sean-k-mooney | so thats the problem in this case | |
| 12:33:17 | sahid | sean-k-mooney: is related to this one, if you have a moment to take a look https://review.opendev.org/c/openstack/nova/+/867324 | |
| 12:33:40 | sahid | basically it's to add ability to set number of retry | |
| 12:34:36 | sahid | originaly the option is Bool, and used to activated or desactive announces | |
| 12:34:43 | sean-k-mooney | ah that patch i saw that breifly fly by | |
| 12:34:58 | sean-k-mooney | honestly i would just add a second config option for the retry | |
| 12:35:04 | sahid | it's now needed to specify a number of retries and i wamted to avoid that we introduce a new option | |
| 12:35:32 | sean-k-mooney | yep but if we do this i think its just clean to add a new option and default to 1 or 3 | |
| 12:35:36 | sahid | yes... as it turn now it's basically what we will have to do | |
| 12:36:44 | sean-k-mooney | ya so workaround options still are treatd like normal config options so the same rules apply | |
| 12:36:46 | sahid | you mean we could harcored the number of retry instead, | |
| 12:37:01 | sean-k-mooney | in this case i would jsu tkeep the enable as a bool and add a retry option | |
| 12:37:17 | sean-k-mooney | well we coudl but im ok with a config option for the reties | |
| 12:37:41 | sean-k-mooney | ill just comment on the patch one sec. | |
| 12:37:50 | sahid | so one option to enable, one option to set the number of retries, and one option to specify the interval | |
| 12:37:59 | sahid | cool thank you | |
| 12:38:11 | sean-k-mooney | yep exactly | |
| 12:38:30 | sean-k-mooney | and we can set teh retires and interval to whatever we think is a good default | |
| 12:39:06 | sahid | ok fairenough :) | |
| 12:43:25 | sean-k-mooney | ok done i was suggesting 1 or 3 because 1 i sthe current behavior and 3 is what qemu defaults too when it sends them | |
| 13:37:58 | kashyap | sean-k-mooney: BTW a small data point on that "mpx" saga: if Nova doesn't break at the first CPU compare in check_cpu_compatibility(), then using "cpu_model_extra_flags=-mpx" works | |
| 13:38:26 | kashyap | (That gives a hint too that the first compare is wrong) | |
| 13:39:34 | sean-k-mooney | the way it should be working is we should be removing the mpx flag form all the modles listed in cpu_models and if any of them pass then we proceed as normal | |
| 13:40:09 | sean-k-mooney | so as long as any of the listed modeles work with the cpu_model_extra_flags option applied then we shoudl boot | |
| 13:40:19 | sean-k-mooney | /boot/start the agent/ | |
| 13:41:00 | sean-k-mooney | although really if any of them are invlied with that combination we shoudl reject it as an error | |
| 16:33:27 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 16:56:35 | opendevreview | Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812 | |
| 16:56:50 | artom | bauzas, ^^ | |
| 16:57:00 | bauzas | artom: ack, will look | |
| 17:39:04 | bauzas | damn, who knows the Launchpad nick of Kirill ? /me needs to paperwork the right ownership of https://blueprints.launchpad.net/nova/+spec/ironic-vnc-console | |
| 17:40:55 | bauzas | anyway, I can live with that | |
| 17:44:24 | bauzas | wow, the numbers of accepted blueprints for Antelope are identical to Yoga | |
| 17:44:37 | bauzas | disclaimer: this is gonna be a productive 5-week | |
| 17:53:13 | sean-k-mooney | ya we have more then we will likely land but we shal see how it goes | |
| 17:53:53 | sean-k-mooney | pci in palcemnt is technially complete we jsut have some cleanups and a bugfix still waiting ot merge | |
| 17:53:58 | sean-k-mooney | but the feature is fully merged | |
| 17:54:44 | sean-k-mooney | im hoping artom's fqdn change, shaids evacuate change and dansmits uuid change will merge in the next week | |
| 17:55:02 | sean-k-mooney | we will see i guess based on review bandwith | |
| 18:40:06 | opendevreview | Merged openstack/nova master: Follow up for the PCI in placement series https://review.opendev.org/c/openstack/nova/+/855654 | |
| 19:44:26 | opendevreview | Merged openstack/nova master: Rename _to_device_spec_conf to _to_list_of_json_str https://review.opendev.org/c/openstack/nova/+/855648 | |
| 23:49:45 | opendevreview | Merged openstack/nova master: Enable new defaults and scope checks by default https://review.opendev.org/c/openstack/nova/+/866218 | |
| 23:49:52 | opendevreview | Merged openstack/nova master: Remove use of removeprefix https://review.opendev.org/c/openstack/nova/+/867788 | |
| 23:56:46 | opendevreview | Merged openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358 | |
| 23:56:54 | opendevreview | Merged openstack/nova master: Reproduce PCI pool filtering bug https://review.opendev.org/c/openstack/nova/+/855649 | |
| #openstack-nova - 2023-01-17 | |||
| 00:06:20 | opendevreview | Merged openstack/nova master: Update Availability zone doc page https://review.opendev.org/c/openstack/nova/+/846463 | |
| 09:17:17 | bauzas | gibi: re: https://bugs.launchpad.net/nova/+bug/2002951 OOM | |
| 09:17:44 | bauzas | gibi: based on the example you gave, those are the tests that were run for the failing worker https://paste.opendev.org/show/bUSshY14qpkpDQ5jraEt/ | |
| 09:18:11 | gibi | nothing really jumps out from that list | |
| 09:18:15 | bauzas | me too | |
| 09:29:10 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 09:30:29 | bauzas | gibi: looks like the test was downloading the image when it stacktraced | |
| 09:32:58 | bauzas | wait, no | |
| 09:33:08 | bauzas | timings don't match | |
| 09:34:40 | gibi | I don't think OOM kill will cause a stack trace, the process will simply disappear | |
| 09:34:53 | bauzas | my bad | |
| 09:34:57 | bauzas | I meant when it was killed | |
| 09:35:17 | gibi | also as we discussed the point where the OOM hit might not be close to the point where the killed process used up the excessive memory | |
| 09:35:27 | bauzas | I'm trying to find where the test was when the worker got killed | |
| 09:41:28 | gibi | from this we can rule out that it is on a specific provider https://paste.opendev.org/show/b1CpIgnpVmLh4YCUOIar/ I see failures on ovh, rax, inmotion | |
| 09:42:04 | bauzas | gibi: TIL how to ask subunit from a CI log : | |
| 09:42:05 | bauzas | (venv) [sbauza@sbauza zuul-logs.9HEwdg]$ cat testrepository.subunit | subunit-filter -s --xfail --with-tag=worker-0 | subunit-ls | |
| 09:42:24 | bauzas | a grep does the same but not by the same manner :D | |
| 09:43:17 | bauzas | gibi: do you have any idea why I'm seeing a tempest call 30 mins before the run is run ? | |
| 09:43:24 | bauzas | before the *test is run ? | |
| 09:43:58 | gibi | TZ difference in log? | |
| 09:44:25 | bauzas | gibi: https://paste.opendev.org/show/bI0yvTNy52PzFSQUsGze/ | |
| 09:46:18 | gibi | maybe job-output.txt rendered after the job failed | |
| 09:46:22 | gibi | hm | |
| 09:47:08 | gibi | I would believe the tempest_log over the job-output.txt about the time steps | |
| 09:47:13 | bauzas | me too | |
| 09:47:26 | bauzas | but look, the image eventually was downloaded | |
| 09:47:30 | bauzas | we can see the log | |
| 09:47:44 | bauzas | which means the HTTP call was done | |
| 09:47:55 | gibi | the OOM hit at 22:31:13 based on syslog | |
| 09:48:01 | gibi | that matches the tempest_log timestamp | |
| 09:48:11 | bauzas | good point then | |
| 09:48:21 | bauzas | gibi: I briefly looked at glance logs | |
| 09:48:43 | bauzas | as I said, the image was apparently fully downloaded in 7-ish secs | |
| 09:51:24 | bauzas | oh wait | |
| 09:52:32 | bauzas | gibi: https://paste.opendev.org/show/bLFZGO2MZTYjdRRV3DCM/ | |
| 09:55:07 | bauzas | looks like we were caching the image | |
| 09:55:36 | bauzas | as we got the new path, and then nothing | |
| 09:56:08 | bauzas | and the timings match this time | |
| 10:16:18 | gibi | hm this is interesting, in all the 16 nova-ceph-multistore jobs that failed in the last 10 days the same test case got killed https://paste.opendev.org/show/bmEzF6rFgucUibd4CqTX/ | |
| 10:20:00 | bauzas | gibi: and I guess we'll see the same, which is we want to get the image | |
| 10:32:36 | bauzas | gibi: I'm curious btw., I've seen you using a logsearch tool | |
| 10:33:01 | bauzas | is that a CLI about https://opensearch.logs.openstack.org/ ? | |
| 10:33:24 | gibi | nope, it is https://github.com/gibizer/zuul-log-search | |
| 10:33:44 | gibi | my homebrew tool for grepping zuul logs | |