| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-16 | |||
| 12:07:32 | kashyap | sean-k-mooney: I'm in agreement with you on definitely using the newer API, as that's a net-benefit. (I was not debating that one.) | |
| 12:11:57 | sahid | o/ guys, do we have a process to convert an option from bool to int? | |
| 12:17:57 | sean-k-mooney | sahid: we do it called not doing it. basically if your changing the type you have to rename the option and deprecat the old one. in this case your going form bool which in our congi is based on string to int | |
| 12:18:42 | sean-k-mooney | you can do that in plcaee befause we accpet true/yes|false/no not just 1|0 | |
| 12:19:16 | sean-k-mooney | sahid: what config option do you want to modify | |
| 12:19:52 | sean-k-mooney | you will basically have to deprecate the old one and add a new one in the new format and support both in the A cycle. | |
| 12:20:44 | sean-k-mooney | supporting both formats is required becasue you are not allowd to requrie config change to upgrade | |
| 12:25:26 | sahid | yes that the point I don't want to break things. | |
| 12:25:58 | sahid | sean-k-mooney: I'm not sure to understand you mean we can update from bool to int transparently as this is using a string to int? | |
| 12:27:17 | sean-k-mooney | we cant do it transparently because its string to int | |
| 12:27:23 | sahid | oh.. but the pb in our case will be that, a True will not be converted to a int | |
| 12:27:28 | sean-k-mooney | in c it would be int to int | |
| 12:28:06 | sean-k-mooney | we still need to accpet yes/y/True ectra in the config | |
| 12:28:18 | sean-k-mooney | and that woudl have to be converted to 1 i guess | |
| 12:28:37 | sean-k-mooney | but you would also have to accpet 1 and any other values you are supproting | |
| 12:29:32 | sean-k-mooney | so ya after your change you still need to be able to handel a config with True in it as valid if it was to be transparent | |
| 12:29:39 | sean-k-mooney | so thats the problem in this case | |
| 12:33:17 | sahid | sean-k-mooney: is related to this one, if you have a moment to take a look https://review.opendev.org/c/openstack/nova/+/867324 | |
| 12:33:40 | sahid | basically it's to add ability to set number of retry | |
| 12:34:36 | sahid | originaly the option is Bool, and used to activated or desactive announces | |
| 12:34:43 | sean-k-mooney | ah that patch i saw that breifly fly by | |
| 12:34:58 | sean-k-mooney | honestly i would just add a second config option for the retry | |
| 12:35:04 | sahid | it's now needed to specify a number of retries and i wamted to avoid that we introduce a new option | |
| 12:35:32 | sean-k-mooney | yep but if we do this i think its just clean to add a new option and default to 1 or 3 | |
| 12:35:36 | sahid | yes... as it turn now it's basically what we will have to do | |
| 12:36:44 | sean-k-mooney | ya so workaround options still are treatd like normal config options so the same rules apply | |
| 12:36:46 | sahid | you mean we could harcored the number of retry instead, | |
| 12:37:01 | sean-k-mooney | in this case i would jsu tkeep the enable as a bool and add a retry option | |
| 12:37:17 | sean-k-mooney | well we coudl but im ok with a config option for the reties | |
| 12:37:41 | sean-k-mooney | ill just comment on the patch one sec. | |
| 12:37:50 | sahid | so one option to enable, one option to set the number of retries, and one option to specify the interval | |
| 12:37:59 | sahid | cool thank you | |
| 12:38:11 | sean-k-mooney | yep exactly | |
| 12:38:30 | sean-k-mooney | and we can set teh retires and interval to whatever we think is a good default | |
| 12:39:06 | sahid | ok fairenough :) | |
| 12:43:25 | sean-k-mooney | ok done i was suggesting 1 or 3 because 1 i sthe current behavior and 3 is what qemu defaults too when it sends them | |
| 13:37:58 | kashyap | sean-k-mooney: BTW a small data point on that "mpx" saga: if Nova doesn't break at the first CPU compare in check_cpu_compatibility(), then using "cpu_model_extra_flags=-mpx" works | |
| 13:38:26 | kashyap | (That gives a hint too that the first compare is wrong) | |
| 13:39:34 | sean-k-mooney | the way it should be working is we should be removing the mpx flag form all the modles listed in cpu_models and if any of them pass then we proceed as normal | |
| 13:40:09 | sean-k-mooney | so as long as any of the listed modeles work with the cpu_model_extra_flags option applied then we shoudl boot | |
| 13:40:19 | sean-k-mooney | /boot/start the agent/ | |
| 13:41:00 | sean-k-mooney | although really if any of them are invlied with that combination we shoudl reject it as an error | |
| 16:33:27 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 16:56:35 | opendevreview | Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812 | |
| 16:56:50 | artom | bauzas, ^^ | |
| 16:57:00 | bauzas | artom: ack, will look | |
| 17:39:04 | bauzas | damn, who knows the Launchpad nick of Kirill ? /me needs to paperwork the right ownership of https://blueprints.launchpad.net/nova/+spec/ironic-vnc-console | |
| 17:40:55 | bauzas | anyway, I can live with that | |
| 17:44:24 | bauzas | wow, the numbers of accepted blueprints for Antelope are identical to Yoga | |
| 17:44:37 | bauzas | disclaimer: this is gonna be a productive 5-week | |
| 17:53:13 | sean-k-mooney | ya we have more then we will likely land but we shal see how it goes | |
| 17:53:53 | sean-k-mooney | pci in palcemnt is technially complete we jsut have some cleanups and a bugfix still waiting ot merge | |
| 17:53:58 | sean-k-mooney | but the feature is fully merged | |
| 17:54:44 | sean-k-mooney | im hoping artom's fqdn change, shaids evacuate change and dansmits uuid change will merge in the next week | |
| 17:55:02 | sean-k-mooney | we will see i guess based on review bandwith | |
| 18:40:06 | opendevreview | Merged openstack/nova master: Follow up for the PCI in placement series https://review.opendev.org/c/openstack/nova/+/855654 | |
| 19:44:26 | opendevreview | Merged openstack/nova master: Rename _to_device_spec_conf to _to_list_of_json_str https://review.opendev.org/c/openstack/nova/+/855648 | |
| 23:49:45 | opendevreview | Merged openstack/nova master: Enable new defaults and scope checks by default https://review.opendev.org/c/openstack/nova/+/866218 | |
| 23:49:52 | opendevreview | Merged openstack/nova master: Remove use of removeprefix https://review.opendev.org/c/openstack/nova/+/867788 | |
| 23:56:46 | opendevreview | Merged openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358 | |
| 23:56:54 | opendevreview | Merged openstack/nova master: Reproduce PCI pool filtering bug https://review.opendev.org/c/openstack/nova/+/855649 | |
| #openstack-nova - 2023-01-17 | |||
| 00:06:20 | opendevreview | Merged openstack/nova master: Update Availability zone doc page https://review.opendev.org/c/openstack/nova/+/846463 | |
| 09:17:17 | bauzas | gibi: re: https://bugs.launchpad.net/nova/+bug/2002951 OOM | |
| 09:17:44 | bauzas | gibi: based on the example you gave, those are the tests that were run for the failing worker https://paste.opendev.org/show/bUSshY14qpkpDQ5jraEt/ | |
| 09:18:11 | gibi | nothing really jumps out from that list | |
| 09:18:15 | bauzas | me too | |
| 09:29:10 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 09:30:29 | bauzas | gibi: looks like the test was downloading the image when it stacktraced | |
| 09:32:58 | bauzas | wait, no | |
| 09:33:08 | bauzas | timings don't match | |
| 09:34:40 | gibi | I don't think OOM kill will cause a stack trace, the process will simply disappear | |
| 09:34:53 | bauzas | my bad | |
| 09:34:57 | bauzas | I meant when it was killed | |
| 09:35:17 | gibi | also as we discussed the point where the OOM hit might not be close to the point where the killed process used up the excessive memory | |
| 09:35:27 | bauzas | I'm trying to find where the test was when the worker got killed | |
| 09:41:28 | gibi | from this we can rule out that it is on a specific provider https://paste.opendev.org/show/b1CpIgnpVmLh4YCUOIar/ I see failures on ovh, rax, inmotion | |
| 09:42:04 | bauzas | gibi: TIL how to ask subunit from a CI log : | |
| 09:42:05 | bauzas | (venv) [sbauza@sbauza zuul-logs.9HEwdg]$ cat testrepository.subunit | subunit-filter -s --xfail --with-tag=worker-0 | subunit-ls | |
| 09:42:24 | bauzas | a grep does the same but not by the same manner :D | |
| 09:43:17 | bauzas | gibi: do you have any idea why I'm seeing a tempest call 30 mins before the run is run ? | |
| 09:43:24 | bauzas | before the *test is run ? | |
| 09:43:58 | gibi | TZ difference in log? | |
| 09:44:25 | bauzas | gibi: https://paste.opendev.org/show/bI0yvTNy52PzFSQUsGze/ | |
| 09:46:18 | gibi | maybe job-output.txt rendered after the job failed | |
| 09:46:22 | gibi | hm | |
| 09:47:08 | gibi | I would believe the tempest_log over the job-output.txt about the time steps | |
| 09:47:13 | bauzas | me too | |
| 09:47:26 | bauzas | but look, the image eventually was downloaded | |
| 09:47:30 | bauzas | we can see the log | |
| 09:47:44 | bauzas | which means the HTTP call was done | |
| 09:47:55 | gibi | the OOM hit at 22:31:13 based on syslog | |
| 09:48:01 | gibi | that matches the tempest_log timestamp | |
| 09:48:11 | bauzas | good point then | |
| 09:48:21 | bauzas | gibi: I briefly looked at glance logs | |
| 09:48:43 | bauzas | as I said, the image was apparently fully downloaded in 7-ish secs | |
| 09:51:24 | bauzas | oh wait | |
| 09:52:32 | bauzas | gibi: https://paste.opendev.org/show/bLFZGO2MZTYjdRRV3DCM/ | |
| 09:55:07 | bauzas | looks like we were caching the image | |
| 09:55:36 | bauzas | as we got the new path, and then nothing | |
| 09:56:08 | bauzas | and the timings match this time | |