| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-26 | |||
| 00:48:32 | melwitt | yeah, it's clear it's not a thing that we want to be permanent and why it wouldn't go into the normal configs | |
| 00:49:16 | sean-k-mooney | i did consider just suggesting hardcoding to 3 with a longer interval | |
| 00:49:34 | sean-k-mooney | but they already had the config option when i review for the interval | |
| 00:50:00 | sean-k-mooney | so i kind fo didnt wnat to have anouther patch tweeking this again later | |
| 00:51:28 | sean-k-mooney | https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.device_detach_attempts and https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.device_detach_timeout are the detach/attach options we added that are kind of like it | |
| 00:51:41 | melwitt | yeah, would've been ideal to hardcode it but if it's that fiddly then I see why we wouldn't want to have to do future tweaks to it | |
| 00:52:50 | sean-k-mooney | it kind of sucks that the workaround is needed but ya i expected when we added the orgianal workaround ot not need to do more then one addtional GARP | |
| 00:53:07 | melwitt | ack, I don't think the concept is itself weird, it's that you get a lot of fine-grained tuning for your one workaround 😆 I'm not suggesting blocking it but just saying it looks quite odd to me | |
| 00:53:07 | melwitt | ack, I don't think the concept is itself weird, it's that you get a lot of fine-grained tuning for your one workaround 😆 I'm not suggesting blocking it but just saying it looks quite odd to me | |
| 00:53:07 | sean-k-mooney | qemu is already doing 3 before we do anything | |
| 00:53:54 | sean-k-mooney | i agree on the odd but its pargmatic | |
| 00:54:23 | sean-k-mooney | its kind of like shouting at the network to say hay i really reallly really am over here now | |
| 00:54:56 | melwitt | lol :) | |
| 00:55:46 | sean-k-mooney | anyway its like 1am for me so im going to sleep came onlien to check something breilfy and got distracted with geting my email inbox to 0 | |
| 00:56:29 | melwitt | k g'night! o/ | |
| 00:56:36 | sean-k-mooney | o/ | |
| 07:57:40 | tobias-urdin | eu people might start their day now :) please have a look at getting a release for the CVE out https://review.opendev.org/c/openstack/releases/+/871802 | |
| 08:43:10 | frickler | gibi: bauzas: sean-k-mooney: ^^ I kind of agree with tobias-urdin that there is a bit of urgency behind this. not sure if release team could skip ptl approval in this case, though? | |
| 08:43:50 | kashyap | Is it just me or the "tempest-integrated-compute" is timing out often for others too? | |
| 08:44:22 | bauzas | frickler: I'm here | |
| 08:44:52 | bauzas | frickler: we said we were trying to also merge another CVE fix before we release this | |
| 08:50:37 | bauzas | ok, the other cve bug is fixed down to xena, so yeah, we can have releases | |
| 08:56:41 | frickler | that other CVE was https://review.opendev.org/c/openstack/nova/+/859315, right? | |
| 09:04:51 | bauzas | yup, just checked, reviewing now the releases patch | |
| 09:09:21 | bauzas | gibi: frickler: I was torn with the proposed semver of https://review.opendev.org/c/openstack/releases/+/871802 which was doing .y releases but I can live with that | |
| 09:10:18 | bauzas | tobias-urdin: ^ anything in particular you had in mind when you did set the releases for a .y bump ? | |
| 09:10:30 | bauzas | the fact that it was exposing a new conf knob, I guess ? | |
| 09:14:03 | gibi | I'm fine with the minor bump | |
| 09:15:05 | bauzas | that seems a bit agressive but thinking out more, that means that distros have to adapt their toolings if they wanna set the new conf knob | |
| 09:15:09 | bauzas | so yeah a .y bump seems ok | |
| 09:15:32 | bauzas | even if semantically, we're sending a wrong signal | |
| 09:17:16 | gibi | we are removing functionality with a knob by default so I'm OK to y bump it | |
| 09:17:53 | gibi | I double checked it seems both cve is in the release (the vnit_type on was in zed when it was master) | |
| 09:18:06 | gibi | so I think we are good to go | |
| 09:20:10 | tobias-urdin | i pretty much followed cinder that bumped minor, I guess it wouldn't hurt indicating to operations that a minor version that should be upgraded to because of CVE | |
| 09:20:38 | tobias-urdin | but yeah we can change if required, just wanted to make it a priority to release it so downstream can start building stuff new versions as well | |
| 09:20:48 | bauzas | gibi: yeah, checked the other CVE, was my main original driver for the check | |
| 09:21:07 | bauzas | tobias-urdin: no worries, as I said, I can live with that | |
| 09:21:24 | gibi | next is stable/wallaby but that needs the tempest pin first. https://review.opendev.org/q/topic:wallaby-pin-tempest | |
| 09:21:32 | bauzas | in theory a CVE fix doesn't require a y versioning but meh | |
| 09:21:46 | tobias-urdin | bauzas: ack, thanks, i will keep that in mind for the future | |
| 09:21:49 | bauzas | gibi: correct, I +2/+Wd a patch this morning | |
| 09:22:43 | bauzas | tobias-urdin: np, not anyone needs to know anything :) but if you wanna know more about semver, this is the reference page https://docs.openstack.org/pbr/latest/user/semver.html | |
| 09:23:10 | bauzas | gibi: do you know if gmann did the tempest patch ? | |
| 09:23:17 | bauzas | I can check, I just didn't had the time yet | |
| 09:23:32 | gibi | bauzas: here is the tempest pin series https://review.opendev.org/q/topic:wallaby-pin-tempest it needs love | |
| 09:23:44 | bauzas | I can surely provide love | |
| 09:23:45 | gibi | the DNM test patches are failing | |
| 09:24:33 | bauzas | then I guess the love has to be on finding why the DNM patches are failing | |
| 09:24:35 | bauzas | lovely | |
| 09:24:51 | bauzas | that's just 12 hours I haven't looked at zuul files | |
| 09:48:18 | bauzas | mmm | |
| 09:48:40 | bauzas | gibi: does those skipttest exceptions in tempest look correct to you ? | |
| 09:48:42 | bauzas | 2023-01-26 01:06:13.455899 | controller | unittest2.case.SkipTest: Identity api v2 is not enabled | |
| 09:48:51 | bauzas | I have seen gmann rechecking on such errors | |
| 09:49:06 | bauzas | https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_0a4/871782/2/check/tempest-full-py3/0a42465/job-output.txt | |
| 09:49:17 | gibi | it does not look good, but maybe gmann rechecked on it as he changed one of the depends-ons | |
| 09:50:31 | bauzas | I just asked for a recheck | |
| 09:50:44 | bauzas | on the devstack patch | |
| 09:51:03 | bauzas | anyway, the changes themselves on both tempest and devstack seem logic to me | |
| 09:51:04 | gibi | we will see | |
| 09:51:12 | bauzas | so I guess this is just a transient issue | |
| 09:52:22 | bauzas | at least keystone was running | |
| 09:53:48 | sahid | o/ sean-k-mooney, bauzas I have noticed your new comments | |
| 09:53:51 | sahid | working on it ! | |
| 09:53:53 | sahid | thanks | |
| 09:54:33 | bauzas | sahid: thanks | |
| 09:54:47 | bauzas | sahid: ping me when you're done with those, and I'll rereview | |
| 09:55:13 | sean-k-mooney[m] | the rpc pin case is not something i orginally tought of but i dont think its a large change to just make sure we dont retrun a 500 from the api and add a test for that so hopefully it wont take too much to adress that | |
| 09:55:33 | bauzas | sahid: to clarify, sorry but you don't need to change the RPC client, just make sure that on the API you can verify it | |
| 09:56:26 | sean-k-mooney[m] | yep so add a functional test that pins to 6.0 | |
| 09:56:32 | sean-k-mooney[m] | use the new microversion | |
| 09:56:46 | sean-k-mooney[m] | and assert an excption is raised | |
| 09:56:53 | sean-k-mooney[m] | ideally it should be a 409 | |
| 09:57:06 | sean-k-mooney[m] | the same as when the compute is not upgraded | |
| 09:57:53 | sean-k-mooney[m] | you should not expose the crrent RPC pin in the excption | |
| 09:58:48 | sean-k-mooney[m] | just that the could does not meet the requirements for the new microversion and you should use the old one like the other exception | |
| 09:59:35 | sean-k-mooney[m] | bauzas actully is there any reason not to use the same excpetion here as in the api when the compute service is to low | |
| 10:00:04 | bauzas | good question | |
| 10:00:29 | bauzas | problem is, evacuation is defined by policy, right? | |
| 10:00:52 | sean-k-mooney[m] | i guess its and admin api and they might want to know but if they manually pinned. perhapse its diffent admins that do upgrade vs day to day | |
| 10:01:04 | sean-k-mooney[m] | well all apis are controlable by policy | |
| 10:01:06 | bauzas | so you can change the policy to have evacuation (without a host param) be supported for endusers | |
| 10:01:28 | bauzas | if so, we could return something like 'sorry, compute service is too low' | |
| 10:01:31 | sean-k-mooney[m] | you could | |
| 10:01:45 | bauzas | I don't know whether it would be a problem for our operators then | |
| 10:01:56 | kashyap | This timeout is blocking a couple of patches. /me is trying to find where exactly is the time_out -- https://zuul.opendev.org/t/openstack/build/453f991eeeb34498b132eb84de3301db/logs | |
| 10:02:43 | sean-k-mooney[m] | sound like just a slow node to be honest | |
| 10:03:05 | sean-k-mooney[m] | and that job might be hitting up agagisnt the timeout anyway but lests see what the normal runtime is | |
| 10:03:25 | kashyap | So only a full recheck is the only option? :( | |
| 10:03:33 | kashyap | (Already did it once) | |
| 10:03:37 | sean-k-mooney[m] | normally 90 mins or so | |
| 10:03:51 | kashyap | sean-k-mooney[m]: For the full recheck? | |
| 10:04:45 | sean-k-mooney[m] | that job normally takes 90 mins | |
| 10:04:46 | sean-k-mooney[m] | https://zuul.opendev.org/t/openstack/builds?job_name=tempest-integrated-compute&project=openstack/nova | |
| 10:05:26 | sean-k-mooney[m] | there have been 5 time outs in the last 300 runs of that job | |
| 10:10:07 | sean-k-mooney[m] | the 3 time outs are form 4 providers so i dont really see any corralation | |
| 10:10:33 | sean-k-mooney[m] | it would be good to see if there is anything odd in the devstack or tempet runs | |
| 10:10:49 | sean-k-mooney[m] | but something took more time then normal | |
| 10:11:31 | sean-k-mooney[m] | so yes a recheck is the way to proceed but it would be good to see if say a lot of swap was used or an image/pacakge download was slow | |