Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-26
00:48:32 melwitt yeah, it's clear it's not a thing that we want to be permanent and why it wouldn't go into the normal configs
00:49:16 sean-k-mooney i did consider just suggesting hardcoding to 3 with a longer interval
00:49:34 sean-k-mooney but they already had the config option when i review for the interval
00:50:00 sean-k-mooney so i kind fo didnt wnat to have anouther patch tweeking this again later
00:51:28 sean-k-mooney https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.device_detach_attempts and https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.device_detach_timeout are the detach/attach options we added that are kind of like it
00:51:41 melwitt yeah, would've been ideal to hardcode it but if it's that fiddly then I see why we wouldn't want to have to do future tweaks to it
00:52:50 sean-k-mooney it kind of sucks that the workaround is needed but ya i expected when we added the orgianal workaround ot not need to do more then one addtional GARP
00:53:07 melwitt ack, I don't think the concept is itself weird, it's that you get a lot of fine-grained tuning for your one workaround 😆 I'm not suggesting blocking it but just saying it looks quite odd to me
00:53:07 melwitt ack, I don't think the concept is itself weird, it's that you get a lot of fine-grained tuning for your one workaround 😆 I'm not suggesting blocking it but just saying it looks quite odd to me
00:53:07 sean-k-mooney qemu is already doing 3 before we do anything
00:53:54 sean-k-mooney i agree on the odd but its pargmatic
00:54:23 sean-k-mooney its kind of like shouting at the network to say hay i really reallly really am over here now
00:54:56 melwitt lol :)
00:55:46 sean-k-mooney anyway its like 1am for me so im going to sleep came onlien to check something breilfy and got distracted with geting my email inbox to 0
00:56:29 melwitt k g'night! o/
00:56:36 sean-k-mooney o/
07:57:40 tobias-urdin eu people might start their day now :) please have a look at getting a release for the CVE out https://review.opendev.org/c/openstack/releases/+/871802
08:43:10 frickler gibi: bauzas: sean-k-mooney: ^^ I kind of agree with tobias-urdin that there is a bit of urgency behind this. not sure if release team could skip ptl approval in this case, though?
08:43:50 kashyap Is it just me or the "tempest-integrated-compute" is timing out often for others too?
08:44:22 bauzas frickler: I'm here
08:44:52 bauzas frickler: we said we were trying to also merge another CVE fix before we release this
08:50:37 bauzas ok, the other cve bug is fixed down to xena, so yeah, we can have releases
08:56:41 frickler that other CVE was https://review.opendev.org/c/openstack/nova/+/859315, right?
09:04:51 bauzas yup, just checked, reviewing now the releases patch
09:09:21 bauzas gibi: frickler: I was torn with the proposed semver of https://review.opendev.org/c/openstack/releases/+/871802 which was doing .y releases but I can live with that
09:10:18 bauzas tobias-urdin: ^ anything in particular you had in mind when you did set the releases for a .y bump ?
09:10:30 bauzas the fact that it was exposing a new conf knob, I guess ?
09:14:03 gibi I'm fine with the minor bump
09:15:05 bauzas that seems a bit agressive but thinking out more, that means that distros have to adapt their toolings if they wanna set the new conf knob
09:15:09 bauzas so yeah a .y bump seems ok
09:15:32 bauzas even if semantically, we're sending a wrong signal
09:17:16 gibi we are removing functionality with a knob by default so I'm OK to y bump it
09:17:53 gibi I double checked it seems both cve is in the release (the vnit_type on was in zed when it was master)
09:18:06 gibi so I think we are good to go
09:20:10 tobias-urdin i pretty much followed cinder that bumped minor, I guess it wouldn't hurt indicating to operations that a minor version that should be upgraded to because of CVE
09:20:38 tobias-urdin but yeah we can change if required, just wanted to make it a priority to release it so downstream can start building stuff new versions as well
09:20:48 bauzas gibi: yeah, checked the other CVE, was my main original driver for the check
09:21:07 bauzas tobias-urdin: no worries, as I said, I can live with that
09:21:24 gibi next is stable/wallaby but that needs the tempest pin first. https://review.opendev.org/q/topic:wallaby-pin-tempest
09:21:32 bauzas in theory a CVE fix doesn't require a y versioning but meh
09:21:46 tobias-urdin bauzas: ack, thanks, i will keep that in mind for the future
09:21:49 bauzas gibi: correct, I +2/+Wd a patch this morning
09:22:43 bauzas tobias-urdin: np, not anyone needs to know anything :) but if you wanna know more about semver, this is the reference page https://docs.openstack.org/pbr/latest/user/semver.html
09:23:10 bauzas gibi: do you know if gmann did the tempest patch ?
09:23:17 bauzas I can check, I just didn't had the time yet
09:23:32 gibi bauzas: here is the tempest pin series https://review.opendev.org/q/topic:wallaby-pin-tempest it needs love
09:23:44 bauzas I can surely provide love
09:23:45 gibi the DNM test patches are failing
09:24:33 bauzas then I guess the love has to be on finding why the DNM patches are failing
09:24:35 bauzas lovely
09:24:51 bauzas that's just 12 hours I haven't looked at zuul files
09:48:18 bauzas mmm
09:48:40 bauzas gibi: does those skipttest exceptions in tempest look correct to you ?
09:48:42 bauzas 2023-01-26 01:06:13.455899 | controller | unittest2.case.SkipTest: Identity api v2 is not enabled
09:48:51 bauzas I have seen gmann rechecking on such errors
09:49:06 bauzas https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_0a4/871782/2/check/tempest-full-py3/0a42465/job-output.txt
09:49:17 gibi it does not look good, but maybe gmann rechecked on it as he changed one of the depends-ons
09:50:31 bauzas I just asked for a recheck
09:50:44 bauzas on the devstack patch
09:51:03 bauzas anyway, the changes themselves on both tempest and devstack seem logic to me
09:51:04 gibi we will see
09:51:12 bauzas so I guess this is just a transient issue
09:52:22 bauzas at least keystone was running
09:53:48 sahid o/ sean-k-mooney, bauzas I have noticed your new comments
09:53:51 sahid working on it !
09:53:53 sahid thanks
09:54:33 bauzas sahid: thanks
09:54:47 bauzas sahid: ping me when you're done with those, and I'll rereview
09:55:13 sean-k-mooney[m] the rpc pin case is not something i orginally tought of but i dont think its a large change to just make sure we dont retrun a 500 from the api and add a test for that so hopefully it wont take too much to adress that
09:55:33 bauzas sahid: to clarify, sorry but you don't need to change the RPC client, just make sure that on the API you can verify it
09:56:26 sean-k-mooney[m] yep so add a functional test that pins to 6.0
09:56:32 sean-k-mooney[m] use the new microversion
09:56:46 sean-k-mooney[m] and assert an excption is raised
09:56:53 sean-k-mooney[m] ideally it should be a 409
09:57:06 sean-k-mooney[m] the same as when the compute is not upgraded
09:57:53 sean-k-mooney[m] you should not expose the crrent RPC pin in the excption
09:58:48 sean-k-mooney[m] just that the could does not meet the requirements for the new microversion and you should use the old one like the other exception
09:59:35 sean-k-mooney[m] bauzas actully is there any reason not to use the same excpetion here as in the api when the compute service is to low
10:00:04 bauzas good question
10:00:29 bauzas problem is, evacuation is defined by policy, right?
10:00:52 sean-k-mooney[m] i guess its and admin api and they might want to know but if they manually pinned. perhapse its diffent admins that do upgrade vs day to day
10:01:04 sean-k-mooney[m] well all apis are controlable by policy
10:01:06 bauzas so you can change the policy to have evacuation (without a host param) be supported for endusers
10:01:28 bauzas if so, we could return something like 'sorry, compute service is too low'
10:01:31 sean-k-mooney[m] you could
10:01:45 bauzas I don't know whether it would be a problem for our operators then
10:01:56 kashyap This timeout is blocking a couple of patches. /me is trying to find where exactly is the time_out -- https://zuul.opendev.org/t/openstack/build/453f991eeeb34498b132eb84de3301db/logs
10:02:43 sean-k-mooney[m] sound like just a slow node to be honest
10:03:05 sean-k-mooney[m] and that job might be hitting up agagisnt the timeout anyway but lests see what the normal runtime is
10:03:25 kashyap So only a full recheck is the only option? :(
10:03:33 kashyap (Already did it once)
10:03:37 sean-k-mooney[m] normally 90 mins or so
10:03:51 kashyap sean-k-mooney[m]: For the full recheck?
10:04:45 sean-k-mooney[m] that job normally takes 90 mins
10:04:46 sean-k-mooney[m] https://zuul.opendev.org/t/openstack/builds?job_name=tempest-integrated-compute&project=openstack/nova
10:05:26 sean-k-mooney[m] there have been 5 time outs in the last 300 runs of that job
10:10:07 sean-k-mooney[m] the 3 time outs are form 4 providers so i dont really see any corralation
10:10:33 sean-k-mooney[m] it would be good to see if there is anything odd in the devstack or tempet runs
10:10:49 sean-k-mooney[m] but something took more time then normal
10:11:31 sean-k-mooney[m] so yes a recheck is the way to proceed but it would be good to see if say a lot of swap was used or an image/pacakge download was slow

Earlier   Later