Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-26
09:54:33 bauzas sahid: thanks
09:54:47 bauzas sahid: ping me when you're done with those, and I'll rereview
09:55:13 sean-k-mooney[m] the rpc pin case is not something i orginally tought of but i dont think its a large change to just make sure we dont retrun a 500 from the api and add a test for that so hopefully it wont take too much to adress that
09:55:33 bauzas sahid: to clarify, sorry but you don't need to change the RPC client, just make sure that on the API you can verify it
09:56:26 sean-k-mooney[m] yep so add a functional test that pins to 6.0
09:56:32 sean-k-mooney[m] use the new microversion
09:56:46 sean-k-mooney[m] and assert an excption is raised
09:56:53 sean-k-mooney[m] ideally it should be a 409
09:57:06 sean-k-mooney[m] the same as when the compute is not upgraded
09:57:53 sean-k-mooney[m] you should not expose the crrent RPC pin in the excption
09:58:48 sean-k-mooney[m] just that the could does not meet the requirements for the new microversion and you should use the old one like the other exception
09:59:35 sean-k-mooney[m] bauzas actully is there any reason not to use the same excpetion here as in the api when the compute service is to low
10:00:04 bauzas good question
10:00:29 bauzas problem is, evacuation is defined by policy, right?
10:00:52 sean-k-mooney[m] i guess its and admin api and they might want to know but if they manually pinned. perhapse its diffent admins that do upgrade vs day to day
10:01:04 sean-k-mooney[m] well all apis are controlable by policy
10:01:06 bauzas so you can change the policy to have evacuation (without a host param) be supported for endusers
10:01:28 bauzas if so, we could return something like 'sorry, compute service is too low'
10:01:31 sean-k-mooney[m] you could
10:01:45 bauzas I don't know whether it would be a problem for our operators then
10:01:56 kashyap This timeout is blocking a couple of patches. /me is trying to find where exactly is the time_out -- https://zuul.opendev.org/t/openstack/build/453f991eeeb34498b132eb84de3301db/logs
10:02:43 sean-k-mooney[m] sound like just a slow node to be honest
10:03:05 sean-k-mooney[m] and that job might be hitting up agagisnt the timeout anyway but lests see what the normal runtime is
10:03:25 kashyap So only a full recheck is the only option? :(
10:03:33 kashyap (Already did it once)
10:03:37 sean-k-mooney[m] normally 90 mins or so
10:03:51 kashyap sean-k-mooney[m]: For the full recheck?
10:04:45 sean-k-mooney[m] that job normally takes 90 mins
10:04:46 sean-k-mooney[m] https://zuul.opendev.org/t/openstack/builds?job_name=tempest-integrated-compute&project=openstack/nova
10:05:26 sean-k-mooney[m] there have been 5 time outs in the last 300 runs of that job
10:10:07 sean-k-mooney[m] the 3 time outs are form 4 providers so i dont really see any corralation
10:10:33 sean-k-mooney[m] it would be good to see if there is anything odd in the devstack or tempet runs
10:10:49 sean-k-mooney[m] but something took more time then normal
10:11:31 sean-k-mooney[m] so yes a recheck is the way to proceed but it would be good to see if say a lot of swap was used or an image/pacakge download was slow
10:18:40 sean-k-mooney[m] looking at a passing run vs failing devstack too ~900 seconds vs ~1200
10:19:29 sean-k-mooney[m] its seams to be pretty even across apt install pip install and osc
10:19:58 sean-k-mooney[m] so i think this is diskio or just general cpu/disk/net performace related
10:20:27 sean-k-mooney[m] if the devstack install is 33% slower the tempest execation will likely be similarly reduced in perfromace
10:20:49 sean-k-mooney[m] and a normal run is at 75% of the build limit so it cant really tollerate that much of a slow down
10:26:33 sahid bauzas, sean-k-mooney[m] humm i may missing somethinhg, this is not what you are looking for? https://review.opendev.org/c/openstack/nova/+/858384/34/nova/tests/unit/api/openstack/compute/test_evacuate.py#431
10:27:16 sahid that is when using microversion 2.95 with a hosts that are not fully upgraded
10:27:22 sean-k-mooney[m] no
10:27:39 sean-k-mooney[m] that is pininng the min compute service version
10:27:57 sean-k-mooney[m] we were talking about the rpc version pin
10:28:47 sean-k-mooney[m] the ones you pin in https://docs.openstack.org/nova/latest/configuration/config.html#upgrade-levels
10:29:26 sean-k-mooney[m] so [upgrade_levels]/compute=6.0
10:30:41 bauzas sahid: the missing unittest I was referring was to verify that you return an exception if you set the parameter and call a old compute
10:31:09 bauzas sahid: for the functest in the other change (the microversion one), yeah, what sean-k-mooney said
10:32:41 sahid bauzas: this part? https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/manager.py#3831
10:33:13 bauzas sahid: no sorry
10:33:15 bauzas sec
10:33:53 bauzas sahid: in https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/rpcapi.py#1108
10:33:59 bauzas sahid: you return an exception
10:34:10 bauzas and I'm eventually OK with it (sorry for the comments)
10:34:37 bauzas sahid: butn
10:36:12 bauzas sahid: I don't see any unittests for verifying the RPC call
10:36:41 sahid yes good point
10:37:20 bauzas and we also need to have unittests for the conductor RPC API, unfortunately :(
10:37:37 bauzas sahid: sec, will find you where we have unittests for both
10:42:52 bauzas sahid: one example of a simple unittest for testing it https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4734-L4752
10:43:13 bauzas gosh, our tests are so horrible to read :(
10:43:32 bauzas https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4754-L4788 are also good examples
10:43:43 bauzas just add two checks there and I'm all good (for the conductor API)
10:43:56 bauzas (even if we apparently missed the other version checks...)
10:44:14 sahid working on it! thank for you help guys
10:45:32 bauzas sahid: for the compute RPC API, here is the existing test https://github.com/openstack/nova/blob/master/nova/tests/unit/compute/test_rpcapi.py#L883-L916
10:45:40 bauzas HTH
10:45:58 bauzas sahid: np, you're next in my review queue once you're all set
10:51:41 kashyap sean-k-mooney[m]: Sorry for my delay, just reading back the scroll of your analysis
10:54:39 sean-k-mooney tl;dr the io on those nodes that timeed out looks like its about 33% slower then normal. we are normally at 75% of the timeout so we dont have the headroom in that case
11:09:18 sean-k-mooney gibi: bauzas can we land gmann's placement service role patch https://review.opendev.org/c/openstack/placement/+/865618
12:41:38 opendevreview Merged openstack/nova master: Handle InstanceInvalidState exception https://review.opendev.org/c/openstack/nova/+/861738
13:15:20 kashyap sean-k-mooney: Thanks for the summary. It's nearly 2-ish hours, and still no sign of a vote - https://review.opendev.org/c/openstack/nova/+/870794
13:16:39 sean-k-mooney its in gate
13:16:49 sean-k-mooney its been runing for 1hr 46mins
13:17:05 sean-k-mooney and tempest-integrated-compute passed already
13:17:15 sean-k-mooney you can watch it here if you like https://zuul.openstack.org/status#870794
13:21:55 kashyap sean-k-mooney: Thank you :)
13:33:22 ierdem Hi, when I try to cold migrate VMs via cli by specifying destination host, it throws an exception after first migrate "No valid host was found" -no more details, just this message-, but destination host has enough resource. Does nova-scheduler cause this? If is, how can I force it to migrate more than one VMs to the same host? Thanks for all your assistance. (I have kolla-ansible stein-eol)
13:35:22 bauzas gmann: sean-k-mooney: dansmith: sorry, but maybe I'm confused but I thought system readers can scope resources of a deployment. If so, why are we returning HTTP403 on https://review.opendev.org/c/openstack/placement/+/865618 ?
13:36:38 bauzas (Alice in the spec, obviously)
13:37:41 kashyap (Third time lucky, the job indeed passed)
13:38:09 kashyap (We'll see if everythig else succeeds.)
13:38:39 bauzas oh wait
13:38:45 bauzas "Alice can list or retrieve specific endpoints. Alice cannot do any project specific operations since her authorization is limited to the deployment system."
13:38:54 bauzas including list, IIRC
13:43:35 sean-k-mooney bauzas: we are remvoing system reader form the placment polices
13:44:40 sean-k-mooney for most project we are geting rid fo scope entirly based on operator feedback in yoga?
13:46:10 bauzas but okay
13:46:18 bauzas I can understand we restrict the roles then
13:46:49 bauzas and which can be updated as long as we implement the features
13:48:04 opendevreview Merged openstack/nova master: libvirt: At start-up rework compareCPU() usage with a workaround https://review.opendev.org/c/openstack/nova/+/870794
13:50:05 kashyap Christ, it merged at last
13:53:05 gibi kashyap: it wasn't that long:)
13:53:40 kashyap gibi: You're right though. I tend to be a bit of a drama queen sometimes; please ignore me :)
13:54:14 kashyap gibi: Thank you both for the help! Now hope the API replacement patch goes through soon
13:54:21 bauzas sean-k-mooney: any docs you could point me out on the new direction for system roles ?
13:54:25 kashyap (Thanks, sahid for rechecking it while I was away :))
13:56:08 sean-k-mooney bauzas: gmann might have the links more redilaly but i think we updated the goal with it too
13:56:22 sean-k-mooney bauzas: like this was a very very big thing that we have talked about before

Earlier   Later