Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-26
10:01:56 kashyap This timeout is blocking a couple of patches. /me is trying to find where exactly is the time_out -- https://zuul.opendev.org/t/openstack/build/453f991eeeb34498b132eb84de3301db/logs
10:02:43 sean-k-mooney[m] sound like just a slow node to be honest
10:03:05 sean-k-mooney[m] and that job might be hitting up agagisnt the timeout anyway but lests see what the normal runtime is
10:03:25 kashyap So only a full recheck is the only option? :(
10:03:33 kashyap (Already did it once)
10:03:37 sean-k-mooney[m] normally 90 mins or so
10:03:51 kashyap sean-k-mooney[m]: For the full recheck?
10:04:45 sean-k-mooney[m] that job normally takes 90 mins
10:04:46 sean-k-mooney[m] https://zuul.opendev.org/t/openstack/builds?job_name=tempest-integrated-compute&project=openstack/nova
10:05:26 sean-k-mooney[m] there have been 5 time outs in the last 300 runs of that job
10:10:07 sean-k-mooney[m] the 3 time outs are form 4 providers so i dont really see any corralation
10:10:33 sean-k-mooney[m] it would be good to see if there is anything odd in the devstack or tempet runs
10:10:49 sean-k-mooney[m] but something took more time then normal
10:11:31 sean-k-mooney[m] so yes a recheck is the way to proceed but it would be good to see if say a lot of swap was used or an image/pacakge download was slow
10:18:40 sean-k-mooney[m] looking at a passing run vs failing devstack too ~900 seconds vs ~1200
10:19:29 sean-k-mooney[m] its seams to be pretty even across apt install pip install and osc
10:19:58 sean-k-mooney[m] so i think this is diskio or just general cpu/disk/net performace related
10:20:27 sean-k-mooney[m] if the devstack install is 33% slower the tempest execation will likely be similarly reduced in perfromace
10:20:49 sean-k-mooney[m] and a normal run is at 75% of the build limit so it cant really tollerate that much of a slow down
10:26:33 sahid bauzas, sean-k-mooney[m] humm i may missing somethinhg, this is not what you are looking for? https://review.opendev.org/c/openstack/nova/+/858384/34/nova/tests/unit/api/openstack/compute/test_evacuate.py#431
10:27:16 sahid that is when using microversion 2.95 with a hosts that are not fully upgraded
10:27:22 sean-k-mooney[m] no
10:27:39 sean-k-mooney[m] that is pininng the min compute service version
10:27:57 sean-k-mooney[m] we were talking about the rpc version pin
10:28:47 sean-k-mooney[m] the ones you pin in https://docs.openstack.org/nova/latest/configuration/config.html#upgrade-levels
10:29:26 sean-k-mooney[m] so [upgrade_levels]/compute=6.0
10:30:41 bauzas sahid: the missing unittest I was referring was to verify that you return an exception if you set the parameter and call a old compute
10:31:09 bauzas sahid: for the functest in the other change (the microversion one), yeah, what sean-k-mooney said
10:32:41 sahid bauzas: this part? https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/manager.py#3831
10:33:13 bauzas sahid: no sorry
10:33:15 bauzas sec
10:33:53 bauzas sahid: in https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/rpcapi.py#1108
10:33:59 bauzas sahid: you return an exception
10:34:10 bauzas and I'm eventually OK with it (sorry for the comments)
10:34:37 bauzas sahid: butn
10:36:12 bauzas sahid: I don't see any unittests for verifying the RPC call
10:36:41 sahid yes good point
10:37:20 bauzas and we also need to have unittests for the conductor RPC API, unfortunately :(
10:37:37 bauzas sahid: sec, will find you where we have unittests for both
10:42:52 bauzas sahid: one example of a simple unittest for testing it https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4734-L4752
10:43:13 bauzas gosh, our tests are so horrible to read :(
10:43:32 bauzas https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4754-L4788 are also good examples
10:43:43 bauzas just add two checks there and I'm all good (for the conductor API)
10:43:56 bauzas (even if we apparently missed the other version checks...)
10:44:14 sahid working on it! thank for you help guys
10:45:32 bauzas sahid: for the compute RPC API, here is the existing test https://github.com/openstack/nova/blob/master/nova/tests/unit/compute/test_rpcapi.py#L883-L916
10:45:40 bauzas HTH
10:45:58 bauzas sahid: np, you're next in my review queue once you're all set
10:51:41 kashyap sean-k-mooney[m]: Sorry for my delay, just reading back the scroll of your analysis
10:54:39 sean-k-mooney tl;dr the io on those nodes that timeed out looks like its about 33% slower then normal. we are normally at 75% of the timeout so we dont have the headroom in that case
11:09:18 sean-k-mooney gibi: bauzas can we land gmann's placement service role patch https://review.opendev.org/c/openstack/placement/+/865618
12:41:38 opendevreview Merged openstack/nova master: Handle InstanceInvalidState exception https://review.opendev.org/c/openstack/nova/+/861738
13:15:20 kashyap sean-k-mooney: Thanks for the summary. It's nearly 2-ish hours, and still no sign of a vote - https://review.opendev.org/c/openstack/nova/+/870794
13:16:39 sean-k-mooney its in gate
13:16:49 sean-k-mooney its been runing for 1hr 46mins
13:17:05 sean-k-mooney and tempest-integrated-compute passed already
13:17:15 sean-k-mooney you can watch it here if you like https://zuul.openstack.org/status#870794
13:21:55 kashyap sean-k-mooney: Thank you :)
13:33:22 ierdem Hi, when I try to cold migrate VMs via cli by specifying destination host, it throws an exception after first migrate "No valid host was found" -no more details, just this message-, but destination host has enough resource. Does nova-scheduler cause this? If is, how can I force it to migrate more than one VMs to the same host? Thanks for all your assistance. (I have kolla-ansible stein-eol)
13:35:22 bauzas gmann: sean-k-mooney: dansmith: sorry, but maybe I'm confused but I thought system readers can scope resources of a deployment. If so, why are we returning HTTP403 on https://review.opendev.org/c/openstack/placement/+/865618 ?
13:36:38 bauzas (Alice in the spec, obviously)
13:37:41 kashyap (Third time lucky, the job indeed passed)
13:38:09 kashyap (We'll see if everythig else succeeds.)
13:38:39 bauzas oh wait
13:38:45 bauzas "Alice can list or retrieve specific endpoints. Alice cannot do any project specific operations since her authorization is limited to the deployment system."
13:38:54 bauzas including list, IIRC
13:43:35 sean-k-mooney bauzas: we are remvoing system reader form the placment polices
13:44:40 sean-k-mooney for most project we are geting rid fo scope entirly based on operator feedback in yoga?
13:46:10 bauzas but okay
13:46:18 bauzas I can understand we restrict the roles then
13:46:49 bauzas and which can be updated as long as we implement the features
13:48:04 opendevreview Merged openstack/nova master: libvirt: At start-up rework compareCPU() usage with a workaround https://review.opendev.org/c/openstack/nova/+/870794
13:50:05 kashyap Christ, it merged at last
13:53:05 gibi kashyap: it wasn't that long:)
13:53:40 kashyap gibi: You're right though. I tend to be a bit of a drama queen sometimes; please ignore me :)
13:54:14 kashyap gibi: Thank you both for the help! Now hope the API replacement patch goes through soon
13:54:21 bauzas sean-k-mooney: any docs you could point me out on the new direction for system roles ?
13:54:25 kashyap (Thanks, sahid for rechecking it while I was away :))
13:56:08 sean-k-mooney bauzas: gmann might have the links more redilaly but i think we updated the goal with it too
13:56:22 sean-k-mooney bauzas: like this was a very very big thing that we have talked about before
13:56:48 sean-k-mooney https://github.com/openstack/governance/commit/1909d4f7a0dc2920fc04ab5bfac112a671547cee
13:56:54 bauzas " In yoga cycle, we redefined this goal with the changes mentioned above so that allowing system administrators to access system level resources APIs only and allow project users to access project-level resource APIs. These changes have been done for nova and neutron. "
13:57:06 bauzas sean-k-mooney: problem is, sometimes my brain splits
13:57:14 sean-k-mooney it was actully in zed
13:57:43 bauzas so it looks to me I probably heard about it, but then it wasn't ringing the bell in my brain when I saw today's change
13:57:45 sean-k-mooney well it depnes but that commit has the relevent info in it
14:01:53 sahid kashyap: sure :-)
14:10:33 sean-k-mooney kashyap: now all you have to do is backport it all the way to train and then cerry pick it downstream :P
14:11:20 sean-k-mooney so for 2 patchs that only 20 ish commits to create and review
14:11:56 sean-k-mooney with that said im not sure we need to backport the second patch
14:12:09 sean-k-mooney the frist we should but using the new api is less imporant
14:18:50 opendevreview Alexey Stupnikov proposed openstack/nova stable/zed: Remove deleted projects from flavor access list https://review.opendev.org/c/openstack/nova/+/870053
14:26:13 opendevreview Sahid Orentino Ferdjaoui proposed openstack/nova master: compute: enhance compute evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858383
14:26:14 opendevreview Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384
14:27:30 sahid ^ fixed compute part, and owrking on the API one
14:29:59 sahid for the API part, is a BadRequest returned sounds good? in case that the rpc is not right and user is using new version?
14:30:17 sahid or do you have in mind an other http status?
14:30:42 sahid sean-k-mooney, bauzas ^ if you have a sec
14:31:23 bauzas sahid: we agreed yesterday on returning a HTTP409 Conflict
14:31:29 sahid ack

Earlier   Later