| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-26 | |||
| 09:56:53 | sean-k-mooney[m] | ideally it should be a 409 | |
| 09:57:06 | sean-k-mooney[m] | the same as when the compute is not upgraded | |
| 09:57:53 | sean-k-mooney[m] | you should not expose the crrent RPC pin in the excption | |
| 09:58:48 | sean-k-mooney[m] | just that the could does not meet the requirements for the new microversion and you should use the old one like the other exception | |
| 09:59:35 | sean-k-mooney[m] | bauzas actully is there any reason not to use the same excpetion here as in the api when the compute service is to low | |
| 10:00:04 | bauzas | good question | |
| 10:00:29 | bauzas | problem is, evacuation is defined by policy, right? | |
| 10:00:52 | sean-k-mooney[m] | i guess its and admin api and they might want to know but if they manually pinned. perhapse its diffent admins that do upgrade vs day to day | |
| 10:01:04 | sean-k-mooney[m] | well all apis are controlable by policy | |
| 10:01:06 | bauzas | so you can change the policy to have evacuation (without a host param) be supported for endusers | |
| 10:01:28 | bauzas | if so, we could return something like 'sorry, compute service is too low' | |
| 10:01:31 | sean-k-mooney[m] | you could | |
| 10:01:45 | bauzas | I don't know whether it would be a problem for our operators then | |
| 10:01:56 | kashyap | This timeout is blocking a couple of patches. /me is trying to find where exactly is the time_out -- https://zuul.opendev.org/t/openstack/build/453f991eeeb34498b132eb84de3301db/logs | |
| 10:02:43 | sean-k-mooney[m] | sound like just a slow node to be honest | |
| 10:03:05 | sean-k-mooney[m] | and that job might be hitting up agagisnt the timeout anyway but lests see what the normal runtime is | |
| 10:03:25 | kashyap | So only a full recheck is the only option? :( | |
| 10:03:33 | kashyap | (Already did it once) | |
| 10:03:37 | sean-k-mooney[m] | normally 90 mins or so | |
| 10:03:51 | kashyap | sean-k-mooney[m]: For the full recheck? | |
| 10:04:45 | sean-k-mooney[m] | that job normally takes 90 mins | |
| 10:04:46 | sean-k-mooney[m] | https://zuul.opendev.org/t/openstack/builds?job_name=tempest-integrated-compute&project=openstack/nova | |
| 10:05:26 | sean-k-mooney[m] | there have been 5 time outs in the last 300 runs of that job | |
| 10:10:07 | sean-k-mooney[m] | the 3 time outs are form 4 providers so i dont really see any corralation | |
| 10:10:33 | sean-k-mooney[m] | it would be good to see if there is anything odd in the devstack or tempet runs | |
| 10:10:49 | sean-k-mooney[m] | but something took more time then normal | |
| 10:11:31 | sean-k-mooney[m] | so yes a recheck is the way to proceed but it would be good to see if say a lot of swap was used or an image/pacakge download was slow | |
| 10:18:40 | sean-k-mooney[m] | looking at a passing run vs failing devstack too ~900 seconds vs ~1200 | |
| 10:19:29 | sean-k-mooney[m] | its seams to be pretty even across apt install pip install and osc | |
| 10:19:58 | sean-k-mooney[m] | so i think this is diskio or just general cpu/disk/net performace related | |
| 10:20:27 | sean-k-mooney[m] | if the devstack install is 33% slower the tempest execation will likely be similarly reduced in perfromace | |
| 10:20:49 | sean-k-mooney[m] | and a normal run is at 75% of the build limit so it cant really tollerate that much of a slow down | |
| 10:26:33 | sahid | bauzas, sean-k-mooney[m] humm i may missing somethinhg, this is not what you are looking for? https://review.opendev.org/c/openstack/nova/+/858384/34/nova/tests/unit/api/openstack/compute/test_evacuate.py#431 | |
| 10:27:16 | sahid | that is when using microversion 2.95 with a hosts that are not fully upgraded | |
| 10:27:22 | sean-k-mooney[m] | no | |
| 10:27:39 | sean-k-mooney[m] | that is pininng the min compute service version | |
| 10:27:57 | sean-k-mooney[m] | we were talking about the rpc version pin | |
| 10:28:47 | sean-k-mooney[m] | the ones you pin in https://docs.openstack.org/nova/latest/configuration/config.html#upgrade-levels | |
| 10:29:26 | sean-k-mooney[m] | so [upgrade_levels]/compute=6.0 | |
| 10:30:41 | bauzas | sahid: the missing unittest I was referring was to verify that you return an exception if you set the parameter and call a old compute | |
| 10:31:09 | bauzas | sahid: for the functest in the other change (the microversion one), yeah, what sean-k-mooney said | |
| 10:32:41 | sahid | bauzas: this part? https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/manager.py#3831 | |
| 10:33:13 | bauzas | sahid: no sorry | |
| 10:33:15 | bauzas | sec | |
| 10:33:53 | bauzas | sahid: in https://review.opendev.org/c/openstack/nova/+/858383/25/nova/compute/rpcapi.py#1108 | |
| 10:33:59 | bauzas | sahid: you return an exception | |
| 10:34:10 | bauzas | and I'm eventually OK with it (sorry for the comments) | |
| 10:34:37 | bauzas | sahid: butn | |
| 10:36:12 | bauzas | sahid: I don't see any unittests for verifying the RPC call | |
| 10:36:41 | sahid | yes good point | |
| 10:37:20 | bauzas | and we also need to have unittests for the conductor RPC API, unfortunately :( | |
| 10:37:37 | bauzas | sahid: sec, will find you where we have unittests for both | |
| 10:42:52 | bauzas | sahid: one example of a simple unittest for testing it https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4734-L4752 | |
| 10:43:13 | bauzas | gosh, our tests are so horrible to read :( | |
| 10:43:32 | bauzas | https://github.com/openstack/nova/blob/master/nova/tests/unit/conductor/test_conductor.py#L4754-L4788 are also good examples | |
| 10:43:43 | bauzas | just add two checks there and I'm all good (for the conductor API) | |
| 10:43:56 | bauzas | (even if we apparently missed the other version checks...) | |
| 10:44:14 | sahid | working on it! thank for you help guys | |
| 10:45:32 | bauzas | sahid: for the compute RPC API, here is the existing test https://github.com/openstack/nova/blob/master/nova/tests/unit/compute/test_rpcapi.py#L883-L916 | |
| 10:45:40 | bauzas | HTH | |
| 10:45:58 | bauzas | sahid: np, you're next in my review queue once you're all set | |
| 10:51:41 | kashyap | sean-k-mooney[m]: Sorry for my delay, just reading back the scroll of your analysis | |
| 10:54:39 | sean-k-mooney | tl;dr the io on those nodes that timeed out looks like its about 33% slower then normal. we are normally at 75% of the timeout so we dont have the headroom in that case | |
| 11:09:18 | sean-k-mooney | gibi: bauzas can we land gmann's placement service role patch https://review.opendev.org/c/openstack/placement/+/865618 | |
| 12:41:38 | opendevreview | Merged openstack/nova master: Handle InstanceInvalidState exception https://review.opendev.org/c/openstack/nova/+/861738 | |
| 13:15:20 | kashyap | sean-k-mooney: Thanks for the summary. It's nearly 2-ish hours, and still no sign of a vote - https://review.opendev.org/c/openstack/nova/+/870794 | |
| 13:16:39 | sean-k-mooney | its in gate | |
| 13:16:49 | sean-k-mooney | its been runing for 1hr 46mins | |
| 13:17:05 | sean-k-mooney | and tempest-integrated-compute passed already | |
| 13:17:15 | sean-k-mooney | you can watch it here if you like https://zuul.openstack.org/status#870794 | |
| 13:21:55 | kashyap | sean-k-mooney: Thank you :) | |
| 13:33:22 | ierdem | Hi, when I try to cold migrate VMs via cli by specifying destination host, it throws an exception after first migrate "No valid host was found" -no more details, just this message-, but destination host has enough resource. Does nova-scheduler cause this? If is, how can I force it to migrate more than one VMs to the same host? Thanks for all your assistance. (I have kolla-ansible stein-eol) | |
| 13:35:22 | bauzas | gmann: sean-k-mooney: dansmith: sorry, but maybe I'm confused but I thought system readers can scope resources of a deployment. If so, why are we returning HTTP403 on https://review.opendev.org/c/openstack/placement/+/865618 ? | |
| 13:36:38 | bauzas | (Alice in the spec, obviously) | |
| 13:37:41 | kashyap | (Third time lucky, the job indeed passed) | |
| 13:38:09 | kashyap | (We'll see if everythig else succeeds.) | |
| 13:38:39 | bauzas | oh wait | |
| 13:38:45 | bauzas | "Alice can list or retrieve specific endpoints. Alice cannot do any project specific operations since her authorization is limited to the deployment system." | |
| 13:38:54 | bauzas | including list, IIRC | |
| 13:43:35 | sean-k-mooney | bauzas: we are remvoing system reader form the placment polices | |
| 13:44:40 | sean-k-mooney | for most project we are geting rid fo scope entirly based on operator feedback in yoga? | |
| 13:46:10 | bauzas | but okay | |
| 13:46:18 | bauzas | I can understand we restrict the roles then | |
| 13:46:49 | bauzas | and which can be updated as long as we implement the features | |
| 13:48:04 | opendevreview | Merged openstack/nova master: libvirt: At start-up rework compareCPU() usage with a workaround https://review.opendev.org/c/openstack/nova/+/870794 | |
| 13:50:05 | kashyap | Christ, it merged at last | |
| 13:53:05 | gibi | kashyap: it wasn't that long:) | |
| 13:53:40 | kashyap | gibi: You're right though. I tend to be a bit of a drama queen sometimes; please ignore me :) | |
| 13:54:14 | kashyap | gibi: Thank you both for the help! Now hope the API replacement patch goes through soon | |
| 13:54:21 | bauzas | sean-k-mooney: any docs you could point me out on the new direction for system roles ? | |
| 13:54:25 | kashyap | (Thanks, sahid for rechecking it while I was away :)) | |
| 13:56:08 | sean-k-mooney | bauzas: gmann might have the links more redilaly but i think we updated the goal with it too | |
| 13:56:22 | sean-k-mooney | bauzas: like this was a very very big thing that we have talked about before | |
| 13:56:48 | sean-k-mooney | https://github.com/openstack/governance/commit/1909d4f7a0dc2920fc04ab5bfac112a671547cee | |
| 13:56:54 | bauzas | " In yoga cycle, we redefined this goal with the changes mentioned above so that allowing system administrators to access system level resources APIs only and allow project users to access project-level resource APIs. These changes have been done for nova and neutron. " | |
| 13:57:06 | bauzas | sean-k-mooney: problem is, sometimes my brain splits | |
| 13:57:14 | sean-k-mooney | it was actully in zed | |
| 13:57:43 | bauzas | so it looks to me I probably heard about it, but then it wasn't ringing the bell in my brain when I saw today's change | |
| 13:57:45 | sean-k-mooney | well it depnes but that commit has the relevent info in it | |
| 14:01:53 | sahid | kashyap: sure :-) | |