| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-12-15 | |||
| 12:33:09 | lyarwood | gibi: n-cpu deletes the volume attachment and that eventually leads c-vol to mark the volume as available again | |
| 12:33:32 | lyarwood | gibi: I just need to trace the cinder side through to confirm the race but n-cpu hasn't logged the call to delete the volume attachment for some reason | |
| 12:33:37 | gibi | ohh, so we would neet to wait for the volume become available before we call delete on it | |
| 12:33:42 | lyarwood | yeah indeed | |
| 12:33:54 | lyarwood | that's awkward but that's life with async APIs | |
| 12:34:06 | gibi | yeah, I think that is OK to add as a fix | |
| 12:36:19 | gibi | lyarwood: so the tearDownClass() needs to be smarter or we need to extend a specific test case? | |
| 12:37:31 | lyarwood | oh wait, actually it looks like n-cpu didn't delete the attachment | |
| 12:45:51 | teoobo_ | stephenfin: thanks for the code review | |
| 12:55:16 | lyarwood | gibi: ah ha! it's actually slow snapshot creation in c-vol that's causing this | |
| 12:55:20 | lyarwood | gibi: that's also async | |
| 12:55:43 | lyarwood | gibi: there's a TODO from Matt in the test that we should sort out to handle this, adding a waiter on the snapshot state etc | |
| 12:56:04 | gibi | lyarwood: I admire you that you was able to found this, I alway lost in the cinder logs | |
| 12:58:51 | gibi | is there a way I can help fixing this in tempest? | |
| 13:00:56 | gibi | I mean is it easy (for somebody like me without much tempest knowledge) to fix in tempest? | |
| 13:04:11 | lyarwood | gibi: haha thanks, I didn't have to sacrifice anything or anyone this time ;) | |
| 13:04:22 | lyarwood | gibi: so I think we just need to wait until the volume snapshot is ready | |
| 13:04:42 | gibi | I can hack on that if you have better things to do | |
| 13:04:45 | lyarwood | gibi: the issue here is that we are using the Nova imageCreate API that creates the volume snapshot for us indirectly | |
| 13:04:57 | lyarwood | gibi: no I can write this up and fix it | |
| 13:05:03 | gibi | cool, thanks | |
| 13:05:15 | lyarwood | gibi: just need to extract the volume snapshot details from the image metadata | |
| 13:08:03 | gibi | so it is not _that_ simple :) | |
| 13:10:21 | lyarwood | https://bugs.launchpad.net/tempest/+bug/1908269 | |
| 13:10:21 | openstack | Launchpad bug 1908269 in tempest "test_snapshot_volume_backed_multiattach not waiting for volume snapshot to become available" [Undecided,New] | |
| 13:10:26 | lyarwood | haha no nothing to do with bdms is | |
| 13:25:15 | gibi | lyarwood: thanks for the bugreport | |
| 13:28:31 | lyarwood | gibi: np, just rebuilding a devstack env, the fix should be trivial | |
| 13:28:44 | lyarwood | famous last words and all | |
| 13:31:19 | gibi | :) | |
| 13:57:05 | gmann | gibi: bauzas checking | |
| 14:01:43 | openstackgerrit | Merged openstack/nova master: Remove outdated comment from tox.ini https://review.opendev.org/c/openstack/nova/+/765534 | |
| 14:10:12 | gmann | gibi: bauzas on policy stuff in placement, I am waiting for the test case to be added and then we can start review. anyways I will check and leave the comment there. | |
| 14:10:29 | bauzas | all cool | |
| 14:11:44 | gibi | gmann: thanks for following that | |
| 14:12:17 | gmann | gibi: I will update the l-c fix comments after QA office hour. | |
| 14:12:26 | gibi | gmann: cool | |
| 14:14:01 | sean-k-mooney | gibi: did you see my ping yesterday | |
| 14:14:55 | sean-k-mooney | oh you did | |
| 14:15:08 | sean-k-mooney | it still might be worht a try setting the timeout | |
| 14:26:48 | gibi | sean-k-mooney: you want get the test to time out if it slow? or you want to prevent the test to time out when it is slow? | |
| 14:28:47 | sean-k-mooney | manila just set a 5 minute timeout to workaround slow nodes | |
| 14:28:59 | sean-k-mooney | vs the default which i think is 60 seconds | |
| 14:30:25 | gibi | sean-k-mooney: we don't have per test timeout I saw these test run and pass after 300 seconds on the gate | |
| 14:30:57 | sean-k-mooney | thats what the manila patch added | |
| 14:31:17 | gibi | hm, I saw these in the nova test suite | |
| 14:31:37 | sean-k-mooney | https://review.opendev.org/c/openstack/manila/+/291397/ | |
| 14:33:27 | gibi | sean-k-mooney: yeah, I saw this patch yesterday, what I say is that in nova we don't have a per test case timeout, and also I see that in nova these tests run for a long time (both in case of failing or passing) | |
| 14:36:00 | gibi | so I don't see how the extra timeout for these tests would help | |
| 14:37:02 | gibi | logstash don't want to help me know to show some examples | |
| 14:37:23 | gibi | I only found one: TestNovaAPIMigrationsWalkMySQL.test_walk_versions [192.038984s] ... FAILED | |
| 14:37:39 | gibi | TestNovaAPIMigrationsWalkPostgreSQL.test_walk_versions [43.027941s] | |
| 14:41:52 | sean-k-mooney | ah ok so its likely not a timeout issue | |
| 14:42:32 | sean-k-mooney | i did see a timeout in one of the logs | |
| 14:42:39 | sean-k-mooney | but i asumed that was a test timeout | |
| 14:43:30 | gibi | sean-k-mooney: link me the timeout case then I can double check that | |
| 14:44:22 | gibi | I see them failing after random amont of time. the only time sensitvity I can see is that when it fails then the execution is slower then when it passes | |
| 14:44:30 | gibi | but I saw slow passing executions too | |
| 14:44:44 | gibi | I just haven't seen a fast failing execution | |
| 14:54:24 | openstackgerrit | Ghanshyam proposed openstack/nova stable/rocky: DNM: testing gate https://review.opendev.org/c/openstack/nova/+/767027 | |
| 14:58:46 | lyarwood | gibi: https://review.opendev.org/c/openstack/tempest/+/767165 btw, should resolve this. | |
| 14:58:52 | lyarwood | gibi: did you write an ER query for this btw? | |
| 14:59:23 | openstackgerrit | Ghanshyam proposed openstack/placement master: Fix l-c job and move to latest hacking 4.0.0 https://review.opendev.org/c/openstack/placement/+/766994 | |
| 14:59:27 | gibi | lyarwood: I haven't as I was not able to find the root cause | |
| 14:59:38 | lyarwood | k np | |
| 14:59:38 | gibi | lyarwood: and I agree that as we have the fix the ER query is less important now | |
| 15:01:33 | gibi | lyarwood: thanks for the tempest fix | |
| 15:01:53 | gibi | I might ping in the future with cinder timeouts :) | |
| 15:03:07 | gibi | I might ping _you_ :) | |
| 15:05:28 | lyarwood | gibi: always happy to help, even if it debugging cinder timeouts ;) | |
| 15:05:34 | lyarwood | it is* | |
| 15:06:19 | gibi | honestly, I'm so lost in cinder that I alway get stuck with these failures, give up and go do something else :) | |
| 15:08:56 | lyarwood | gibi: I'm the same with Neutron and PCI stuff tbh | |
| 15:09:06 | lyarwood | actually not the PCI stuff anymore | |
| 15:09:11 | gibi | lyarwood: you can ping me with PCI stuff if needed | |
| 15:09:13 | gibi | ohh | |
| 15:09:14 | gibi | bumer | |
| 15:09:17 | lyarwood | but the Neutron flows still confuse the hell out of me | |
| 15:09:30 | gibi | I have nothing to offer :D | |
| 15:10:09 | bauzas | hmmm, TIL I learned that AZ is mandatory even if the doc says it's optional https://docs.openstack.org/nova/latest/reference/api-microversion-history.html#id70 | |
| 15:10:14 | bauzas | gibi: others ^ | |
| 15:11:01 | bauzas | https://docs.openstack.org/api-ref/compute/?expanded=unshelve-restore-shelved-server-unshelve-action-detail#unshelve-restore-shelved-server-unshelve-action | |
| 15:11:14 | bauzas | availability_zone (Optional) | |
| 15:11:14 | bauzas | body | |
| 15:11:14 | bauzas | ||
| 15:11:14 | bauzas | string | |
| 15:11:15 | bauzas | ||
| 15:11:16 | bauzas | The availability zone name. Specifying an availability zone is only allowed when the server status is SHELVED_OFFLOADED otherwise a 409 HTTPConflict response is returned. | |
| 15:11:17 | bauzas | New in version 2.77 | |
| 15:11:22 | bauzas | what the heck it is | |
| 15:12:26 | bauzas | so, basically, if you use the latest microversion, we break you | |
| 15:12:35 | bauzas | you need to pass a specific AZ | |
| 15:12:39 | bauzas | when unshelving | |
| 15:13:34 | lyarwood | https://github.com/openstack/nova/blob/240ee3091c5ec458753983afa90e3a0cc11dc322/nova/api/openstack/compute/shelve.py#L95-L98 | |
| 15:13:37 | gibi | bauzas: is it an api ref doc bug or a code bug? | |
| 15:13:45 | lyarwood | so wouldn't {'unshelve': null} still work? | |
| 15:14:15 | bauzas | lyarwood: no | |
| 15:14:22 | bauzas | see https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/schemas/shelve.py#L31 | |
| 15:14:35 | bauzas | gibi: well, the doc says it's an optional field | |
| 15:14:53 | bauzas | and honestly, I don't know why it should be a mandatory field | |