Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-12
10:14:25 gibi case
10:15:13 zigo Hard to tell...
10:16:05 zigo Last output was: https://paste.opendev.org/show/bsDsX7d0gw1SDAnLcRMA/
10:16:20 zigo Then it hangged ...
10:16:36 sean-k-mooney whats that form?
10:16:49 zigo sean-k-mooney: building oslo.concurrency.
10:17:23 sean-k-mooney oh i disconnected and reconnected for a bit
10:17:32 sean-k-mooney missed the start of your conversation with gibi i think
10:18:11 zigo Problem is: I have autopkgtest issues in Eventlet ... 0.33.x :/
10:18:47 sean-k-mooney https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L184-L201
10:20:10 sean-k-mooney i see
10:20:28 sean-k-mooney so thats failing whiel building in some cases
10:20:58 sean-k-mooney im not really sure who/why
10:21:01 gibi zigo: could you point to how you run the unti tests?
10:21:31 sean-k-mooney the est is just concorting an instnace of an exeption class
10:21:58 sean-k-mooney asserting when you call str() on it that it contians the message
10:22:10 gibi there are unit tests in the repo that can be run with and without eventlet monkey patching but there are tests that can only run in eventlet
10:22:13 gibi hence the https://github.com/openstack/oslo.concurrency/blob/01cf2ffdf48c21f886b2aa3f766be5d268248c18/tox.ini#L14-L15
10:22:30 sean-k-mooney maybe this print is the issue https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L203
10:22:49 zigo I simply do this:
10:22:49 zigo PYTHON=python3 stestr run --subunit | subunit2pyunit
10:23:11 gibi I can imagine that if you run the eventlet aware test without eventlet monkey patching then the eventlet only test might missbehave
10:23:38 zigo I'll try further and let you know where it leads me.
10:24:07 sean-k-mooney if you do things like eventlet.spwawn directly without monkeypatching
10:24:17 sean-k-mooney you need to manually invoke the event loop to have it run
10:24:33 sean-k-mooney we saw that in the nova-api when we added scater gather
10:25:20 gibi zigo: based on that command line you run without eventlet monkey patching but you still run tests from https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/unit/test_lockutils_eventlet.py
10:26:07 gibi zigo: you can try not running those test to see if that resolve the hang
10:26:30 sean-k-mooney you could fix them by adding eventlet.sleep(seconds=0)
10:26:48 sean-k-mooney i think that will make it work if not monkeypatched
10:27:05 sean-k-mooney but ya not runnign them would be better
10:27:24 gibi there is the place where the test monkey patches selectively https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/__init__.py
10:28:10 sean-k-mooney https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_lockutils_eventlet.py#L49
10:28:23 sean-k-mooney so if we also do that on lin 51 before the pool.waitall()
10:28:29 sean-k-mooney i think that might actully work in either case
10:28:55 sean-k-mooney but i might make sense for use to use the skip funciton in the classs
10:29:02 sean-k-mooney to have it skip in not monkey patched
10:35:57 sean-k-mooney gibi: is there any reason not to check if we are monkey patched in the setup function and call skipTest
10:40:21 gibi I don't know we might need to consult with oslo cores
10:40:59 gibi because there was eventlet specific tests before I added mine I thought it is handled centrally not to run them in a non patched env
10:41:06 gibi this might not be the case
10:41:31 sean-k-mooney i dont see that logic genericlly
10:41:44 sean-k-mooney its certenly possible to add
11:20:24 noonedeadpunk hello folks! I was wondering - does reverting resize failure rings anybody a bell? Ie - create server, resize server, revert resize -> VM is "stuck" in REVERT_RESIZE until message timeouts, then it goes back to VERIFY_RESIZE but with original flavor, and then nova-compute shutdown VM on hypervisor at all. The only way to recover is to reset state
11:21:51 sean-k-mooney not that i recall but i can see that poteilaly happening if we raise an excption in the revert path and cant proceed
11:22:17 sean-k-mooney like if the souce host was down or soemthing like that we would not be able to revert
11:22:58 noonedeadpunk paste: https://paste.openstack.org/show/bh3kML89sPFYDN9HkDSd/
11:23:54 noonedeadpunk sean-k-mooney: to have that said, out of 100 tempest runs of tempest.api.compute.servers.test_server_actions.ServerActionsTestJSON.test_resize_server_revert 57 has failed
11:24:47 noonedeadpunk the stack trace I've spotted: https://paste.openstack.org/show/bytEWO0CHe8cSVktE3tv/
11:25:43 noonedeadpunk I assumed it can be related to the heartbeat_in_pthread thing, as in the region it was "default" setting on Xena (which is enabled), but disabling it didn't fix that
11:27:33 sean-k-mooney do you have any timeouts form ovsdbapp
11:27:52 sean-k-mooney in the nova-compute logs
11:30:01 noonedeadpunk sean-k-mooney: um, nope
11:30:10 noonedeadpunk it's ovs, not ovn fwiw
11:30:22 sean-k-mooney ya it would be the same either way
11:30:48 sean-k-mooney https://bugzilla.redhat.com/show_bug.cgi?id=2085583
11:30:56 sean-k-mooney i was wondering if it was related to that
11:32:33 sean-k-mooney those are still pending backport upstream https://review.opendev.org/c/openstack/os-vif/+/841771
11:32:39 sean-k-mooney https://review.opendev.org/c/openstack/os-vif/+/841772/1
11:33:08 sean-k-mooney if you are using the native os-vif backend then the compute agent can hang if the connection to the ovs db drops
11:33:13 sean-k-mooney and that can cause timeouts
11:33:25 sean-k-mooney the workaround is to just use the vsctl backend
11:34:30 sean-k-mooney you will see something that looks like thi
11:34:33 sean-k-mooney 2022-09-01 16:33:38.583 2 DEBUG ovsdbapp.backend.ovs_idl.vlog [-] [POLLIN] on fd 18 __log_wakeup /usr/lib64/python3.9/site-packages/ovs/poller.py:263
11:34:34 noonedeadpunk in neutron-ovs-agent the only string referencing port is `neutron.agent.common.ovs_lib [req-3aa28a82-1d30-4713-989a-e5093e55f7ab - - - - -] Port 200e682f-b9af-487c-aa8c-605e20b99002 not present in bridge br-int`
11:34:35 sean-k-mooney 2022-09-01 16:33:38.584 2 DEBUG ovsdbapp.backend.ovs_idl.vlog [-] tcp:127.0.0.1:6640: entering ACTIVE _transition /usr/lib64/python3.9/site-packages/ovs/reconnect.py:519
11:34:37 sean-k-mooney 2022-09-01 16:33:40.874 2 DEBUG ovsdbapp.backend.ovs_idl.vlog [-] [POLLIN] on fd 18 __log_wakeup /usr/lib64/python3.9/site-packages/ovs/poller.py:263
11:34:39 sean-k-mooney 2022-09-01 16:33:43.584 2 DEBUG ovsdbapp.backend.ovs_idl.vlog [-] [POLLIN] on fd 18 __log_wakeup /usr/lib64/python3.9/site-packages/ovs/poller.py:263
11:34:41 sean-k-mooney 2022-09-01 16:33:48.587 2 DEBUG ovsdbapp.backend.ovs_idl.vlog [-] 4999-ms timeout __log_wakeup /usr/lib64/python3.9/site-packages/ovs/poller.py:248
11:34:47 sean-k-mooney noonedeadpunk: not in neutron in the nova-compute agent
11:35:08 noonedeadpunk ah, I would need to have DEBUG enabled
11:35:33 noonedeadpunk let me enable it and reproduce :)
11:35:37 sean-k-mooney noonedeadpunk: am well that just means that the logical port is defiend in the nortdb
11:35:58 sean-k-mooney but not actully added to the br-int bridge yet
11:36:00 sean-k-mooney i think
11:36:28 sean-k-mooney so not northdb but more or less the same
11:36:44 sean-k-mooney neutron know that the port should exist bug nova has not added it yet
11:37:34 sean-k-mooney noonedeadpunk: if you dont have debug enabled the symthom to look for is large gaps in the log of 5+ seconds without output
11:38:02 sean-k-mooney that however is not helpful if the system is idel or not activly spwaning a vm
11:38:22 sean-k-mooney since that is what you woudl expect it to look like unless you asked nova to do something
11:40:49 noonedeadpunk yeah, well, delay between `Updating port 73862d0c-b31b-4b75-b833-8e29d8066b9a with attributes {'binding:host_id': 'cc-compute04-tky1', 'device_owner': 'compute:nova'}` and stack trace is exactly 5 sec. But it's exactly the timeout
11:42:11 sean-k-mooney that is a little sus yes
11:44:19 sean-k-mooney the tl;dr of https://bugs.launchpad.net/os-vif/+bug/1929446 is that in the ovs python bindign they backislly make a blocking call to select.poll() which blocks the main thread
11:44:35 sean-k-mooney *basically
11:45:02 sean-k-mooney http://patchwork.ozlabs.org/project/openvswitch/patch/20210611142923.474384-1-twilson@redhat.com/
11:47:36 noonedeadpunk ah
11:54:09 sean-k-mooney by the way we have seen ValueError: Circular reference detected before
11:54:22 sean-k-mooney but i dont think we ever figured out where that comes form or how
11:55:00 sean-k-mooney any time we have seen it there has always been some other error too
11:55:17 sean-k-mooney and fixing the othe rerror has caused both to go away
11:56:34 noonedeadpunk oh yes, I do see `DEBUG ovsdbapp.backend.ovs_idl.vlog [-] [POLLIN] on fd 31 __log_wakeup /openstack/venvs/nova-24.0.0/lib/python3.8/site-packages/ovs/poller.py:263` when revert_resize is failed
11:57:08 noonedeadpunk thanks sean-k-mooney, I will check out if patch will work :)
11:57:51 sean-k-mooney its oke to see it if you dont see large timeouts or see it ocationally
11:58:16 noonedeadpunk I see it only when revert is failing
11:58:17 sean-k-mooney otherwise etiehr backport that patch or set os vif to use vsctl
11:58:23 sean-k-mooney ack
11:58:59 sean-k-mooney [os_vif_ovs]/ovsdb_interface=vsctl
11:59:24 sean-k-mooney setting that in your nova.conf will also workaround it
12:00:58 noonedeadpunk ah, undocumented option?:)
12:01:12 noonedeadpunk let me try it out then

Earlier   Later