Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-19
15:35:02 artom I'll just remove it from the commit message :P
15:35:14 sean-k-mooney ill see if i can figure out why
15:48:11 sean-k-mooney dansmith: so it soudn like its hitting the code for https://github.com/openstack/neutron/commit/8a55f091925fd5e6742fb92783c524450843f5a0
15:50:38 sean-k-mooney hum so at the time of the port bidning
15:51:30 sean-k-mooney there are no errro in the metadta aganet log but there are gaps for 3-6 seconds at a tiem and its interacting with both ovs and privsep
16:12:18 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Re-order and parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
16:12:19 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Call Neutron immediately upon _post_live_migration() start https://review.opendev.org/c/openstack/nova/+/883682
16:15:42 sean-k-mooney dansmith: so my best guess is its related to thsi change https://github.com/openstack/neutron/commit/628442aed7400251f12809a45605bd717f494c4e
16:16:16 sean-k-mooney 7 mounts ago they started trying to spread the agent heatbeats
16:16:42 sean-k-mooney im seeing logs to the effect fo delaying update to the cachs table for 10 seconds
16:17:02 sean-k-mooney around when the agent prior to the agent being detected as dead
16:17:24 sean-k-mooney my guess is if the agent is doign somthign like writing to the ovs db
16:17:36 sean-k-mooney it can miss the heatbeat
16:18:04 sean-k-mooney Delaying updating chassis table for 23 seconds {{(pid=38857) run /opt/stack/neutron/neutron/agent/ovn/metadata/agent.py:243}}
16:18:23 sean-k-mooney im seeign quite a spread
16:20:08 dansmith artom: ack I figured, probably better to make it accurate though yeah :)
16:20:23 dansmith sean-k-mooney: ah, interesting
16:20:37 dansmith sean-k-mooney: so like under heavy load they're missing some heartbeats maybe
16:20:53 sean-k-mooney ya perhaps
16:21:21 sean-k-mooney im goign to put up a tiny patch to change that form cfg.CONF.agent_down_time // 2 to cfg.CONF.agent_down_time // 3
16:21:33 sean-k-mooney that will make it heat beat a little more often
16:22:02 dansmith ack cool
16:22:06 sean-k-mooney that was recently done for rabbit 2 -> 3 for similar reasons
16:26:50 sean-k-mooney oh its not merged yet https://review.opendev.org/c/openstack/oslo.messaging/+/875615
16:32:16 sean-k-mooney dansmith: i assume there isnt a bug currently
16:32:27 dansmith sean-k-mooney: not that I've opened
16:32:49 sean-k-mooney ok ill file one quickly with some of the errors i was seeing
16:33:00 sean-k-mooney the logs are not super helpful
16:40:59 dansmith sweet thanks
16:44:56 sean-k-mooney https://bugs.launchpad.net/neutron/+bug/2020215
16:45:15 sean-k-mooney i will push a patch once i run the unit/functional tests and see what breaks
16:59:50 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
17:11:11 sean-k-mooney dansmith: i think https://review.opendev.org/c/openstack/neutron/+/883687 will help but its hard to tell if not then https://bugs.launchpad.net/neutron/+bug/2020215 might give the neutron folks another idea
17:11:27 dansmith ack thanks for chasing that
17:11:45 sean-k-mooney im going to finish there for today o/
17:11:55 dansmith thanks, enjoy the weekend
17:25:40 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
20:06:25 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
23:01:58 opendevreview Artom Lifshitz proposed openstack/nova master: POC: Parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
#openstack-nova - 2023-05-20
00:36:55 opendevreview David Hill proposed openstack/nova stable/train: Allow an operator to override the default proto type of a VF https://review.opendev.org/c/openstack/nova/+/883732
00:38:49 opendevreview David Hill proposed openstack/nova stable/train: Allow an operator to override the default proto type of a VF https://review.opendev.org/c/openstack/nova/+/883732
01:54:40 opendevreview Artom Lifshitz proposed openstack/nova master: Call Neutron immediately upon _post_live_migration() start https://review.opendev.org/c/openstack/nova/+/883682
01:54:41 opendevreview Artom Lifshitz proposed openstack/nova master: Parallelize calls to Neutron and Cinder in post_live_migration https://review.opendev.org/c/openstack/nova/+/883678
21:11:02 opendevreview David Hill proposed openstack/nova stable/train: Allow an operator to override the default proto type of a VF https://review.opendev.org/c/openstack/nova/+/883732
21:11:47 opendevreview David Hill proposed openstack/nova master: Allow an operator to override the default proto type of a VF https://review.opendev.org/c/openstack/nova/+/883745
21:16:13 opendevreview David Hill proposed openstack/nova master: Allow an operator to override the default proto type of a VF https://review.opendev.org/c/openstack/nova/+/883745
#openstack-nova - 2023-05-22
07:27:12 bauzas good morning
08:05:55 gibi o/
08:06:00 ykarel sean-k-mooney[m], gibi can you please check https://bugs.launchpad.net/neutron/+bug/2015065 comment 7/8
08:07:38 ykarel randomly one of nova-api worker just get's stuck when doing requests to neutron(not sure if same is seen with any other service yet)
08:42:37 gibi ykarel: quickly looked at the bug. Thanks for collecting all that data. When the nova-api stuck in calling neutronclient's show_security_group do you see that the actualy API request to neutron was sent but never received by neutron-server? Or nova-api is stuck on sending the message?
08:49:24 ykarel gibi, i don't see the request received on neutron side, not sure where to check if it's stuck on sending
08:51:27 gibi ykarel: ack
09:03:20 gibi ykarel:
09:03:47 gibi i feel like we are seeing an interesting interaction between multiple things
09:04:15 gibi I'm trying to follow the stack trace from the latest comment from the bug to see where the neutronclient got stuck
09:04:33 gibi the firts interesting point is
09:04:33 gibi /usr/local/lib/python3.10/dist-packages/urllib3/util/connection.py:28 in is_connection_dropped
09:04:58 gibi https://github.com/urllib3/urllib3/blob/a5b29ac1025f9bb30f2c9b756f3b171389c2c039/src/urllib3/connectionpool.py#L272
09:05:32 gibi so urllib try to check if the existing client connection is still usable or got disconnected
09:05:49 gibi https://github.com/urllib3/urllib3/blob/a5b29ac1025f9bb30f2c9b756f3b171389c2c039/src/urllib3/util/connection.py#L28
09:05:54 gibi wait_for_read(sock, timeout=0.0)
09:06:06 gibi os it checks if it can read from the socket with 0.0 timeout
09:06:40 gibi https://github.com/urllib3/urllib3/blob/a5b29ac1025f9bb30f2c9b756f3b171389c2c039/src/urllib3/util/wait.py#L84-L85
09:06:58 gibi that 0.0 timeout is passed to python's select.select
09:07:07 gibi https://docs.python.org/3.10/library/select.html#select.select
09:07:23 gibi "The optional timeout argument specifies a time-out as a floating point number in seconds. When the timeout argument is omitted the function blocks until at least one file descriptor is ready. A time-out value of zero specifies a poll and never blocks."
09:07:47 gibi so that select.select called with 0.0 should never block
09:07:52 gibi BUT
09:08:22 gibi in our env the envtlet monkey patching is changing python's select.select
09:08:25 gibi /usr/local/lib/python3.10/dist-packages/eventlet/green/select.py:80 in select
09:08:46 gibi and redirects it to implement the envtlet switching mechanism
09:11:16 gibi https://github.com/eventlet/eventlet/blob/88ec603404b2ed25c610dead75d4693c7b3e8072/eventlet/green/select.py#L30-L80C32
09:12:34 gibi looking at that code it seems enventlet sets a timer with the timeout value
09:12:45 gibi via hub.schedule_call_global
09:17:05 gibi here I'm getting lost in the eventlet code but I assume sheduling a timer with 0.0 timeout in eventlet can be racy
09:17:31 gibi based on the comment in https://github.com/eventlet/eventlet/blob/88ec603404b2ed25c610dead75d4693c7b3e8072/eventlet/green/select.py#L62-L69
09:21:38 gibi one could argue that what we see is an eventlet bug as select.select with timeout=0.0 should not ever block but it does block in our case.
09:28:48 opendevreview suzhengwei proposed openstack/nova master: rename 'recreate' to 'evacuate' https://review.opendev.org/c/openstack/nova/+/883810
10:24:19 ykarel Thanks gibi for checking, anyway the issue can be fixed/worked around on nova side?
10:25:38 gibi I'm trying to open an issue on eventlet and see if the maintainer agrees with my analysis or not. I don't see now any easy workaround. Maybe sean-k-mooney or melwitt can see some
10:25:59 gibi ykarel: I will update the launchpad bug
10:26:05 ykarel Thanks gibi
10:28:04 sean-k-mooney gibi: sorry i missed the start of this what is the issue
10:29:47 gibi nova-api using neutron client to call neutron API but get stuck for ever
10:30:12 gibi based on the stack trace it stuck checking if the previous connection is still usable
10:30:34 gibi we end up in eventlet monkeypatched select.select on a socket
10:30:40 gibi with a timeout 0.0
10:31:09 gibi based on the stdlib doc timeout 0.0 means non blocking but we still block
10:31:21 gibi so I assume eventlet not properly handles timeout 0.0 in the eventlet select impl
10:31:22 sean-k-mooney i see
10:31:34 sean-k-mooney that or its python version specirif
10:31:46 sean-k-mooney but ya sound like a api compaitblity bug
10:31:54 gibi details are here https://bugs.launchpad.net/neutron/+bug/2015065
10:36:40 sean-k-mooney Changed in version 3.7: The method no longer toggles SOCK_NONBLOCK flag on socket.type.
10:36:46 sean-k-mooney https://docs.python.org/3/library/socket.html#socket.socket.settimeout
10:37:28 sean-k-mooney ykarel: gibi: it looks like we shoudl not be using 0.0 to make it non-blocking after 3.7
10:38:00 sean-k-mooney we should be using socket.setblocking(false)
10:38:03 gibi sean-k-mooney: https://docs.python.org/3.10/library/select.html#select.select for select.select timeout=0.0 still means
10:38:06 gibi non blocking

Earlier   Later