Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-15
09:13:01 bauzas yeah, confirmed, the timings match between the job-output and the n-api log
09:13:18 bauzas in n-api I can see Feb 14 18:18:57.185675 np0033093378 devstack@n-api.service[74363]: DEBUG nova.api.openstack.wsgi [None req-de265685-f44e-4d63-98a7-2e38964ce455 tempest-AttachVolumeNegativeTest-1605485622 tempest-AttachVolumeNegativeTest-1605485622-project] Calling method '<bound method ServersController.show of <nova.api.openstack.compute.servers.ServersController object at 0x7f3bd48f7dc0>>' {{(pid=74363) _process_stack /opt/s
09:13:19 bauzas tack/nova/nova/api/openstack/wsgi.py:513}}
09:13:47 bauzas that correlates with job-output : 2023-02-14 18:22:39.100791 | controller | 2023-02-14 18:18:57,782 92653 INFO [tempest.lib.common.rest_client] Request (AttachVolumeNegativeTest:test_attach_attached_volume_to_different_server): 200 GET https://10.176.197.146/compute/v2.1/servers/6a265379-ebfd-4aea-a081-8b271f32c0ea 0.621s
09:14:31 gibi one thing we can do before bump the ssh timeout is to add "os-getConsoleOutput" nova API call to tempest when ssh timeouts to see the guest log at that point in time
09:14:44 bauzas gibi: I dunno what happens but yeah, the discover happens way earlier than the ssh timeout AFAICU
09:15:22 bauzas or
09:15:53 bauzas between the last console timestamp that says [12...] and the 3rd discover, then 2min15s lasted
09:17:52 bauzas https://bugs.launchpad.net/cirros/+bug/1273159
09:18:08 bauzas "It only sends up to 3 DHCP discover packets with a 60 second pause between." :)
09:18:21 bauzas so yeah, that explains
09:18:57 bauzas that explains why we don't see the dhcp lease failing, but that doesn't explain why after 2mins, the dhcp server wasn't providing a lease
09:25:07 gibi_ bauzas: this was the last thing I saw
09:25:08 gibi_ 10:18 < bauzas> so yeah, that explains
09:25:33 bauzas (10:18:57) bauzas: that explains why we don't see the dhcp lease failing, but that doesn't explain why after 2mins, the dhcp server wasn't providing a lease
09:25:50 gibi_ yeah
09:26:05 gibi_ so it might be a neutron dhcp issue
09:26:17 bauzas gibi: I'm inclined to say that increasing the ssh timeout wouldn't help
09:26:49 gibi_ bauzas: it might prove that the 3rd discover fails too and then udhcp gives up
09:27:03 gibi_ but I agree that it is unlikely that the 3rd discover will succeed
09:27:16 gibi_ (except if this is a slow worker node situation)
09:28:00 bauzas well
09:28:30 bauzas I wouldn't say that a ssh connection successful after 3 mins would be a good user experience
09:28:46 bauzas we asked to bind the port way earlier
09:29:07 bauzas and the network cache was updated
09:29:15 bauzas so this doesn't look a vif plugging issue
09:29:29 bauzas gibi: welcome back :)
09:29:40 bauzas anyway, I gonna drop my investigations
09:29:54 bauzas we identified the problem
09:30:12 bauzas now we now it can somehow be kinda-related to https://bugs.launchpad.net/nova/+bug/2006467
09:32:35 gibi_ bauzas: would it make sense to ping #openstack-neutron with ^^
09:33:14 bauzas gibi_: seems to me a good thing to do + adding neutron to the list of projects impacted by that bug
09:43:38 opendevreview Merged openstack/nova-specs master: Amend FQDN in hostname spec to reflect implementation https://review.opendev.org/c/openstack/nova-specs/+/872422
10:15:12 sean-k-mooney bauzas: its not really a openstack bug is it
10:15:39 sean-k-mooney bauzas: like its not a nova bug we have nothing to do with dhcp
10:16:05 sean-k-mooney bauzas: perhaps we can finally bump the verion of cirros in devstack
10:16:20 sean-k-mooney to a 6.x version
10:16:31 sean-k-mooney that should fix some of the kernel panics we see too
10:16:56 stephenfin Do folks think dropping the legacy migrations and fully supporting SQLA 2.0 is viable during the A cycle? https://review.opendev.org/c/openstack/nova/+/872428/
10:17:22 sean-k-mooney you mean by tommorow right
10:17:37 stephenfin oof, I thought we had longer for services
10:17:54 sean-k-mooney i saw you submit that and also use sdk seriese
10:17:54 stephenfin then I guess that answers my question 0:)
10:18:28 sean-k-mooney well we might be able too but i have not looked at the patch yet
10:18:32 stephenfin yeah, the use SDK series is a nice-to-have and highlights some SDK gaps. This one's a _little_ more urgent
10:18:52 bauzas stephenfin: while I understand your concern, I'm afraid of merging it 2 days before FF, given the huge number of CI failures we havbe
10:18:55 sean-k-mooney i skimed it when you first submitted it and the second patch failed zuul
10:19:04 sean-k-mooney look liek it passed after a recheck
10:19:24 sean-k-mooney bauzas: why
10:19:37 sean-k-mooney this code is not used currently
10:20:15 sean-k-mooney if i recall correctly its only used if your coming form pre train
10:20:40 sean-k-mooney actully pre wallaby based on the commit message
10:21:04 sean-k-mooney so thsi wont impact any ci jobs
10:21:31 sean-k-mooney stephenfin: for what its worth we can land them early after RC1 if we dont land them now
10:21:52 bauzas sean-k-mooney: because we already have a lot of changes to be merged
10:22:04 bauzas and I'm done with the CI failures
10:22:07 stephenfin yeah, early in RC1 would work also. We just don't want to drag our feet on this.
10:22:56 sean-k-mooney well i ment after RC1 so in bobcat
10:23:17 sean-k-mooney i would be ok with proceedign with this in Antelope however
10:23:26 sean-k-mooney if ci does not object
10:25:00 sean-k-mooney what is the sate of ci in general currently. has the tempest issue been fixed
10:25:09 bauzas yes
10:25:17 bauzas but now we also have other issues
10:29:16 opendevreview Alexey Stupnikov proposed openstack/nova stable/victoria: Clean up when queued live migration aborted https://review.opendev.org/c/openstack/nova/+/845754
10:30:11 sean-k-mooney stephenfin: im +2 on both of those if others care to review
10:30:30 sean-k-mooney otherwise we can try and land it again in 2 weeks or so
11:35:04 sean-k-mooney bauzas: is https://review.opendev.org/c/openstack/nova/+/873584 just there so we can merge it before RC1 or have ye found the issue in the tests?
11:35:26 bauzas sean-k-mooney: we merged it yesterday
11:35:36 bauzas I mean the logging patch
11:35:47 bauzas so I'd prefer to keep it logging for more than one day
11:37:29 sean-k-mooney ack i was just asking if they issue had been found or not
11:37:35 opendevreview Rajesh Tailor proposed openstack/nova master: Fix case-sensitivity for metadata keys https://review.opendev.org/c/openstack/nova/+/873901
11:37:42 sean-k-mooney we can leave it there for a week or so until we get clsoe to rc1
12:25:56 opendevreview Merged openstack/nova master: cpu: interfaces for managing state and governor https://review.opendev.org/c/openstack/nova/+/868236
13:50:08 opendevreview Alexey Stupnikov proposed openstack/nova stable/victoria: Cleanup old resize instances dir before resize https://review.opendev.org/c/openstack/nova/+/864730
14:09:07 bauzas gibi: another interesting usual suspect kinda related https://ae59d1e8526fa7671728-240e4b572b6f89b26c1b0e70b1c00c17.ssl.cf1.rackcdn.com/872413/3/check/nova-multi-cell/5e89e48/testr_results.html
14:09:25 bauzas somethink looks wrong to me with the userplane
14:09:41 bauzas this time, we got a lease directly
14:09:49 bauzas but when adding the route, it failed
14:10:07 bauzas and when calling the metadata API, we got a failure
14:25:26 bauzas gibi: oh, ralonsoh did an update on the udhcpc bug https://review.opendev.org/c/openstack/neutron/+/871272
14:26:18 bauzas I guess we need to update our zuul config, lemme check
14:26:28 gibi bauzas: I don't know if the root of this is the warning in the route add, or that just a red herring and we have a failure in metadata request handling in neutron or in nova
14:26:59 bauzas the 'route add default gw' command failed
14:27:08 bauzas so basically there is no default route set
14:27:08 gibi with a WARN :D
14:27:21 sean-k-mooney bauzas: that happens somethimes
14:27:24 sean-k-mooney its normal
14:27:28 bauzas that's why I guess the call to the metadata API is failing
14:27:39 bauzas since the route to the network isn't told
14:27:56 sean-k-mooney adding the default route often fails because its already there
14:28:00 bauzas sean-k-mooney: I'm just checking the 'normality' in logsearch
14:28:14 bauzas route: SIOCADDRT: File exists WARN: failed: route add -net "0.0.0.0/0" gw "10.1.0.1"
14:28:15 sean-k-mooney i litrally have been seeing that for years
14:28:28 bauzas technically there is indeed a default route set
14:28:45 bauzas because of the 'file exists'
14:29:52 sean-k-mooney i dont think this is related
14:30:15 sean-k-mooney as i said im used to seeing htat warning. its possible however i think its unlikely
14:30:42 bauzas well, OK
14:31:03 bauzas I already spent my month investigating

Earlier   Later