| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-16 | |||
| 15:29:13 | bauzas | ralonsoh: something new https://zuul.opendev.org/t/openstack/build/76f29afe3f134f139a48b537de7029dc | |
| 15:29:14 | fungi | so you're probably right | |
| 15:29:23 | bauzas | fungi: doh | |
| 15:29:28 | fungi | sadly | |
| 15:29:35 | bauzas | ok, I was wanting to query all the recheck messages I wrote | |
| 15:29:44 | bauzas | and since I haven't followed a clear pattern... | |
| 15:30:07 | fungi | could script that through the rest api, but it would be a bit of work | |
| 15:30:19 | fungi | depends on how much you want it, i guess | |
| 15:31:28 | bauzas | well, it was just for a quick lool | |
| 15:31:35 | bauzas | look* even | |
| 15:31:37 | bauzas | nevermind | |
| 15:32:16 | ralonsoh | bauzas, I don't know where this message is coming, but this is the OVS local service | |
| 15:32:38 | bauzas | yup on n-cpu | |
| 15:32:55 | bauzas | but apparently it's enough serious to do a reschedule | |
| 15:33:19 | bauzas | ... which on an AIO doesn't help | |
| 15:33:22 | bauzas | :) | |
| 15:34:05 | ralonsoh | bauzas, actually this is something normal, coming from the python ovs bindings | |
| 15:34:05 | ralonsoh | def __log_wakeup(self, events): | |
| 15:34:05 | ralonsoh | if not events: | |
| 15:34:06 | ralonsoh | vlog.dbg("%d-ms timeout" % self.timeout) | |
| 15:34:52 | ralonsoh | https://github.com/openvswitch/ovs/blob/master/python/ovs/poller.py#L246# | |
| 15:37:00 | bauzas | hmmmm | |
| 15:37:23 | bauzas | I'm able to find another patchset that got the exact failure from the same job on the same test tempest.api.compute.servers.test_list_servers_negative.ListServersNegativeTestJSON | |
| 15:37:27 | bauzas | https://d2746e36843633ae266c-ad135f72b22a132e11a904324fbc4e60.ssl.cf1.rackcdn.com/872413/3/check/tempest-integrated-compute-enforce-scope-new-defaults/5204f72/job-output.txt | |
| 15:41:05 | ralonsoh | bauzas, I'm checking https://44f9259a9cd22acee92d-000061e1666ecf9c52f0643ab3c391ab.ssl.cf1.rackcdn.com/873934/1/check/nova-next/785cc57 | |
| 15:41:15 | ralonsoh | for example the first test case failing | |
| 15:41:24 | ralonsoh | test_attach_scsi_disk_with_config_drive | |
| 15:42:01 | ralonsoh | I see the DHCP agent configuring dnsmasq process, I see n-cpu creating the interface, OVS agent receiving this creation event | |
| 15:42:17 | ralonsoh | but I see nowhere in the DHCP agent logs the DHCPREQUEST for this fixed IP | |
| 15:42:43 | ralonsoh | and the tempest test doesn't retrieve the VM logs | |
| 15:43:26 | dansmith | yeah I'm not sure why we're not logging the console in that case | |
| 15:43:35 | dansmith | because we should see that, to know if it's a guest crash | |
| 15:46:15 | ralonsoh | dansmith, qq, this is about the image resources | |
| 15:46:25 | ralonsoh | Feb 15 15:24:26.291000 np0033111254 nova-compute[39923]: DEBUG nova.compute.resource_tracker [None req-a80484d7-0e98-4086-a01f-afda9d54f06b None None] Instance 749a5ad0-e7b8-47f2-8e9c-78896c92b80e actively managed on this compute host and has allocations in placement: {'resources': {'VCPU': 1, 'MEMORY_MB': 128, 'DISK_GB': 1}}. {{(pid=39923) _remove_deleted_instances_allocations /opt/stack/nova/nova/compute/resource_tra | |
| 15:46:25 | ralonsoh | cker.py:1632}} | |
| 15:46:32 | bauzas | despite, none of them show the console log | |
| 15:46:38 | ralonsoh | why are you using 128MB? | |
| 15:46:49 | ralonsoh | cirros should be using 256, right? | |
| 15:47:12 | bauzas | hmmm | |
| 15:47:14 | dansmith | ralonsoh: idk, maybe that was the result of a resize? | |
| 15:47:29 | bauzas | that's a good catch | |
| 15:47:35 | ralonsoh | dansmith, no, this is the VM booting | |
| 15:47:41 | bauzas | I haven't checked the flavors that were related to the cirros failures I saw | |
| 15:47:42 | ralonsoh | first log in the compute agent | |
| 15:47:54 | dansmith | flavors are 42, and 84 | |
| 15:48:21 | ralonsoh | 128 and 192 | |
| 15:48:26 | ralonsoh | I'll check neutron CI | |
| 15:49:52 | dansmith | 42 is the default, and that's 128M | |
| 15:50:19 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: Test aborting queued live migration https://review.opendev.org/c/openstack/nova/+/873575 | |
| 15:50:34 | dansmith | 84 is 192M yeah | |
| 15:51:06 | dansmith | so perhaps we're flying too close to the sun with 128M | |
| 15:51:38 | ralonsoh | dansmith, I'm checking what we are using in Neutron | |
| 15:52:54 | dansmith | AFAIK, these are devstack defaults | |
| 15:53:10 | ralonsoh | we use the default values too, 128M | |
| 15:53:14 | dansmith | bumping the flavor memory is likely to result in other instabilities, I fear | |
| 15:53:16 | ralonsoh | {'resources': {'DISK_GB': 1, 'MEMORY_MB': 128, 'VCPU': 1}} | |
| 15:53:20 | dansmith | yeah | |
| 15:54:20 | ralonsoh | if you could retrieve the console logs, that could help | |
| 15:54:32 | ralonsoh | at least to know that the VM tried to request an IP | |
| 15:54:43 | ralonsoh | and, somewhere, this request was dropped | |
| 15:55:27 | dansmith | yeah, we should chat with gmann about it. I can go look at the tempest stuff to figure out why, | |
| 15:55:32 | dansmith | but he probably knows off the top of his head | |
| 16:04:19 | bauzas | ralonsoh: I have a few logs where the guest failed to acquire a lease, sec | |
| 16:04:36 | bauzas | and some where the guest panickjed | |
| 16:06:11 | sean-k-mooney | for what its worht the every increaseign memory requirement for cirrios was one of the reasons i looked at moving us to alpine a few years ago | |
| 16:06:20 | sean-k-mooney | longterm i still think that would be a better approch | |
| 16:06:58 | dansmith | is alpine really going to be smaller than cirros? I mean, that seems odd to me | |
| 16:07:10 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/victoria: Test aborting queued live migration https://review.opendev.org/c/openstack/nova/+/845748 | |
| 16:07:11 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/victoria: Add functional tests to reproduce bug #1960412 https://review.opendev.org/c/openstack/nova/+/845753 | |
| 16:07:22 | dansmith | so I think we get the automatic console stuff when we use specific waiters for servers to be available | |
| 16:07:28 | dansmith | so that test must use a different one | |
| 16:07:52 | bauzas | oh damn, internal meeting | |
| 16:08:13 | ralonsoh | bauzas, do you have some links? in any case, this is not Nova nor Neutron fault, I think | |
| 16:08:14 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/victoria: Clean up when queued live migration aborted https://review.opendev.org/c/openstack/nova/+/845754 | |
| 16:08:16 | ralonsoh | at least in this case | |
| 16:08:47 | bauzas | ralonsoh: yup, which kind of failures do you want to see ? for the dhcp query? | |
| 16:09:04 | ralonsoh | yes and the kernel panic | |
| 16:20:57 | bauzas | ralonsoh: one for segfaults https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_da2/821228/7/check/nova-multi-cell/da2689f/job-output.txt | |
| 16:23:22 | bauzas | ralonsoh: that one for kernel panicking https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_e4a/821228/7/gate/nova-next/e4ab52f/job-output.txt | |
| 16:24:47 | bauzas | ralonsoh: that one is an interesting case where the default route is already present and metadata goes into weeds https://ae59d1e8526fa7671728-240e4b572b6f89b26c1b0e70b1c00c17.ssl.cf1.rackcdn.com/872413/3/check/nova-multi-cell/5e89e48/job-output.txt | |
| 16:27:11 | ralonsoh | I'll check this last one | |
| 16:28:05 | ralonsoh | bauzas, eh hold on, this could be a problem with the OVN version in Jammy | |
| 16:28:27 | ralonsoh | --> https://review.opendev.org/c/openstack/neutron/+/873684 | |
| 16:28:59 | ralonsoh | there is an issue with ovn v22.03.0, included in jammy | |
| 16:29:05 | ralonsoh | and some missing flows for the metadata | |
| 16:29:17 | ralonsoh | Yatin found it and we are skipping those tests | |
| 16:29:38 | ralonsoh | actually we are going to test using a compiled version of OVN | |
| 16:29:41 | ralonsoh | https://review.opendev.org/c/openstack/neutron/+/874112/1 | |
| 16:30:52 | bauzas | ack ok | |
| 16:51:31 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: Test aborting queued live migration https://review.opendev.org/c/openstack/nova/+/873575 | |
| 16:51:32 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: Add functional tests to reproduce bug #1960412 https://review.opendev.org/c/openstack/nova/+/873576 | |
| 17:01:18 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: Clean up when queued live migration aborted https://review.opendev.org/c/openstack/nova/+/873577 | |
| 17:09:43 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: Clean up when queued live migration aborted https://review.opendev.org/c/openstack/nova/+/873577 | |
| 17:25:16 | fungi | ralonsoh: any idea if the fixes have been backported to v22 such that ubuntu could do an sru to patch their packages? | |
| 17:25:57 | ralonsoh | fungi, Yatin opened a bug today: https://bugs.launchpad.net/ubuntu/+source/ovn/+bug/2003056 | |
| 17:26:02 | fungi | awesome | |
| 17:26:12 | ralonsoh | (sorry, not today) | |
| 17:26:31 | fungi | i'd hate to see devstack using a bespoke ovn build long-term | |
| 17:26:42 | fungi | here's hoping they're able to patch it | |
| 17:27:15 | ralonsoh | from https://bugs.launchpad.net/ubuntu/+source/ovn/+bug/2003056/comments/4, there should be a new version now (22.04.1, instead of 22.03) | |