| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-16 | |||
| 09:53:25 | gibi | the rest is automatic | |
| 09:53:47 | bauzas | aren't we checking the number of traits we support ? | |
| 09:53:47 | gibi | placement loads all the trait defs from whathever os-traits lib it founds at startup | |
| 09:53:49 | bauzas | I thought so | |
| 09:53:59 | gibi | I think we fixed that check | |
| 09:55:35 | gibi | https://review.opendev.org/c/openstack/placement/+/851966 | |
| 09:56:55 | bauzas | gibi: great | |
| 10:14:25 | bauzas | gibi: hah, fun, we forgot to bump the reqs for 2.9.0 https://github.com/openstack/placement/blob/master/requirements.txt#L29 | |
| 10:15:02 | bauzas | gibi: but since we don't pin on a release, we should get 2.10 | |
| 10:15:32 | bauzas | so technically, it's more about saying that for 2023.1 Placement will always support those traits are bare min | |
| 10:15:39 | bauzas | as* bare | |
| 10:24:39 | gibi | bauzas: yeah that is still on me to have a limited lower constraint job running on placement and nova to catch these as we only test with upper today | |
| 10:25:05 | gibi | I tried to do that after the last PTG but it was non trivial so I never finished it | |
| 10:25:09 | bauzas | cool, we haven't branched RC1 yet so we're on time | |
| 10:25:31 | bauzas | for the moment, we need to just make sure we document this | |
| 10:27:42 | gibi | I believe more in systems that enforces rules than documentation that we don't read when we should | |
| 10:29:51 | opendevreview | Sylvain Bauza proposed openstack/placement master: Update 2023.1 reqs to support os-traits 2.10 as min version https://review.opendev.org/c/openstack/placement/+/874080 | |
| 10:30:02 | bauzas | we have the PTL docs that I personnally enforce :) | |
| 10:30:44 | bauzas | gibi: sean-k-mooney: time for a placement review https://review.opendev.org/c/openstack/placement/+/874080 | |
| 10:30:57 | gibi | bauzas: +@ | |
| 10:30:58 | gibi | bauzas: +2 | |
| 10:56:46 | bauzas | nova-next is getting me mad | |
| 10:57:26 | bauzas | https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_e4a/821228/7/gate/nova-next/e4ab52f/testr_results.html is anthologic | |
| 10:57:37 | bauzas | volume timeouts + guest kernel tainting | |
| 11:02:02 | opendevreview | Tobias Urdin proposed openstack/nova master: libvirt: update description for live_migration_completion_timeout https://review.opendev.org/c/openstack/nova/+/874083 | |
| 12:21:20 | gibi | bauzas: is this with the new cirros version? | |
| 12:23:10 | bauzas | gibi: nope, we haven't merged yet the change | |
| 12:23:41 | bauzas | I haven't verified which cirros image we were having with the test, but I guess it's still 0.5 | |
| 12:24:06 | bauzas | 2023-02-15 20:08:15.854905 | controller | ++ stackrc:source:692 : DEFAULT_IMAGE_NAME=cirros-0.5.2-x86_64-disk 2023-02-15 20:08:15.856966 | controller | ++ stackrc:source:693 : DEFAULT_IMAGE_FILE_NAME=cirros-0.5.2-x86_64-disk.img | |
| 12:24:10 | bauzas | indeed | |
| 12:26:53 | sean-k-mooney | i said this a few days ago but i do think we should look at going to 0.6.x | |
| 12:27:17 | sean-k-mooney | there are a few kernel bugs in the 0.5.2 image that cause ocational kernel panics related to the apic in the guest | |
| 12:27:41 | sean-k-mooney | the 0.6.2 image is built on the ubuntu 22.04 kernel | |
| 12:29:10 | gibi | bauzas: thanks, then I guess yet another guest kernel bug | |
| 12:30:06 | bauzas | sean-k-mooney: I have a change for this | |
| 12:30:17 | bauzas | sec. | |
| 12:30:32 | bauzas | sean-k-mooney: https://review.opendev.org/c/openstack/nova/+/873934 | |
| 12:30:56 | bauzas | we could also use 0.6.2 if we want, I'm not against | |
| 12:31:08 | sean-k-mooney | i would prefer ot change it in devstack | |
| 12:31:26 | sean-k-mooney | we can do it per job | |
| 12:31:37 | sean-k-mooney | but it would be better to use the same cirror image in all jobs | |
| 12:32:00 | bauzas | sean-k-mooney: but we can test the new cirros image by nova-next and if we see it works, then we could indeed use it for all our jobs | |
| 12:32:15 | bauzas | nova-next is also there for testing new stuff | |
| 12:32:46 | bauzas | I'm just afraid of changing all our jobs by once without correctly testing first | |
| 12:33:01 | sean-k-mooney | i guess but given we have known kernel issue in the 5.2 image nad we have been trying to change it for years i would prefer to do the change in Antilepe if we can | |
| 14:02:40 | opendevreview | Amit Uniyal proposed openstack/nova master: Added a lock_unlock dcorator for instance https://review.opendev.org/c/openstack/nova/+/873648 | |
| 14:21:20 | opendevreview | Amit Uniyal proposed openstack/nova master: Added context manager for instance lock https://review.opendev.org/c/openstack/nova/+/873648 | |
| 14:42:28 | bauzas | aaaaaand now I see more and more cirros guest segfaulting... https://review.opendev.org/c/openstack/nova/+/821228/7 | |
| 14:42:38 | bauzas | https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_da2/821228/7/check/nova-multi-cell/da2689f/job-output.txt | |
| 14:44:02 | opendevreview | Takashi Natsume proposed openstack/nova master: doc: mark the max microversion for 2023.1 Antelope https://review.opendev.org/c/openstack/nova/+/874103 | |
| 14:50:30 | dansmith | bauzas: where does that scenario:dhcp_client thing come in? | |
| 14:50:56 | bauzas | dansmith: https://review.opendev.org/c/openstack/nova/+/873934 | |
| 14:51:06 | bauzas | dansmith: and https://launchpad.net/bugs/2006467 | |
| 14:51:09 | dansmith | bauzas: right, what does scenario:dhcp_client affect | |
| 14:51:39 | dansmith | does that end up with cirros behaving differently? or does it make tempest do something inside the guest? | |
| 14:51:45 | bauzas | dansmith: ralonsoh told me to use that way b/c of https://review.opendev.org/c/openstack/neutron/+/871272/1/zuul.d/tempest-multinode.yaml | |
| 14:52:02 | bauzas | and I shamelessly copy/pasteed | |
| 14:52:46 | ralonsoh | this is the dhcp client used in cirros 0.6.1 | |
| 14:53:19 | ralonsoh | more info here: https://review.opendev.org/c/openstack/neutron/+/871272 | |
| 14:53:28 | ralonsoh | (in the commit message) | |
| 14:53:55 | dansmith | ralonsoh: but if the cirros client just uses it, what does the tempest config change? | |
| 14:54:39 | ralonsoh | dansmith, https://review.opendev.org/c/openstack/tempest/+/871270/2/tempest/common/utils/linux/remote_client.py | |
| 14:54:49 | ralonsoh | we choose what is the VM dhcp client | |
| 14:55:11 | ralonsoh | --> https://review.opendev.org/c/openstack/tempest/+/871270/2/tempest/config.py | |
| 14:55:21 | dansmith | ralonsoh: that's just for a manual renew, but doesn't affect how the guest works on first boot.. is that what you mean? | |
| 14:55:50 | ralonsoh | dansmith, no, that should not affect how the VM boots | |
| 14:56:06 | ralonsoh | the OS will use the exiting dhcp client | |
| 14:56:12 | dansmith | ralonsoh: gotcha okay, so I guess most of the issues I see getting an IP seem to be related to initial boot | |
| 14:56:17 | ralonsoh | right | |
| 14:56:38 | dansmith | okay cool, just making sure I understand | |
| 15:05:14 | bauzas | ralonsoh: and to clarify, cirros-0.6.x switched its dhch client to dhcpcd ? | |
| 15:05:25 | ralonsoh | yes | |
| 15:06:20 | bauzas | ack | |
| 15:06:32 | bauzas | then I understand what I wrote, huzzah :D | |
| 15:08:03 | dansmith | ralonsoh: does neutron record an event if the IP gets actually leased? | |
| 15:08:32 | dansmith | meaning, on failure can we query to neutron to see if the guest ever pulled its IP, to distinguish between "we can't ssh to the guest because of network problems" vs. "the guest is not alive and never pulled its ip" ? | |
| 15:08:34 | ralonsoh | dansmith, let me check, maybe in the syslog | |
| 15:08:52 | ralonsoh | understood, let me check | |
| 15:10:36 | ralonsoh | dansmith, neutron builds (adds/deletes) the leases file but we don't log this event. This is done by dnsmasq, you should be able to see that in syslog | |
| 15:10:47 | bauzas | oh good point | |
| 15:10:48 | sean-k-mooney | dansmith: i dont think neutron does but dnsmacq might | |
| 15:11:02 | bauzas | I forgot to look at dnsmasq, fucking shit | |
| 15:11:18 | bauzas | my ops skills become rusty | |
| 15:11:26 | dansmith | it's too bad because it might be a nice API to be able to poke that remotely.. i.e. instead of sshing forever, have a reasonably short timeout for the is-it-leased | |
| 15:11:33 | sean-k-mooney | bauzas: is this ovn | |
| 15:11:40 | fungi | so are the cirros kernel panics similar to one another, or random excuses? | |
| 15:11:42 | sean-k-mooney | because if its ovn we are not useing dnsmasq | |
| 15:11:54 | dansmith | and for reporting on failure, so we can say what forensics have been done | |
| 15:11:56 | sean-k-mooney | this is being handeled by openflow rules added by ovn | |
| 15:12:37 | dansmith | also, all three of those failed tests are volume-related | |
| 15:12:41 | sean-k-mooney | fungi: if they are related to acpi then its a know issue withthe cirrus 5.2 kernel | |
| 15:12:50 | dansmith | so I still wouldn't write-off it being a volume problem | |
| 15:13:02 | fungi | sean-k-mooney: sounds like a good reason to switch to 0.6.1 then | |
| 15:13:35 | dansmith | fungi: that's what I said, but 0.6.1 bringing other changes could be more destabilizing | |
| 15:13:37 | sean-k-mooney | fungi: i started working on alpine based image 2 years ago after i found out that the kernel bug was fixed in a later ubuntu kernel and cirro was just not updated | |
| 15:13:40 | ralonsoh | dansmith, what is the backend? OVS or OVN? | |
| 15:13:48 | dansmith | ralonsoh: I dunno | |
| 15:13:49 | ralonsoh | is this nova-next, right? | |
| 15:13:55 | dansmith | ralonsoh: yes | |
| 15:14:00 | ralonsoh | ok, let me check | |