Earlier  
Posted Nick Remark
#openstack-nova - 2020-10-30
15:23:33 melwitt stephenfin: argh ok, checking my headset settings
15:24:16 melwitt hm looks like it's all set right
15:24:37 lyarwood actually sounds really cool, I wouldn't worry
15:26:52 melwitt sounds cool xD heh
15:44:41 lyarwood gmann: https://review.opendev.org/#/c/755525/ - can you remove your -W on this btw, the `fix` landed in stable/victoria
15:45:01 openstackgerrit Lee Yarwood proposed openstack/nova stable/ussuri: libvirt: Increase incremental and max sleep time during device detach https://review.opendev.org/757306
15:45:57 gmann lyarwood: done,
15:53:38 lyarwood gmann: thanks
16:19:46 johnsom Hi Nova community. The Octavia team is seeing some strange behavior with libvirtd. https://storage.gra.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_c77/759973/3/check/octavia-v2-dsvm-scenario/c77fe63/controller/logs/libvirt/libvirtd_log.txt
16:19:52 johnsom Warning, that is a 136MB log file
16:20:08 johnsom It seems to go into some sort of tight loop:
16:20:26 lyarwood ops, opening that in chrome was a mistake
16:20:30 johnsom https://www.irccloud.com/pastebin/c6NcVeeF/
16:20:54 johnsom Almost 90% of that log file is the above
16:20:59 lyarwood kashyap: ^ if you have time
16:21:05 lyarwood ring any bells?
16:21:38 johnsom Eventually it seems to finally exit. The issue we have is while it's stuck in this loop other VMs are not going ACTIVE in nova, so causes all sorts of timeouts.
16:21:39 kashyap lyarwood: Just a sec; lost in two other convo threads :-(
16:22:15 lyarwood kashyap: np
16:22:26 kashyap johnsom: As a quick question - in what scenario do you see that?
16:22:33 lyarwood johnsom: might be useful to write that up in a nova bug that kashyap can follow up with
16:22:37 lyarwood ops sorry
16:22:38 kashyap johnsom: That "End of file" ... simply means libvirt lost connection to QEMU
16:22:52 johnsom It appears that nova was trying to delete the instance
16:22:55 kashyap lyarwood: No worries; a bug also helps - as a record
16:23:45 johnsom Here is a chunk of the log right before the loop:
16:23:48 johnsom https://www.irccloud.com/pastebin/LRErlmPf/
16:26:35 kashyap johnsom: Yeah, but that doesn't tell us what events led to that (they're in the 130MB log file). :-) Is this reproducible?
16:27:12 johnsom It is reproducible, but intermittent. Not every job triggers it.
16:28:05 gibi #nova now refresh connection_info / avoid storing stale connection_info in Nova
16:28:54 johnsom We are seeing it while we run our scenario test suites.
16:30:06 kashyap johnsom: A write-up in a bug would be nice to track. Along w/ a description of a test case that's hitting it. (I'm almost out of neurons tonight, and my mind now has the "thousand yard stare")
16:31:10 johnsom Yeah, I'm writing one up. I will link to the long, but I don't think I should attach it given it's size, so it will expire off the object store.
16:31:18 johnsom I will paste in a few snippets.
16:31:56 kashyap You can xz-compress the log; as the CI one might get cleaned up any moment
16:32:35 kashyap johnsom: Also, paste-bins expore; please add relevant snippets as a _text_ file. Because, LaunchPad, annoyingly enough, breaks formatting
16:32:46 kashyap Thanks :)
16:33:12 johnsom Ok
16:40:17 kashyap johnsom: (Sorry for being pedanctic.)
16:40:38 johnsom Oh no worries. I want to make sure you get what you need to help us out.
16:54:51 johnsom kashyap https://bugs.launchpad.net/nova/+bug/1902276
16:54:51 openstack Launchpad bug 1902276 in OpenStack Compute (nova) "libvirtd going into a tight loop causing instances to not transition to ACTIVE" [Undecided,New]
16:59:59 kashyap johnsom: Thanks; another question - so the instances are not coming up online?
17:02:11 kashyap johnsom: Disregard me; you actually say that "eventually things go back to normal"
17:03:15 johnsom Yep
17:10:24 kashyap johnsom: I'm a bit baffled with it all :) I asked a couple of libvirt devs to see if the lead in and the end of the loop makes any further sense to them
17:10:43 johnsom Ok, thank you
17:15:46 kashyap johnsom: Haha, you write: "We seen this regularly, but it's intermittent" — now, which is it?
17:16:19 johnsom kashyap We see it on jobs daily, but it's not every run
17:16:54 johnsom Maybe that makes more sense? grin
17:18:10 johnsom It's been a long week, words are hard. grin
17:18:11 kashyap Ah, yes. That makes more sense
17:18:24 kashyap But something seems fishy; what changed so that things "go back to normal"
17:18:33 kashyap To be continued ... next week :)
17:18:34 johnsom Yeah, I have no idea
17:20:43 johnsom It repeats 276713 times in the log, so not a nice even number really.
18:10:35 melwitt dansmith: train backport of cherry pick check fix if you get a moment https://review.opendev.org/759118
22:02:10 openstackgerrit Merged openstack/nova stable/train: Follow up for cherry-pick check for merge patch https://review.opendev.org/759118
22:22:45 openstackgerrit melanie witt proposed openstack/nova stable/stein: Follow up for cherry-pick check for merge patch https://review.opendev.org/760673
#openstack-nova - 2020-10-31
06:31:41 openstackgerrit Merged openstack/nova master: Fix virsh domifstat to get vhostuser vif statistics https://review.opendev.org/757448
07:38:23 openstackgerrit Takashi Natsume proposed openstack/nova master: Update contributor guide for Wallaby https://review.opendev.org/754427
08:14:33 openstackgerrit Merged openstack/nova master: Add regression test for bug #1899649 https://review.opendev.org/757893
08:14:33 openstack bug 1899649 in OpenStack Compute (nova) "Volume marked as available after a failure to build" [Undecided,In progress] https://launchpad.net/bugs/1899649 - Assigned to Lee Yarwood (lyarwood)
08:26:21 openstackgerrit zhanghao proposed openstack/nova stable/victoria: Fix virsh domifstat to get vhostuser vif statistics https://review.opendev.org/760684
08:28:38 openstackgerrit Jorhson Deng proposed openstack/nova master: limit the instance to backup while task state is not None https://review.opendev.org/760685
10:54:34 noonedeadpunk hey, everyone. still having issue with isolated aggregates here... I've just got instances spawned on the isolated aggregate without any traits provided. and I ythink I have like 1-2 vms being spawned that way per month, so it's not something that I can reproduce but somehow that happens
10:56:20 noonedeadpunk nothing to look at in logs, but I don't have debug enabled since it prod env....
11:20:22 openstackgerrit Merged openstack/nova master: objects: Fix issue in exception type https://review.opendev.org/756069
11:47:08 noonedeadpunk can it be because of https://opendev.org/openstack/nova/src/branch/master/nova/scheduler/manager.py#L146 ?
11:47:38 noonedeadpunk so when user clicks rebuild instance it gets spawned on the isolated aggregate?
11:51:11 noonedeadpunk any reason not to use placement for rebuilds?
17:28:44 legochen hi nova experts, do nova support us to set instance/cpu/mem/disk quota by AZ.
17:29:59 legochen The AZ is compute resource pool that visible to end users. Users use AZ to provision VMs in a specific environment or compute resources.
17:31:24 legochen Currently, seems like we only can set nova quota globally, like I only allow projectA users to create 10 instances in AZ A, but they still can use this quota to create instances in AZ B.
17:54:08 gmann dansmith: this is ready. glance's copy_image policy things we discussed on qa channel. - https://review.opendev.org/#/c/760422/1
21:27:53 prometheanfire what do nova people think about updating mock? https://review.opendev.org/760057
#openstack-nova - 2020-11-02
06:26:03 openstackgerrit Nobuhiro MIKI proposed openstack/nova-specs master: Add IP address to libvirt guest metadata https://review.opendev.org/760750
06:54:45 openstackgerrit Nobuhiro MIKI proposed openstack/nova-specs master: Add IP address to libvirt guest metadata https://review.opendev.org/760750
08:35:15 bauzas good morning Nova
08:35:57 gibi good morning
08:45:56 lyarwood Morning
08:48:42 bauzas fwiw, our sanitary protocols become stricter for schools (and my kids are back in school after 2 weeks), so I'll need to disappear like 30 mins every 3 hours
09:45:19 alex_xu gibi: stephenfin, If I have a feature have to do numa affinity, then we must implement numa in placement first, right? As my understand, the two phase commit is only for resolving the race problem, we still need back to numa in placement when we need affinity, is it right?
09:52:34 gibi alex_xu: I don't think this is a must any more. I think we agreed that we all would prefere having the numa modelled in placement. but it did not happened in the last cycles. so we are not sure it will ever happen. If you have time and willingnes to work on the numa in placement implementation then I think that is welcomed.
09:53:19 gibi alex_xu: but if you propose a numa affinity feature without numa in placement then it is also OK at least to discuss
09:54:28 alex_xu gibi: yea, I see that, I will get that back to team and pm, see if we have resource
09:54:38 alex_xu gibi: thanks, I clear the path now
09:54:51 gibi OK, cool
10:56:59 hemanth_n Hi Nova folks, I am looking for the following review to get merged https://review.opendev.org/#/c/749175/ unless there is something else to take care of.. it is already +2ed couple of time
10:58:20 openstackgerrit Balazs Gibizer proposed openstack/nova master: Add functional-py39 tox target https://review.opendev.org/760884
11:34:49 gibi sean-k-mooney: hi! it is just me or the os-vif unit test os_vif.tests.unit.internal.ip.linux.test_impl_pyroute2.TestIpCommand.test_add_exit_code fails consistently (at least locally for me)
11:37:57 openstackgerrit Balazs Gibizer proposed openstack/os-vif master: DNM https://review.opendev.org/760891
11:39:10 gibi sean-k-mooney: never mind, it was a local issue for me
11:51:58 slaweq sean-k-mooney: gibi: hi, can You take a look at https://bugs.launchpad.net/nova/+bug/1902516 when You will have few minutes? I saw it at least couple of times in the ci in last few days, maybe You know already what is the root cause of it :)
11:51:58 openstack Launchpad bug 1902516 in OpenStack Compute (nova) "[CI] Libvirt error "TCG doesn't support requested feature: CPUID.01H:ECX.vmx" cause failure while spawning instance" [Undecided,New]
11:54:28 gibi slaweq: ack, I haven't seen this yet
11:54:38 slaweq gibi: thx
11:58:10 gibi slaweq: thanks for reporting it
11:58:49 slaweq gibi: yw :)
13:46:23 lyarwood gibi: https://review.opendev.org/#/c/758928/ - would you mind taking a look at this when you have a chance? melwitt++

Earlier   Later