| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-10-30 | |||
| 16:20:59 | lyarwood | kashyap: ^ if you have time | |
| 16:21:05 | lyarwood | ring any bells? | |
| 16:21:38 | johnsom | Eventually it seems to finally exit. The issue we have is while it's stuck in this loop other VMs are not going ACTIVE in nova, so causes all sorts of timeouts. | |
| 16:21:39 | kashyap | lyarwood: Just a sec; lost in two other convo threads :-( | |
| 16:22:15 | lyarwood | kashyap: np | |
| 16:22:26 | kashyap | johnsom: As a quick question - in what scenario do you see that? | |
| 16:22:33 | lyarwood | johnsom: might be useful to write that up in a nova bug that kashyap can follow up with | |
| 16:22:37 | lyarwood | ops sorry | |
| 16:22:38 | kashyap | johnsom: That "End of file" ... simply means libvirt lost connection to QEMU | |
| 16:22:52 | johnsom | It appears that nova was trying to delete the instance | |
| 16:22:55 | kashyap | lyarwood: No worries; a bug also helps - as a record | |
| 16:23:45 | johnsom | Here is a chunk of the log right before the loop: | |
| 16:23:48 | johnsom | https://www.irccloud.com/pastebin/LRErlmPf/ | |
| 16:26:35 | kashyap | johnsom: Yeah, but that doesn't tell us what events led to that (they're in the 130MB log file). :-) Is this reproducible? | |
| 16:27:12 | johnsom | It is reproducible, but intermittent. Not every job triggers it. | |
| 16:28:05 | gibi | #nova now refresh connection_info / avoid storing stale connection_info in Nova | |
| 16:28:54 | johnsom | We are seeing it while we run our scenario test suites. | |
| 16:30:06 | kashyap | johnsom: A write-up in a bug would be nice to track. Along w/ a description of a test case that's hitting it. (I'm almost out of neurons tonight, and my mind now has the "thousand yard stare") | |
| 16:31:10 | johnsom | Yeah, I'm writing one up. I will link to the long, but I don't think I should attach it given it's size, so it will expire off the object store. | |
| 16:31:18 | johnsom | I will paste in a few snippets. | |
| 16:31:56 | kashyap | You can xz-compress the log; as the CI one might get cleaned up any moment | |
| 16:32:35 | kashyap | johnsom: Also, paste-bins expore; please add relevant snippets as a _text_ file. Because, LaunchPad, annoyingly enough, breaks formatting | |
| 16:32:46 | kashyap | Thanks :) | |
| 16:33:12 | johnsom | Ok | |
| 16:40:17 | kashyap | johnsom: (Sorry for being pedanctic.) | |
| 16:40:38 | johnsom | Oh no worries. I want to make sure you get what you need to help us out. | |
| 16:54:51 | johnsom | kashyap https://bugs.launchpad.net/nova/+bug/1902276 | |
| 16:54:51 | openstack | Launchpad bug 1902276 in OpenStack Compute (nova) "libvirtd going into a tight loop causing instances to not transition to ACTIVE" [Undecided,New] | |
| 16:59:59 | kashyap | johnsom: Thanks; another question - so the instances are not coming up online? | |
| 17:02:11 | kashyap | johnsom: Disregard me; you actually say that "eventually things go back to normal" | |
| 17:03:15 | johnsom | Yep | |
| 17:10:24 | kashyap | johnsom: I'm a bit baffled with it all :) I asked a couple of libvirt devs to see if the lead in and the end of the loop makes any further sense to them | |
| 17:10:43 | johnsom | Ok, thank you | |
| 17:15:46 | kashyap | johnsom: Haha, you write: "We seen this regularly, but it's intermittent" — now, which is it? | |
| 17:16:19 | johnsom | kashyap We see it on jobs daily, but it's not every run | |
| 17:16:54 | johnsom | Maybe that makes more sense? grin | |
| 17:18:10 | johnsom | It's been a long week, words are hard. grin | |
| 17:18:11 | kashyap | Ah, yes. That makes more sense | |
| 17:18:24 | kashyap | But something seems fishy; what changed so that things "go back to normal" | |
| 17:18:33 | kashyap | To be continued ... next week :) | |
| 17:18:34 | johnsom | Yeah, I have no idea | |
| 17:20:43 | johnsom | It repeats 276713 times in the log, so not a nice even number really. | |
| 18:10:35 | melwitt | dansmith: train backport of cherry pick check fix if you get a moment https://review.opendev.org/759118 | |
| 22:02:10 | openstackgerrit | Merged openstack/nova stable/train: Follow up for cherry-pick check for merge patch https://review.opendev.org/759118 | |
| 22:22:45 | openstackgerrit | melanie witt proposed openstack/nova stable/stein: Follow up for cherry-pick check for merge patch https://review.opendev.org/760673 | |
| #openstack-nova - 2020-10-31 | |||
| 06:31:41 | openstackgerrit | Merged openstack/nova master: Fix virsh domifstat to get vhostuser vif statistics https://review.opendev.org/757448 | |
| 07:38:23 | openstackgerrit | Takashi Natsume proposed openstack/nova master: Update contributor guide for Wallaby https://review.opendev.org/754427 | |
| 08:14:33 | openstackgerrit | Merged openstack/nova master: Add regression test for bug #1899649 https://review.opendev.org/757893 | |
| 08:14:33 | openstack | bug 1899649 in OpenStack Compute (nova) "Volume marked as available after a failure to build" [Undecided,In progress] https://launchpad.net/bugs/1899649 - Assigned to Lee Yarwood (lyarwood) | |
| 08:26:21 | openstackgerrit | zhanghao proposed openstack/nova stable/victoria: Fix virsh domifstat to get vhostuser vif statistics https://review.opendev.org/760684 | |
| 08:28:38 | openstackgerrit | Jorhson Deng proposed openstack/nova master: limit the instance to backup while task state is not None https://review.opendev.org/760685 | |
| 10:54:34 | noonedeadpunk | hey, everyone. still having issue with isolated aggregates here... I've just got instances spawned on the isolated aggregate without any traits provided. and I ythink I have like 1-2 vms being spawned that way per month, so it's not something that I can reproduce but somehow that happens | |
| 10:56:20 | noonedeadpunk | nothing to look at in logs, but I don't have debug enabled since it prod env.... | |
| 11:20:22 | openstackgerrit | Merged openstack/nova master: objects: Fix issue in exception type https://review.opendev.org/756069 | |
| 11:47:08 | noonedeadpunk | can it be because of https://opendev.org/openstack/nova/src/branch/master/nova/scheduler/manager.py#L146 ? | |
| 11:47:38 | noonedeadpunk | so when user clicks rebuild instance it gets spawned on the isolated aggregate? | |
| 11:51:11 | noonedeadpunk | any reason not to use placement for rebuilds? | |
| 17:28:44 | legochen | hi nova experts, do nova support us to set instance/cpu/mem/disk quota by AZ. | |
| 17:29:59 | legochen | The AZ is compute resource pool that visible to end users. Users use AZ to provision VMs in a specific environment or compute resources. | |
| 17:31:24 | legochen | Currently, seems like we only can set nova quota globally, like I only allow projectA users to create 10 instances in AZ A, but they still can use this quota to create instances in AZ B. | |
| 17:54:08 | gmann | dansmith: this is ready. glance's copy_image policy things we discussed on qa channel. - https://review.opendev.org/#/c/760422/1 | |
| 21:27:53 | prometheanfire | what do nova people think about updating mock? https://review.opendev.org/760057 | |
| #openstack-nova - 2020-11-02 | |||
| 06:26:03 | openstackgerrit | Nobuhiro MIKI proposed openstack/nova-specs master: Add IP address to libvirt guest metadata https://review.opendev.org/760750 | |
| 06:54:45 | openstackgerrit | Nobuhiro MIKI proposed openstack/nova-specs master: Add IP address to libvirt guest metadata https://review.opendev.org/760750 | |
| 08:35:15 | bauzas | good morning Nova | |
| 08:35:57 | gibi | good morning | |
| 08:45:56 | lyarwood | Morning | |
| 08:48:42 | bauzas | fwiw, our sanitary protocols become stricter for schools (and my kids are back in school after 2 weeks), so I'll need to disappear like 30 mins every 3 hours | |
| 09:45:19 | alex_xu | gibi: stephenfin, If I have a feature have to do numa affinity, then we must implement numa in placement first, right? As my understand, the two phase commit is only for resolving the race problem, we still need back to numa in placement when we need affinity, is it right? | |
| 09:52:34 | gibi | alex_xu: I don't think this is a must any more. I think we agreed that we all would prefere having the numa modelled in placement. but it did not happened in the last cycles. so we are not sure it will ever happen. If you have time and willingnes to work on the numa in placement implementation then I think that is welcomed. | |
| 09:53:19 | gibi | alex_xu: but if you propose a numa affinity feature without numa in placement then it is also OK at least to discuss | |
| 09:54:28 | alex_xu | gibi: yea, I see that, I will get that back to team and pm, see if we have resource | |
| 09:54:38 | alex_xu | gibi: thanks, I clear the path now | |
| 09:54:51 | gibi | OK, cool | |
| 10:56:59 | hemanth_n | Hi Nova folks, I am looking for the following review to get merged https://review.opendev.org/#/c/749175/ unless there is something else to take care of.. it is already +2ed couple of time | |
| 10:58:20 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add functional-py39 tox target https://review.opendev.org/760884 | |
| 11:34:49 | gibi | sean-k-mooney: hi! it is just me or the os-vif unit test os_vif.tests.unit.internal.ip.linux.test_impl_pyroute2.TestIpCommand.test_add_exit_code fails consistently (at least locally for me) | |
| 11:37:57 | openstackgerrit | Balazs Gibizer proposed openstack/os-vif master: DNM https://review.opendev.org/760891 | |
| 11:39:10 | gibi | sean-k-mooney: never mind, it was a local issue for me | |
| 11:51:58 | slaweq | sean-k-mooney: gibi: hi, can You take a look at https://bugs.launchpad.net/nova/+bug/1902516 when You will have few minutes? I saw it at least couple of times in the ci in last few days, maybe You know already what is the root cause of it :) | |
| 11:51:58 | openstack | Launchpad bug 1902516 in OpenStack Compute (nova) "[CI] Libvirt error "TCG doesn't support requested feature: CPUID.01H:ECX.vmx" cause failure while spawning instance" [Undecided,New] | |
| 11:54:28 | gibi | slaweq: ack, I haven't seen this yet | |
| 11:54:38 | slaweq | gibi: thx | |
| 11:58:10 | gibi | slaweq: thanks for reporting it | |
| 11:58:49 | slaweq | gibi: yw :) | |
| 13:46:23 | lyarwood | gibi: https://review.opendev.org/#/c/758928/ - would you mind taking a look at this when you have a chance? melwitt++ | |
| 13:50:22 | gibi | lyarwood: will check | |
| 13:51:16 | lyarwood | many thanks | |
| 13:58:02 | gmann | gibi: it will be good to add gate job too for py3.9 along with tox env. what you say? | |
| 13:58:07 | gmann | this one https://review.opendev.org/#/c/760884/1 | |
| 13:59:07 | gibi | gmann: ohh, now I got efried's comment in the review. Sure I can try to add functional jobs to the zuul.yaml | |
| 13:59:36 | gmann | because functional jobs does not come form generic template, it is added in project side explicitly | |
| 13:59:43 | gmann | thanks | |
| 13:59:49 | efried | ^ | |
| 14:00:48 | gibi | efried: hey! can I say that welcome back? | |
| 14:01:14 | efried | Afraid not. I just happened to see your email in the ML and thought, "Hey, there's an easy review for me!" | |
| 14:01:20 | gibi | :) | |
| 14:01:23 | gibi | how is life? | |
| 14:01:34 | efried | Life is *good*. | |
| 14:02:05 | gibi | I'm glad to hear that | |