Earlier  
Posted Nick Remark
#openstack-nova - 2020-09-18
12:20:31 sean-k-mooney this sound like its not a nova issue for what its worth
12:21:41 sean-k-mooney but check the qemu log to confim there is not reboot of the instace when it happened. if so then you need to look at the l2 agent logs if that does not show anyting then you need to check the dhcp agent and neutron server logs for the port uuid
12:21:50 sean-k-mooney to track down why the event was sent
12:26:33 ygk_12345 ok
12:28:19 ygk_12345 sean-k-mooney even when the vnics have disappeared, those ports are being shown as active in nova
12:28:28 ygk_12345 how can this be ?
12:31:10 sean-k-mooney without logs i cant really help you resolve this. can you file a bug and attach some logs specificly the nova compute and l2 agent logs for aroudn the time when this happened
12:31:49 ygk_12345 ok
12:32:02 sean-k-mooney what your describing does not really match to any known broken behviaor that comes to mind
12:32:43 sean-k-mooney the only way to revmoe an interface form a vm that is runnign other then a nova interface detach which is not present in the event logs
12:32:50 sean-k-mooney is to delete the neutron port
12:33:05 sean-k-mooney that will send a network-vif-deleted event to nova
12:33:19 sean-k-mooney but presuably you did not delete the nuetorn port
12:33:32 sean-k-mooney so im not aware of any code path that could result in this
12:33:52 ygk_12345 strange
12:34:03 sean-k-mooney the info cache currption issue requires a hard reboot for the instance to be lost or a move operations
12:44:59 ygk_12345 sean-k-mooney when the ping is unreachable for the vm , I dont see any messages for that port in the neutron-server except when I do a manual reboot
12:46:03 lyarwood ubuntu--
12:46:13 openstackgerrit Lee Yarwood proposed openstack/nova master: test_evacuate.sh: Support libvirt-bin and libvirtd systemd services https://review.opendev.org/752650
12:46:14 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_FILE_BACKED_DISCARD_VERSION https://review.opendev.org/746982
12:46:14 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Bump MIN_{LIBVIRT,QEMU}_VERSION and NEXT_MIN_{LIBVIRT,QEMU}_VERSION https://review.opendev.org/746981
12:46:15 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_BETTER_SIGKILL_HANDLING https://review.opendev.org/746984
12:46:15 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_{LIBVIRT,QEMU}_NATIVE_TLS_VERSION https://review.opendev.org/746983
12:46:16 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_{LIBVIRT,QEMU}_PMEM_SUPPORT https://review.opendev.org/746986
12:46:16 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_VIDEO_MODEL_VERSIONS https://review.opendev.org/746985
12:47:09 stephenfin lyarwood: https://media.tenor.com/images/a53dc07bf0a6f8822f16c0760299f915/tenor.gif
12:49:13 gibi :)
12:51:35 lyarwood yup sorry all
12:53:56 lyarwood back this evening to clean up any fallout caused by all of this
12:56:33 gibi ack
12:58:54 openstack Launchpad bug 1896226 in OpenStack Compute (nova) "The vnics are disappearing in the vm" [Undecided,New]
12:58:54 ygk_12345 sean-k-mooney i have posted this bug https://bugs.launchpad.net/nova/+bug/1896226
13:02:59 kashyap lyarwood: Hey, so do you know how many CPU cores did that reproducer host had?
13:03:26 kashyap lyarwood: I'm asking because, depending on that I may need to fork my CPU workload to spread across multiple (CPU) cores
13:05:10 kashyap Meanwhile ... lovely: I just created a fresh F32 VM on Ubuntu Focal, and the bootloader is filling up little yellow squares on my screen.
13:10:15 gibi kashyap: I think it is 8 https://zuul.opendev.org/t/openstack/build/4c56def513884c5eb3ba7b0adf7fa260/log/zuul-info/host-info.controller.yaml#440
13:11:01 gibi kashyap: at least this is the CI job run linked to the bug report
13:17:50 bauzas gibi: thanks
13:17:50 kashyap gibi: So that's the "host VM" where Tempest is running, yeaH?
13:18:00 kashyap gibi: I'm trying to recreate the _exact_ damn setup :D
13:18:03 bauzas (for the TC ping)
13:18:46 gibi kashyap: yes I think that is the "host" where tempest runs
13:18:58 sean-k-mooney kashyap: it has 8 i think
13:19:09 kashyap gibi: Okay, 8 vCPUs, and 7599 MB of memory
13:19:22 kashyap gibi: And do you see what's the disk size?
13:19:30 sean-k-mooney yep ist my small flavor
13:20:12 kashyap sean-k-mooney: The disk size is ~16GB?
13:20:13 gibi size: 80.00 GB
13:20:21 kashyap gibi: Oh, right
13:20:28 gibi https://zuul.opendev.org/t/openstack/build/4c56def513884c5eb3ba7b0adf7fa260/log/zuul-info/host-info.controller.yaml#259-293
13:20:39 kashyap Okay; I need to tweak the setup so that I'm also doing QEMU-on-KVM nested and not KVM-on-KVM nested
13:20:42 kashyap Thanks; /me back in a few
13:20:55 gibi thanks kashyap
13:22:20 openstack bug 1894804 in qemu (Ubuntu) "Second DEVICE_DELETED event missing during virtio-blk disk device detach" [Undecided,New] https://launchpad.net/bugs/1894804
13:22:20 openstackgerrit Merged openstack/nova master: releasenote: Add known issue for bug #1894804 https://review.opendev.org/752654
13:33:44 sean-k-mooney kashyap: you dont need 80G of disk by the way 25-40G is typeically more then enough
13:38:55 kashyap sean-k-mooney: Yeah, I have something like that anyway
13:40:21 kashyap sean-k-mooney: Do we have the full guest command-line of the Focal VM "host" upstream is using?
13:40:23 sean-k-mooney if you want an account on my home cloud by the way let me know and i can create one for you
13:40:46 sean-k-mooney no bug its just an openstack vm
13:41:27 sean-k-mooney with 8G of ram and 8 cpus and 80G fo disk
13:41:42 sean-k-mooney im not sure that really matters too much
13:42:02 kashyap sean-k-mooney: Okay; so what you gave me suffices
13:42:26 kashyap sean-k-mooney: For now, stable access (and it _is_ stable) to this VM is sufficient :) Thx!
13:42:29 sean-k-mooney ya my small flaovr is basicaly a proxy for the ci vms
13:43:19 sean-k-mooney except i add hw:cpu_sockets=2 hw:cpu_threads=2 hw:numa_nodes=2 and hw:mem_page_size=large
13:43:52 sean-k-mooney that should not change the behvior in this case but just an fyi
13:44:49 kashyap sean-k-mooney: One more: what's the CirrOS version being run here?
13:45:03 kashyap "here" as in, in the Tempest env
13:45:06 sean-k-mooney the default one form devstack so like 4.2 ish
13:45:23 kashyap I'd like to know the exact version, although it shouldn't matter _that_ much
13:45:25 sean-k-mooney 0.5.1
13:45:32 sean-k-mooney https://github.com/openstack/devstack/blob/master/stackrc#L670
13:45:42 kashyap Excellent
13:46:13 sean-k-mooney it hasnt been updated in 7 months https://github.com/openstack/devstack/commit/7a0fa4fd9e5db7253fee0820fc002703d43bca3c so that should be what its using
13:46:32 sean-k-mooney the fine i think shoudl still be in /opt/stack/data i think
14:08:19 kashyap Also, annoyingly enough, CirrOS from 0.4.0 onwards, it's only initrd/kernel, so I can't use 'guestfish' to look around in it: https://paste.centos.org/view/10bee1e1
14:21:36 openstack bug 1894966 in OpenStack Compute (nova) ussuri "Create servergroup failed with unexpected error" [Undecided,Confirmed] https://launchpad.net/bugs/1894966
14:21:36 openstackgerrit Stephen Finucane proposed openstack/nova stable/ussuri: tests: Add regression test for bug 1894966 https://review.opendev.org/752371
14:21:37 openstackgerrit Stephen Finucane proposed openstack/nova stable/ussuri: api: Set min, maxItems for server_group.policies field https://review.opendev.org/752702
14:31:58 openstack bug 1894966 in OpenStack Compute (nova) ussuri "Create servergroup failed with unexpected error" [Undecided,In progress] https://launchpad.net/bugs/1894966 - Assigned to Stephen Finucane (stephenfinucane)
14:31:58 openstackgerrit Stephen Finucane proposed openstack/nova stable/train: tests: Add regression test for bug 1894966 https://review.opendev.org/752706
14:31:59 openstackgerrit Stephen Finucane proposed openstack/nova stable/train: api: Set min, maxItems for server_group.policies field https://review.opendev.org/752707
14:46:50 mnaser i'm curious as to why nova needs to ssh for cold migrations if using shared storage
14:47:03 mnaser maybe no one just worked on that test path or?
14:47:26 mnaser aka if shared storage: skip ssh
14:54:24 gibi mnaser: I might be wrong but I remember nova detects that the instance is on share path by checking that the same file exists on both the source and the dest
14:54:44 mnaser gibi: yeah i think that might be the only reason it still does it
14:55:04 mnaser i think breaking out early if the VM is using volumes only would make sense cause that would eliminate the whole ssh key thing for us
14:55:09 mnaser the only reason we still have it is just for that
14:55:14 gmann gibi: bauzas lyarwood what is consensus for Focal migration? same as we discussed in meeting?
14:55:24 sean-k-mooney mnaser: we could do that via an rpc i guess
14:55:56 mnaser sean-k-mooney: can one n-cpu rpc to another n-cpui
14:56:03 sean-k-mooney yes
14:56:12 sean-k-mooney it happens alot in move operations
14:56:26 sean-k-mooney well not a lot but we do rpcs between teh computes invovled
14:57:12 gibi mnaser: hm, only volume instances still can have config drives
14:57:29 mnaser gibi: right but i think we dont actually move that we just rebuild it on the target with a cold migrate
14:57:30 mnaser i _think_
14:57:59 gibi most probably yes

Earlier   Later