| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-09-18 | |||
| 12:15:50 | sean-k-mooney | simialrly if the admin state on the newon port was changed it would also show up there | |
| 12:16:38 | sean-k-mooney | ygk_12345: the other thing that you cold check before going into the agent logs is the qemu instance log | |
| 12:16:48 | ygk_12345 | oh ok | |
| 12:16:50 | sean-k-mooney | and confirm that there are no restarts | |
| 12:17:31 | sean-k-mooney | ygk_12345: do you enabel the watchdog or qemu guest agent in teh image that is affected | |
| 12:18:00 | ygk_12345 | no idea | |
| 12:18:15 | sean-k-mooney | they would be listed in the image metadata in glance | |
| 12:20:00 | sean-k-mooney | ygk_12345: those event typically mean 1 of two things. eighet the tap device was removed and added to ovs, or the port was updated (host-id or admin status) | |
| 12:20:10 | ygk_12345 | i dont think they are set for that image | |
| 12:20:17 | sean-k-mooney | ok | |
| 12:20:31 | sean-k-mooney | this sound like its not a nova issue for what its worth | |
| 12:21:41 | sean-k-mooney | but check the qemu log to confim there is not reboot of the instace when it happened. if so then you need to look at the l2 agent logs if that does not show anyting then you need to check the dhcp agent and neutron server logs for the port uuid | |
| 12:21:50 | sean-k-mooney | to track down why the event was sent | |
| 12:26:33 | ygk_12345 | ok | |
| 12:28:19 | ygk_12345 | sean-k-mooney even when the vnics have disappeared, those ports are being shown as active in nova | |
| 12:28:28 | ygk_12345 | how can this be ? | |
| 12:31:10 | sean-k-mooney | without logs i cant really help you resolve this. can you file a bug and attach some logs specificly the nova compute and l2 agent logs for aroudn the time when this happened | |
| 12:31:49 | ygk_12345 | ok | |
| 12:32:02 | sean-k-mooney | what your describing does not really match to any known broken behviaor that comes to mind | |
| 12:32:43 | sean-k-mooney | the only way to revmoe an interface form a vm that is runnign other then a nova interface detach which is not present in the event logs | |
| 12:32:50 | sean-k-mooney | is to delete the neutron port | |
| 12:33:05 | sean-k-mooney | that will send a network-vif-deleted event to nova | |
| 12:33:19 | sean-k-mooney | but presuably you did not delete the nuetorn port | |
| 12:33:32 | sean-k-mooney | so im not aware of any code path that could result in this | |
| 12:33:52 | ygk_12345 | strange | |
| 12:34:03 | sean-k-mooney | the info cache currption issue requires a hard reboot for the instance to be lost or a move operations | |
| 12:44:59 | ygk_12345 | sean-k-mooney when the ping is unreachable for the vm , I dont see any messages for that port in the neutron-server except when I do a manual reboot | |
| 12:46:03 | lyarwood | ubuntu-- | |
| 12:46:13 | openstackgerrit | Lee Yarwood proposed openstack/nova master: test_evacuate.sh: Support libvirt-bin and libvirtd systemd services https://review.opendev.org/752650 | |
| 12:46:14 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_FILE_BACKED_DISCARD_VERSION https://review.opendev.org/746982 | |
| 12:46:14 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Bump MIN_{LIBVIRT,QEMU}_VERSION and NEXT_MIN_{LIBVIRT,QEMU}_VERSION https://review.opendev.org/746981 | |
| 12:46:15 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_BETTER_SIGKILL_HANDLING https://review.opendev.org/746984 | |
| 12:46:15 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_{LIBVIRT,QEMU}_NATIVE_TLS_VERSION https://review.opendev.org/746983 | |
| 12:46:16 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_{LIBVIRT,QEMU}_PMEM_SUPPORT https://review.opendev.org/746986 | |
| 12:46:16 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Remove MIN_LIBVIRT_VIDEO_MODEL_VERSIONS https://review.opendev.org/746985 | |
| 12:47:09 | stephenfin | lyarwood: https://media.tenor.com/images/a53dc07bf0a6f8822f16c0760299f915/tenor.gif | |
| 12:49:13 | gibi | :) | |
| 12:51:35 | lyarwood | yup sorry all | |
| 12:53:56 | lyarwood | back this evening to clean up any fallout caused by all of this | |
| 12:56:33 | gibi | ack | |
| 12:58:54 | openstack | Launchpad bug 1896226 in OpenStack Compute (nova) "The vnics are disappearing in the vm" [Undecided,New] | |
| 12:58:54 | ygk_12345 | sean-k-mooney i have posted this bug https://bugs.launchpad.net/nova/+bug/1896226 | |
| 13:02:59 | kashyap | lyarwood: Hey, so do you know how many CPU cores did that reproducer host had? | |
| 13:03:26 | kashyap | lyarwood: I'm asking because, depending on that I may need to fork my CPU workload to spread across multiple (CPU) cores | |
| 13:05:10 | kashyap | Meanwhile ... lovely: I just created a fresh F32 VM on Ubuntu Focal, and the bootloader is filling up little yellow squares on my screen. | |
| 13:10:15 | gibi | kashyap: I think it is 8 https://zuul.opendev.org/t/openstack/build/4c56def513884c5eb3ba7b0adf7fa260/log/zuul-info/host-info.controller.yaml#440 | |
| 13:11:01 | gibi | kashyap: at least this is the CI job run linked to the bug report | |
| 13:17:50 | bauzas | gibi: thanks | |
| 13:17:50 | kashyap | gibi: So that's the "host VM" where Tempest is running, yeaH? | |
| 13:18:00 | kashyap | gibi: I'm trying to recreate the _exact_ damn setup :D | |
| 13:18:03 | bauzas | (for the TC ping) | |
| 13:18:46 | gibi | kashyap: yes I think that is the "host" where tempest runs | |
| 13:18:58 | sean-k-mooney | kashyap: it has 8 i think | |
| 13:19:09 | kashyap | gibi: Okay, 8 vCPUs, and 7599 MB of memory | |
| 13:19:22 | kashyap | gibi: And do you see what's the disk size? | |
| 13:19:30 | sean-k-mooney | yep ist my small flavor | |
| 13:20:12 | kashyap | sean-k-mooney: The disk size is ~16GB? | |
| 13:20:13 | gibi | size: 80.00 GB | |
| 13:20:21 | kashyap | gibi: Oh, right | |
| 13:20:28 | gibi | https://zuul.opendev.org/t/openstack/build/4c56def513884c5eb3ba7b0adf7fa260/log/zuul-info/host-info.controller.yaml#259-293 | |
| 13:20:39 | kashyap | Okay; I need to tweak the setup so that I'm also doing QEMU-on-KVM nested and not KVM-on-KVM nested | |
| 13:20:42 | kashyap | Thanks; /me back in a few | |
| 13:20:55 | gibi | thanks kashyap | |
| 13:22:20 | openstack | bug 1894804 in qemu (Ubuntu) "Second DEVICE_DELETED event missing during virtio-blk disk device detach" [Undecided,New] https://launchpad.net/bugs/1894804 | |
| 13:22:20 | openstackgerrit | Merged openstack/nova master: releasenote: Add known issue for bug #1894804 https://review.opendev.org/752654 | |
| 13:33:44 | sean-k-mooney | kashyap: you dont need 80G of disk by the way 25-40G is typeically more then enough | |
| 13:38:55 | kashyap | sean-k-mooney: Yeah, I have something like that anyway | |
| 13:40:21 | kashyap | sean-k-mooney: Do we have the full guest command-line of the Focal VM "host" upstream is using? | |
| 13:40:23 | sean-k-mooney | if you want an account on my home cloud by the way let me know and i can create one for you | |
| 13:40:46 | sean-k-mooney | no bug its just an openstack vm | |
| 13:41:27 | sean-k-mooney | with 8G of ram and 8 cpus and 80G fo disk | |
| 13:41:42 | sean-k-mooney | im not sure that really matters too much | |
| 13:42:02 | kashyap | sean-k-mooney: Okay; so what you gave me suffices | |
| 13:42:26 | kashyap | sean-k-mooney: For now, stable access (and it _is_ stable) to this VM is sufficient :) Thx! | |
| 13:42:29 | sean-k-mooney | ya my small flaovr is basicaly a proxy for the ci vms | |
| 13:43:19 | sean-k-mooney | except i add hw:cpu_sockets=2 hw:cpu_threads=2 hw:numa_nodes=2 and hw:mem_page_size=large | |
| 13:43:52 | sean-k-mooney | that should not change the behvior in this case but just an fyi | |
| 13:44:49 | kashyap | sean-k-mooney: One more: what's the CirrOS version being run here? | |
| 13:45:03 | kashyap | "here" as in, in the Tempest env | |
| 13:45:06 | sean-k-mooney | the default one form devstack so like 4.2 ish | |
| 13:45:23 | kashyap | I'd like to know the exact version, although it shouldn't matter _that_ much | |
| 13:45:25 | sean-k-mooney | 0.5.1 | |
| 13:45:32 | sean-k-mooney | https://github.com/openstack/devstack/blob/master/stackrc#L670 | |
| 13:45:42 | kashyap | Excellent | |
| 13:46:13 | sean-k-mooney | it hasnt been updated in 7 months https://github.com/openstack/devstack/commit/7a0fa4fd9e5db7253fee0820fc002703d43bca3c so that should be what its using | |
| 13:46:32 | sean-k-mooney | the fine i think shoudl still be in /opt/stack/data i think | |
| 14:08:19 | kashyap | Also, annoyingly enough, CirrOS from 0.4.0 onwards, it's only initrd/kernel, so I can't use 'guestfish' to look around in it: https://paste.centos.org/view/10bee1e1 | |
| 14:21:36 | openstack | bug 1894966 in OpenStack Compute (nova) ussuri "Create servergroup failed with unexpected error" [Undecided,Confirmed] https://launchpad.net/bugs/1894966 | |
| 14:21:36 | openstackgerrit | Stephen Finucane proposed openstack/nova stable/ussuri: tests: Add regression test for bug 1894966 https://review.opendev.org/752371 | |
| 14:21:37 | openstackgerrit | Stephen Finucane proposed openstack/nova stable/ussuri: api: Set min, maxItems for server_group.policies field https://review.opendev.org/752702 | |
| 14:31:58 | openstack | bug 1894966 in OpenStack Compute (nova) ussuri "Create servergroup failed with unexpected error" [Undecided,In progress] https://launchpad.net/bugs/1894966 - Assigned to Stephen Finucane (stephenfinucane) | |
| 14:31:58 | openstackgerrit | Stephen Finucane proposed openstack/nova stable/train: tests: Add regression test for bug 1894966 https://review.opendev.org/752706 | |
| 14:31:59 | openstackgerrit | Stephen Finucane proposed openstack/nova stable/train: api: Set min, maxItems for server_group.policies field https://review.opendev.org/752707 | |
| 14:46:50 | mnaser | i'm curious as to why nova needs to ssh for cold migrations if using shared storage | |
| 14:47:03 | mnaser | maybe no one just worked on that test path or? | |
| 14:47:26 | mnaser | aka if shared storage: skip ssh | |
| 14:54:24 | gibi | mnaser: I might be wrong but I remember nova detects that the instance is on share path by checking that the same file exists on both the source and the dest | |
| 14:54:44 | mnaser | gibi: yeah i think that might be the only reason it still does it | |
| 14:55:04 | mnaser | i think breaking out early if the VM is using volumes only would make sense cause that would eliminate the whole ssh key thing for us | |
| 14:55:09 | mnaser | the only reason we still have it is just for that | |