| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-03-22 | |||
| 11:22:18 | ralonsoh | I don't mind, there is no problem | |
| 14:04:55 | opendevreview | Merged openstack/nova stable/zed: Reproducer for bug 1951656 https://review.opendev.org/c/openstack/nova/+/866151 | |
| 15:05:24 | opendevreview | René Ribaud proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| 15:05:24 | opendevreview | René Ribaud proposed openstack/nova master: Reproducers for bug 1869804 https://review.opendev.org/c/openstack/nova/+/877772 | |
| 15:33:55 | opendevreview | Merged openstack/nova stable/zed: Handle mdev devices in libvirt 7.7+ https://review.opendev.org/c/openstack/nova/+/866152 | |
| 16:41:32 | opendevreview | Dan Smith proposed openstack/nova master: Make scheduler lazy-load the placement client https://review.opendev.org/c/openstack/nova/+/878238 | |
| 17:47:03 | dansmith | bauzas: do you think we need to do much discussing of this during the ptg? https://review.opendev.org/c/openstack/nova-specs/+/877291 | |
| 17:47:12 | dansmith | or can/should we try to get it approved earlier? | |
| 17:47:27 | dansmith | it's basically what we already have in the backlog spec for the next step | |
| 17:47:47 | bauzas | dansmith: I can do a round of reviews tomorrow | |
| 17:47:53 | dansmith | ack thanks | |
| 17:48:02 | bauzas | I actually *should* do it before the PTG anyway | |
| 17:48:19 | dansmith | yeah, I was thinking it might be good to better surface what actually needs a lot of discussion | |
| 18:55:58 | gmann | dansmith: 1 comment for skip-level-always job change (in case you did not see) https://review.opendev.org/c/openstack/nova/+/875773 | |
| 18:56:13 | dansmith | gmann: ack will look later | |
| 18:56:20 | gmann | sure, thanks | |
| 18:58:25 | opendevreview | René Ribaud proposed openstack/nova master: Reproducers for bug 1869804 https://review.opendev.org/c/openstack/nova/+/877772 | |
| 18:58:26 | opendevreview | René Ribaud proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| #openstack-nova - 2023-03-23 | |||
| 09:16:02 | auniyal | Hi sean-k-mooney | |
| 09:16:09 | auniyal | can you please review this - https://review.opendev.org/c/openstack/nova/+/790447/ | |
| 09:44:53 | zigo | Is it known? Is there a way to fix? Is it related to the version of libvirt or qemu? | |
| 09:44:53 | zigo | https://paste.opendev.org/show/bYlTfz7fxnQtzVhpVf91/ | |
| 09:44:53 | zigo | I'm currently doing routine upgrade of compute nodes in a cluster (running Victoria), and I'm getting live-migration errors of VMs like this one: | |
| 09:44:53 | zigo | Hi there! | |
| 09:49:53 | bauzas | looking | |
| 09:52:47 | bauzas | zigo: good question I guess you've seen the libvirt error | |
| 09:52:49 | bauzas | 2023-03-23 09:37:52.682 3209246 INFO nova.compute.manager [req-5fda88d8-510a-4943-9703-9b47e865a89f - - - - -] [instance: 17672112-c416-494a-88f8-fd7cfa85453b] VM Resumed (Lifecycle Event) 2023-03-23 09:37:52.694 3209246 ERROR nova.virt.libvirt.driver [-] [instance: 17672112-c416-494a-88f8-fd7cfa85453b] Live Migration failure: internal error: qemu unexpectedly closed the monitor: 2023-03-23T09:37:52.215715Z qemu-system-x86_64: VQ | |
| 09:52:51 | bauzas | 0 size 0x80 < last_avail_idx 0x0 - used_idx 0x44 2023-03-23T09:37:52.215742Z qemu-system-x86_64: Failed to load virtio-balloon:virtio 2023-03-23T09:37:52.215745Z qemu-system-x86_64: error while loading state for instance 0x0 of device '0000:00:05.0/virtio-balloon' | |
| 09:53:06 | zigo | Yeah, I do. But then, where do I look? | |
| 09:53:09 | zigo | Libvirt logs? | |
| 09:54:07 | bauzas | it reminds me this bug report https://bugs.launchpad.net/cloud-archive/+bug/1848497 | |
| 09:55:19 | zigo | I haven't see anything doing a tail of /var/log/libvirt/qemu/*.log | |
| 09:55:58 | zigo | On Bullseye, I'm running with qemu 1:5.2+dfsg-11+deb11u2 | |
| 09:57:45 | bauzas | and you're not seeing anything with qemu logs ? | |
| 09:58:07 | bauzas | the error is reported by qemu process, not by libvirtd | |
| 09:58:18 | bauzas | so I'd say check the qemu logs | |
| 10:03:23 | bauzas | zigo: so it seems a qemu migration *to* a node with a version of 4.0 or higher is problematic | |
| 10:03:34 | bauzas | which qemu version the source is running ? | |
| 10:03:51 | bauzas | I assume you're not mixing releases | |
| 10:03:58 | bauzas | but I wanted to double-check | |
| 10:04:53 | zigo | Same version of qemu and libvirt in both source and dest. | |
| 10:05:27 | zigo | It's a plain Bullseye, so I use whatever is in Debian Stable (minus the security upgrades that I'm trying to perform). | |
| 10:06:01 | bauzas | and what migration flags are you using ? | |
| 10:06:25 | bauzas | are the vms paused ? | |
| 10:06:39 | bauzas | or suspended ? | |
| 10:08:45 | zigo | They are ACTIVE. | |
| 10:08:56 | zigo | Is there migration flags I can set?!? :) | |
| 10:08:59 | zigo | Where do I look? | |
| 10:09:41 | zigo | I've just done "nova host-evacuate-live <hostname>" ... | |
| 10:10:05 | zigo | (not sure if there's a way to do this with python-openstackclient from Victoria...) | |
| 10:11:11 | zigo | What's weird, is that MANY VMs on the same host are live-migrating without a glitch. Then on average, 2 VMs on each compute can't live-migrate ... | |
| 10:11:49 | sean-k-mooney | zigo: there isnet and that intentional | |
| 10:12:05 | sean-k-mooney | host-evacuate-live is not somthing we recomend operators use | |
| 10:12:14 | zigo | sean-k-mooney: What should I use then? | |
| 10:12:15 | sean-k-mooney | we are intentioally not supporting it in osc | |
| 10:12:33 | bauzas | wait | |
| 10:12:41 | bauzas | evacuate or live-migrate ? | |
| 10:12:45 | bauzas | I'm lost here | |
| 10:12:46 | sean-k-mooney | zigo: you should write yoru won code to live migrate all the vms forma host that actully has error handeling | |
| 10:12:53 | bauzas | oh | |
| 10:12:58 | bauzas | host-evacuate-live | |
| 10:13:08 | bauzas | damn old unspported CLIs | |
| 10:13:24 | sean-k-mooney | technially deprecated rather then unsuppported | |
| 10:13:35 | sean-k-mooney | until we can remvoe it in C/D | |
| 10:13:47 | sean-k-mooney | we need everyone ot use the sdk first | |
| 10:15:11 | sean-k-mooney | zigo: it wont help now but you used to be able to set https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.mem_stats_period_seconds to 0 to disable the memory ballon | |
| 10:15:48 | sean-k-mooney | can you confirm if that is set to 0 on either host and that the vm has a memory ballon | |
| 10:16:26 | zigo | It's set to default on that host (ie: 10 ...). | |
| 10:16:31 | sean-k-mooney | ack | |
| 10:16:48 | zigo | Should I set it to zero and try again then? | |
| 10:16:50 | sean-k-mooney | i was wondering if having it enabel and disable on differnt hsot coudl cause issues | |
| 10:17:26 | sean-k-mooney | i think this is one of those thigns that you cant change with runnign vms | |
| 10:17:40 | sean-k-mooney | zigo: i assume you cant just cold migrate them | |
| 10:17:43 | zigo | Oh ... :/ | |
| 10:17:43 | bauzas | yup, or it would require a vm recycle | |
| 10:17:59 | sean-k-mooney | /recycle/restart/ | |
| 10:18:24 | zigo | sean-k-mooney: Well, to cold-migrate, I must get in touch with customers to at least warn them about the operation, and let them know their VM will reboot. | |
| 10:18:35 | zigo | That's kind of very annoying with 50 computes and 2k+ VMs ... | |
| 10:18:39 | sean-k-mooney | to answer your orgianl question no im not aware of live migration issues related to memory baloons | |
| 10:20:53 | bauzas | sean-k-mooney: I remember we had some old qemu-4 live migration issues with the qemu balloons | |
| 10:21:02 | bauzas | but zigo isn't impacted | |
| 10:27:50 | sean-k-mooney | zigo: i assume this is persistent | |
| 10:28:03 | sean-k-mooney | i.e. a secodn live migration of the vm has the same error | |
| 10:28:23 | zigo | Right. | |
| 10:28:35 | sean-k-mooney | have you tried migratign to a differnt host? | |
| 10:29:11 | zigo | I didn't try to specify the dest host, but I can try. I'll let you know... | |
| 10:29:24 | sean-k-mooney | im trying to fiture out is the a thing tha that is speciic to the vm, the destion host ectra | |
| 10:32:03 | sean-k-mooney | https://bugzilla.redhat.com/show_bug.cgi?id=1923881 | |
| 10:32:06 | sean-k-mooney | that sound like it | |
| 10:34:13 | sean-k-mooney | similar for virtio-blk https://bugs.launchpad.net/nova/+bug/1737625 | |
| 10:34:55 | sean-k-mooney | ""Dave notes that we get this "guest index inconsistent" error when the migrated RAM is inconsistent with the migrated 'virtio' device state. And a common case is where a 'virtio' device does an operation after the vCPU is stopped and after RAM has been transmitted.""" | |
| 10:36:13 | sean-k-mooney | zigo: are you using post-copy or autoconverge by the way | |
| 10:37:19 | sean-k-mooney | bauzas: this is the qemu 4.0 issue right https://lore.kernel.org/all/156517411102.26464.1302440989654328620.launchpad@gac.canonical.com/T/ | |
| 10:37:42 | sean-k-mooney | https://bugs.launchpad.net/qemu/+bug/1838569 | |
| 10:38:36 | zigo | Not using post-copy (it's set to default, ie: false) | |
| 10:39:11 | zigo | Same for live_migration_permit_auto_converge (set to default: false) | |
| 10:39:22 | zigo | I'm using TLS though... | |
| 10:39:31 | zigo | libvirt over TLS. | |
| 10:40:34 | zigo | I probably should set live_migration_permit_auto_converge to true though, as sometimes, I have to manualy do a live-migration-force-complete ... | |
| 10:40:53 | bauzas | sean-k-mooney: unrelated, was it you who wrote https://etherpad.opendev.org/p/nova-bobcat-ptg#L75 ? | |