| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-02-06 | |||
| 22:58:17 | mriedem | http://logs.openstack.org/62/634962/5/check/legacy-grenade-dsvm-neutron-multinode-live-migration/982e90b/ | |
| 23:01:40 | hogepodge | mriedem: sean-k-mooney: thanks | |
| 23:03:26 | sean-k-mooney | mriedem: so is that related to the block migration issue where it was complaing it was not on shared shared storage | |
| 23:03:39 | sean-k-mooney | ah it is | |
| 23:03:48 | sean-k-mooney | i was looking for https://bugs.launchpad.net/nova/+bug/1813216 yesterday | |
| 23:03:49 | openstack | Launchpad bug 1813216 in OpenStack Compute (nova) "legacy-grenade-dsvm-neutron-multinode-live-migration failing with "is not on shared storage: Shared storage live-migration requires either shared storage or boot-from-volume with no local disks." since Jan 21" [Medium,In progress] - Assigned to Matt Riedemann (mriedem) | |
| 23:05:21 | mriedem | yeah nova-compute during the ceph run didn't have the rbd config in it | |
| 23:05:59 | sean-k-mooney | mriedem: its not really a block migration if you are using the RBD iamge backend | |
| 23:07:27 | sean-k-mooney | its not a bfv instance either but when using RBD it really is a shared filesystem. | |
| 23:08:09 | sean-k-mooney | do we run the block migration tests on any multinode job without ceph/nfs | |
| 23:08:33 | mriedem | yes | |
| 23:08:37 | mriedem | well, i think so, | |
| 23:08:42 | mriedem | neutron multinode or whatever | |
| 23:09:30 | mriedem | http://logs.openstack.org/44/603844/19/check/tempest-multinode-full/698a95c/job-output.txt.gz#_2019-02-05_22_28_06_654011 | |
| 23:10:22 | sean-k-mooney | ah this one https://github.com/openstack/nova/blob/master/.zuul.yaml#L252-L254 | |
| 23:10:32 | mriedem | we don't have anything that tests the crazy environment lyarwood keeps trying to fix which is rbd shared storage for the root disk but local storage for the instance files | |
| 23:10:36 | sean-k-mooney | ok i was wondering if that came in form one of the templates | |
| 23:10:48 | mriedem | nova, for the most part, assumes that if you're using rbd you're shared everywhere | |
| 23:10:55 | mriedem | which apparently isn't true for some people | |
| 23:11:27 | sean-k-mooney | i dont understand the "local storage for the instance files" bit | |
| 23:11:39 | sean-k-mooney | i can ask lyarwood | |
| 23:11:41 | mriedem | the stuff under $instances_path | |
| 23:11:51 | mriedem | console logs, the image, config drive (i think) | |
| 23:11:52 | mriedem | etc | |
| 23:11:59 | sean-k-mooney | oh the libivt logs and stuff | |
| 23:12:28 | mriedem | yes this crazy thing https://docs.openstack.org/nova/latest/configuration/config.html#workarounds.ensure_libvirt_rbd_instance_dir_cleanup | |
| 23:12:57 | sean-k-mooney | i assumed if you set then image backend to ceph those where on ceph too well the config drive and swap at least | |
| 23:13:19 | mriedem | that's what the libvirt driver assumes as well in some places | |
| 23:13:55 | mriedem | but the rbd image backend just controls the disks, not the other files associated with the instance on the host | |
| 23:14:05 | sean-k-mooney | that explain some of my issue on my old newton cloud... | |
| 23:14:05 | mriedem | it's very fun | |
| 23:14:26 | sean-k-mooney | well the more you know | |
| 23:14:30 | mriedem | could probably add that to the pile-o-problems for cdent to work on shared storage providers... | |
| 23:15:12 | sean-k-mooney | i know the file that is created when you suspend an instace was written to the compute node but ya | |
| 23:15:27 | sean-k-mooney | i wonder how many people over look that as i did | |
| 23:16:31 | sean-k-mooney | oh i was talking to lyarwood about a multi attach live migration bug earlier in the week | |
| 23:16:37 | mriedem | well, we don't really help anyone with no mention of this in the docs really, or config options for that matter | |
| 23:17:18 | mriedem | we should probably have a "things to consider if you plan to use shared storage" section in the admin docs | |
| 23:17:28 | sean-k-mooney | assuming we dont already have a tempest senario test i think im going to write one and add it to the tempest-full-multinode job | |
| 23:17:57 | mriedem | there are no tempest tests for live migration + volume multiattach that i'm aware of | |
| 23:18:05 | mriedem | there are for resize though | |
| 23:19:17 | sean-k-mooney | mriedem: lyarwood was looking at a posibly race. i dont have the context to had but i was thinking a simple test that boots 2 vms with the same multi attach volume and migrate one of the vms back and forth | |
| 23:21:22 | sean-k-mooney | i think it was related to concurent update of the attachnets | |
| 23:21:40 | sean-k-mooney | mriedem: oh lyarwood filled an upstream bug https://bugs.launchpad.net/nova/+bug/1814245 | |
| 23:21:41 | openstack | Launchpad bug 1814245 in OpenStack Compute (nova) "_disconnect_volume incorrectly called for multiattach volumes during post_live_migration" [Undecided,In progress] - Assigned to Lee Yarwood (lyarwood) | |
| 23:24:26 | mriedem | so it disconnects the volume from the other server that wasn't live migrated | |
| 23:24:32 | mriedem | if both servers start on the same source host | |
| 23:24:33 | mriedem | yeah? | |
| 23:24:42 | sean-k-mooney | ya | |
| 23:25:27 | mriedem | i guess that bug says https://review.openstack.org/#/c/551302/ fixes it, | |
| 23:25:31 | mriedem | but ^ needs to be updated | |
| 23:27:23 | mriedem | sean-k-mooney: i don't know that a tempest test would catch this unless that test also ssh's into the guest and tries to verify the volume is still connected on the source host | |
| 23:27:40 | mriedem | and and connected to the live migrated server on the dest host | |
| 23:27:59 | sean-k-mooney | mriedem: yes i planned to ssh in touch a file migrate and check i could read the file on both vms | |
| 23:28:02 | mriedem | this is similar to https://review.openstack.org/#/c/548356/ | |
| 23:28:17 | mriedem | maybe you want to work on cleaning up ^ first | |
| 23:28:22 | mriedem | i'm sure steve is long gone by now | |
| 23:28:56 | mriedem | he was only on loan from oracle until they got multiattach support in queens | |
| 23:29:13 | sean-k-mooney | am i can take a look. i have not worked with tempest in a while so i kind of wanted to use this to get familar with it again | |
| 23:30:55 | sean-k-mooney | if that is isshing into the vm it really should be a senario test not an api test... but that is a different issue | |
| 23:31:56 | sean-k-mooney | the api tests "should" be able to run with the nova fake or libvirt fake driver | |
| 23:42:42 | openstackgerrit | Merged openstack/nova master: cleanup *.pyc files in docs tox envs https://review.openstack.org/635210 | |
| 23:51:01 | mshaikh99_ | I am using OpenStack Queens Release and running Ubuntu 16.04 VM on it, The disk IO seems to be very slow with default cache setting as none for Qemu based virt type | |
| 23:51:18 | mshaikh99_ | # time dd if=/dev/zero of=/tmp/test1 bs=20M count=100 conv=fsync;rm /tmp/test1 100+0 records in 100+0 records out 2097152000 bytes (2.1 GB) copied, 296.352 s, 7.1 MB/s real 4m56.559s user 0m0.100s sys 4m19.716s | |
| 23:51:35 | mshaikh99_ | Where is the compute node has the IO as: root@compute1:~# time dd if=/dev/zero of=/tmp/test1 bs=20M count=100 conv=fsync;rm /tmp/test1 100+0 records in 100+0 records out 2097152000 bytes (2.1 GB, 2.0 GiB) copied, 20.724 s, 101 MB/s | |
| 23:52:29 | mshaikh99_ | Based on the openstack forum link, I tried to chagne the default cache from none to writeback as suggested https://ask.openstack.org/en/question/6800/qemu-poor-disk-performance/ but it does not change and after the VM hardreboot I still see it as none | |
| 23:53:02 | mshaikh99_ | http://paste.openstack.org/show/744655/ | |
| 23:53:03 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Change live_migration_wait_for_vif_plug=True by default https://review.openstack.org/635360 | |
| 23:54:48 | mshaikh99_ | which file and which node? | |
| 23:56:04 | sean-k-mooney | mshaikh99_: did you do the hard reboot after restarting nova compute | |
| 23:57:33 | mshaikh99 | Hi | |
| 23:57:44 | mshaikh99 | Is there an update on the question I asked | |
| 23:57:46 | sean-k-mooney | mshaikh99: if the qcow file is not preaccloated that could be part of the problem | |
| 23:58:28 | sean-k-mooney | mshaikh99: https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.preallocate_images | |
| 23:59:29 | mshaikh99 | ok | |
| #openstack-nova - 2019-02-07 | |||
| 00:00:23 | sean-k-mooney | by the way did you do the hard reboot after you restart the nova compute service | |
| 00:01:17 | mshaikh99 | I did hard reboot of VM | |
| 00:01:25 | mshaikh99 | after nova compute service is restarted | |
| 00:02:26 | mshaikh99 | Although after hardreboot of VM, I do see time dd if=/dev/zero of=/tmp/test1 bs=20M count=100 conv=fsync;rm /tmp/test1 100+0 records in 100+0 records out 2097152000 bytes (2.1 GB, 2.0 GiB) copied, 42.8637 s, 48.9 MB/s | |
| 00:03:25 | mshaikh99 | so I guess the option for disk_cachemodes=file=writebac is not used any more. I can remove it from /etc/nova/nova-compute.conf [DEFAULT] compute_driver=libvirt.LibvirtDriver disk_cachemodes=file=writebac | |
| 00:03:59 | sean-k-mooney | well it still exists | |
| 00:04:01 | sean-k-mooney | https://docs.openstack.org/nova/latest/configuration/config.html#libvirt.disk_cachemodes | |
| 00:04:06 | sean-k-mooney | what release are you using again | |
| 00:05:00 | mshaikh99 | Queens | |
| 00:05:03 | sean-k-mooney | file=directsync would be safter then writeback by the way and may give better performance then writethrough | |
| 00:05:08 | mshaikh99 | So I only need to hard reboot and that would have increased the IO | |
| 00:06:08 | mshaikh99 | or I should keep the disk_cachemodes=file=directsync option in DEFAULT section of nova compute config file? | |
| 00:06:51 | melwitt | we had an old bug around that, but it was fixed in queens, so you should have the fix as long as you're running >= 17.0.0.0b2 https://bugs.launchpad.net/nova/+bug/1727558 | |
| 00:06:52 | openstack | Launchpad bug 1727558 in OpenStack Compute (nova) pike "libvirt driver ignores 'disk_cachemodes' configuration setting" [High,Fix committed] - Assigned to melanie witt (melwitt) | |
| 00:07:24 | mshaikh99 | I don't see that writebac option was picked up by the nova | |
| 00:07:26 | mshaikh99 | libvirt+ 27386 1 16 15:01 ? 00:03:06 /usr/bin/qemu-system-x86_64 -name guest=instance-00000002,debug-threads=on -S -object secret,id=masterKey0,format=raw,file=/var/lib/libvirt/qemu/domain-22-instance-00000002/master-key.aes -machine pc-i440fx-bionic,accel=tcg,usb=off,dump-guest-core=off -cpu EPYC,acpi=on,ss=on,hypervisor=on,erms=on,mpx=on,pcommit=on,clwb=on,pku=on,ospke=on,3dnowext=on,3dnow=on,vme=off,fma=off,avx | |
| 00:07:28 | sean-k-mooney | mshaikh99: disk_cachemode is part of the [libvirt] sechont not [DEFAULT] | |
| 00:08:16 | sean-k-mooney | mshaikh99: i would set [libvirt] disk_cachemodes=file=directsync | |
| 00:08:21 | mshaikh99 | ok let me try it then | |
| 00:09:40 | mshaikh99 | I have changed it and hard reboot my VM | |
| 00:09:58 | sean-k-mooney | and restarted nova-compute agent | |
| 00:11:49 | mshaikh99 | there is no nova-compute agent runnnig | |
| 00:12:20 | sean-k-mooney | the nova compute agent and nova compute service are the same thing | |
| 00:14:16 | mshaikh99 | ok now I see time dd if=/dev/zero of=/tmp/test1 bs=20M count=100 conv=fsync;rm /tmp/test1 100+0 records in 100+0 records out 2097152000 bytes (2.1 GB, 2.0 GiB) copied, 44.5152 s, 47.1 MB/s real 0m44.655s user 0m0.072s sys 0m40.564s | |
| 00:14:30 | mshaikh99 | about 47.1 MB | |