Earlier  
Posted Nick Remark
#openstack-nova - 2020-08-26
17:02:59 openstackgerrit Merged openstack/nova stable/train: libvirt: Provide VIR_MIGRATE_PARAM_PERSIST_XML during live migration https://review.opendev.org/747973
17:03:15 openstackgerrit Merged openstack/nova master: releasenotes: Detail support for server ops with vTPM https://review.opendev.org/748215
17:04:26 noonedeadpunk sean-k-mooney: just one more stupid question... trying to find host-evacuate api call in https://docs.openstack.org/api-ref/compute/ but don't see for some reason (only server evacuate)
17:05:03 yoctozepto sean-k-mooney: https://docs.openstack.org/api-ref/compute/?expanded=evacuate-server-evacuate-action-detail#evacuate-server-evacuate-action ?
17:05:10 yoctozepto noonedeadpunk: ^
17:05:28 yoctozepto it's server (instance/vm) that's getting evacuated
17:06:17 noonedeadpunk and how does https://docs.openstack.org/nova/rocky/admin/evacuate.html#evacuate-all-instances work? It gets list of instances and evacuate one by one ?
17:09:07 yoctozepto noonedeadpunk: that's what I would assume, could use checking the code
17:09:47 sean-k-mooney noonedeadpunk: host-evacuate is not an api action
17:09:55 sean-k-mooney noonedeadpunk: its a client command
17:09:59 noonedeadpunk ok, got it
17:10:00 noonedeadpunk yeah
17:10:24 sean-k-mooney noonedeadpunk: http://www.danplanet.com/blog/2016/03/03/evacuate-in-nova-one-command-to-confuse-us-all/
17:11:14 sean-k-mooney dansmith's blog that is basically mandatory reading on host-evacuate and how it related to every thing else
17:11:47 noonedeadpunk oh, ok, now I know under what conditions our client loose all of their data from ephemeral drives...
17:11:48 sean-k-mooney yoctozepto: host-evacuate does not call teh evacuate api
17:11:55 noonedeadpunk when masakari calls instance evacuate
17:11:55 sean-k-mooney yoctozepto: it does cold migrate
17:12:33 sean-k-mooney actully no it does evacuate
17:12:39 sean-k-mooney i should read the blog more often
17:13:21 noonedeadpunk "The core of the evacuate process in nova is actually rebuild, which in many cases is a destructive operation"
17:13:31 noonedeadpunk so it's not evacuate and you;re right
17:13:40 sean-k-mooney from the blog
17:13:42 sean-k-mooney "The nova host-evacuate command does not translate directly to a server-side operation, but is more of a client-side macro or “meta operation.” When you call this command, you provide a hypervisor hostname, which the client uses to list and trigger evacuate operations on each instance running on that hypervisor. You would use this command post-failure (just like the single-instance evacuate
17:13:45 sean-k-mooney command) to trigger evacuations of all the instances on a failed compute host."
17:14:10 sean-k-mooney noonedeadpunk: no it is evacuate
17:14:26 sean-k-mooney noonedeadpunk: but evacuate does not mean what you think it does
17:14:57 noonedeadpunk yeah
17:15:20 noonedeadpunk so when masakari calls instance evacuate when this instance is not volume based - it get's wiped out
17:15:21 sean-k-mooney evacuate only preseved data if you use bfv or are on shared storage
17:15:51 sean-k-mooney noonedeadpunk: unless nova is using ceph via the rbd imageages type or /var/lib/nova is on nfs
17:15:53 noonedeadpunk and I guess even with rbd drive rebuild is destructove for non-bfv?
17:16:19 noonedeadpunk hm...
17:16:19 sean-k-mooney noonedeadpunk: no its not with ceph we detech its on shared storage
17:16:46 noonedeadpunk hm.....
17:17:23 sean-k-mooney dansmith: its proably in your blog post
17:17:29 noonedeadpunk then I still don't get why I got situations when ephemeral got lost or wiped out during node crush....
17:17:38 noonedeadpunk but whatever)
17:17:45 noonedeadpunk I know how to disable this :p
17:18:06 sean-k-mooney dansmith: i keep it bookmarked as a reference document
17:18:20 dansmith yeah, I should charge admission
17:19:16 sean-k-mooney noonedeadpunk: ootnote: In the case of volume-backed instances, the root disk of the instance is usually in a common location such as on a SAN device. In this case, the root disk is not destroyed, but any other instance state is recreated (which includes memory, ephemeral disk, swap disk, etc).
17:19:47 noonedeadpunk yeah, so it's re-created from the image?
17:19:52 noonedeadpunk as rebuild do
17:20:35 noonedeadpunk which is destructive and all changes made on vm are gone which is eventually as intended?
17:20:36 sean-k-mooney the root disk shoudl not be but if you have extra ephermeral disk they proably are i dont know how addtional ephemeral disks work with ceph. i now swap is store as a ceph volume
17:20:43 yoctozepto sean-k-mooney, noonedeadpunk: I guess masakari should include a warning in big, red font about the necessity to have HA storage first before running masakari
17:21:00 noonedeadpunk I guess so...
17:21:05 yoctozepto it's obvious to us (well, me at least) but years of practice has proven it's not entirely that obvious to all of users
17:21:10 sean-k-mooney yoctozepto: if masikari automate evacuate yes
17:21:20 sean-k-mooney or keep all your data on a cinder data volume
17:21:29 yoctozepto yeah, that's what I meant
17:21:29 noonedeadpunk I mean I have ceph for everything.... But still non bfv instances get's wiped out during evacuations...
17:21:59 noonedeadpunk ok, whatever)
17:22:03 sean-k-mooney noonedeadpunk: ill check the ceph code quickly but i tought we dedect ceph and preserved it.
17:22:23 noonedeadpunk maybe it's some recent change? Ie T or U?
17:23:07 sean-k-mooney we pass an on_shared_storage flag internllay which should be set to true for ceph and nfs
17:24:07 sean-k-mooney thats set here https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/api/openstack/compute/evacuate.py#L97
17:24:30 sean-k-mooney oh you have to pass it as an option
17:25:07 noonedeadpunk sean-k-mooney: yoctozepto seems like another patch to masakari?:)
17:25:15 sean-k-mooney noonedeadpunk: https://docs.openstack.org/nova/latest/reference/api-microversion-history.html#id12
17:25:32 sean-k-mooney so before 2.14 you had to pas it
17:25:38 sean-k-mooney after that its automatic
17:26:19 noonedeadpunk 2.12 was soooooo long ago....
17:27:32 sean-k-mooney is that what massikari is using
17:27:51 sean-k-mooney https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/compute/manager.py#L3504 we now ask the dirver if the insance is on shared storage
17:28:00 noonedeadpunk hm. I think I'm reading it wrong.. But it returns None for 2.14+ ? https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/api/openstack/compute/evacuate.py#L45
17:28:04 sean-k-mooney form 2.14 on
17:28:25 sean-k-mooney noonedeadpunk: yes 2.14 on it returns none and we ask the driver
17:28:31 noonedeadpunk ah
17:29:10 sean-k-mooney which does self.image_backend.backend().
17:29:13 sean-k-mooney is_shared_block_storage()
17:29:38 sean-k-mooney so unless you have differen image_backend on different hosts that should be correct
17:30:07 sean-k-mooney rbd returns true https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/virt/libvirt/imagebackend.py#L958
17:30:45 sean-k-mooney so you should see this log message when you evacuate https://github.com/openstack/nova/blob/20459e3e88cb8382d450c7fdb042e2016d5560c5/nova/compute/manager.py#L3514-L3516
17:30:58 sean-k-mooney and we should keep the disk content
17:31:30 noonedeadpunk looks like this, yes
17:31:57 sean-k-mooney what api microversion is masikari using
17:33:22 noonedeadpunk 2.53 for master...
17:33:49 sean-k-mooney it looks like it does not set one https://github.com/openstack/masakari/blob/cd95f6660c14f506603f0864ca24dd25c278a6a8/masakari/engine/drivers/taskflow/host_failure.py#L262-L263
17:34:08 sean-k-mooney so nova client will default to the latest
17:34:17 noonedeadpunk ok, on rocky it was 2.14
17:35:04 yoctozepto we care about stein+ atm
17:35:16 noonedeadpunk ++
17:35:36 sean-k-mooney 2.14 was mitaka
17:35:42 sean-k-mooney so you should be good
17:36:11 sean-k-mooney we have been preseviing shared storage vms for a very long time
18:00:01 noonedeadpunk sean-k-mooney: just tested and during evacuate data is really preserved
18:00:27 sean-k-mooney good that would have been a nasty bug if that was regressed
18:00:41 openstackgerrit Stephen Finucane proposed openstack/nova master: Change default num_retries for glance to 3 https://review.opendev.org/740389
18:02:50 noonedeadpunk TIL
18:02:56 openstackgerrit Merged openstack/nova stable/ussuri: compute: Validate a BDMs disk_bus when provided https://review.opendev.org/744550
18:03:31 sean-k-mooney noonedeadpunk: the api does not guarentee that data is preserved unless you are using bfv
18:03:57 sean-k-mooney noonedeadpunk: so unless you deployed the cloud and know the vm is on shared storage when you dont use bfv dont rely on that behavior
18:06:27 noonedeadpunk I see, yeah, I know that bfv is really preferable. But it's hard to explain in terms of public clouds, when even with no non-0 disk flavors ppl still do use ephemeral. But good part of them, is that they can handle ISO's while cinder wasn't able to boot ISO image properly last time I've checked...
18:06:40 openstackgerrit Merged openstack/nova master: rbd: Move rbd_utils out of libvirt driver under nova.storage https://review.opendev.org/746904
18:06:56 sean-k-mooney noonedeadpunk: well i have no problem iwht non bfv instance
18:07:07 sean-k-mooney i like the ceph backend for example
18:07:24 noonedeadpunk yes, sure ceph backend is everywhere for me:)
18:07:24 sean-k-mooney but its jsut driver devied if evauate preserves data or not

Earlier   Later