Earlier  
Posted Nick Remark
#openstack-nova - 2019-11-20
21:44:23 efried Yeah, so the question is whether we should keep the vTPM
21:44:39 dansmith except when it's evacuate
21:44:40 dansmith we should, I would think
21:44:44 dansmith rebuild is akin to putting the Windows 98 disc in the drive and rebooting
21:44:51 sean-k-mooney artom: so ya im happy with those one thing i notices is we shoudl update teh job templates to drop python 2 and 35 for the tempest plugin and just use the ussuri template instead
21:45:11 sean-k-mooney that can be in a follow up patch
21:45:15 efried I kinda want to say the vTPM should ride with the image
21:45:55 efried that would make for consistency with the backup/restore and shelve/unshelve type cases.
21:46:09 sean-k-mooney we do not nuke addtional ephemeral disks on rebuild right?
21:46:15 sean-k-mooney just the root disk
21:46:23 dansmith vtpm should survive a rebuild
21:46:25 dansmith for sure
21:46:31 sean-k-mooney you could look at the vtpm the same way
21:46:48 efried i.e. any time you create an instance from an image, we check the image for the metadata that points to the vtpm swift store.
21:47:07 efried dansmith: so what if your instance has a vTPM, and then you rebuild with an image that has ^ metadata?
21:47:09 mriedem efried: https://docs.openstack.org/nova/latest/contributor/evacuate-vs-rebuild.html ?!!??! :)
21:47:12 sean-k-mooney efried: i dont think that is right
21:47:26 dansmith efried: instance create is like creating a new server box, rebuild is like reinstalling the OS
21:47:30 sean-k-mooney efried: i think the vtpm lifetime shoudl be tied to the instance not the image
21:47:33 dansmith TPM is hardware, therefore it stays during a rebuild, IMHO
21:47:38 dansmith sean-k-mooney: ++
21:47:56 efried So what if your instance has a vTPM, and then you rebuild with an image that has pointer-to-vtpm-in-swift metadata?
21:48:09 efried "they better match, or we punt"?
21:48:20 dansmith efried: you mean if you rebuild from a snapshot?
21:48:29 dansmith remember the pointer to swift stays with the instance not the snapshot
21:48:43 dansmith oh you mean for the backup case.. that's why it doesn't work for the backup case :)
21:48:45 efried no, you convinced me earlier that snapshot and restore should work.
21:48:58 dansmith no, I didn't, I said "you're effed on one of these two things depending on what you decide :)
21:49:15 sean-k-mooney yes and the point would be stored in the nova system meatadat not the snapshot metadata
21:49:32 efried but that doesn't work for backup/restore
21:49:35 sean-k-mooney *pointer to the tpm snapshot
21:49:38 efried because there's no instance meta
21:49:46 efried there is only image meta.
21:50:10 sean-k-mooney oh right we were talking about shelve and unshelve before
21:50:32 efried if we go on the philosophy that TPM is like hardware, it's not unreasonable that backup/restore would lose it.
21:50:45 sean-k-mooney well for vpmem and ephermeral disk we dont snapshot them currenlty
21:51:01 efried if I do a backup of my laptop and then restore that backup on a different shell, my hardware changed.
21:51:36 sean-k-mooney and if you reinstall the os on the same laptop the tpm data is preserved
21:51:42 dansmith efried: but that's why VMs are better than hardware
21:51:45 efried heh
21:52:22 efried Okay, does this sound reasonable:
21:53:04 efried On rebuild, if your instance had a vTPM before, and the image specifies a vTPM also, they must be the same vTPM, or we fail.
21:53:37 sean-k-mooney not really
21:53:45 efried what should we do in that case?
21:54:02 sean-k-mooney im not sure the image shoudl be able to reference a vtpm
21:54:11 efried then we can't do backup/restore.
21:54:19 sean-k-mooney of the tpm
21:54:35 dansmith no
21:55:04 sean-k-mooney we can still do backup and resotre of the vm root disk like a normal snapshot
21:55:08 dansmith efried: on rebuild you don't have a tpm in swift because you have it on disk, it hasn't gone anywhere
21:55:25 dansmith efried: so if the snapshot specifies one, you wouldn't use it anyway
21:55:26 efried so if the image designates a vTPM, ignore it
21:55:43 dansmith the thing I'm more worried about if you do this,
21:55:54 dansmith is you snapshot an instance and then start 100 copies for it
21:55:57 dansmith *of it
21:56:01 dansmith you just sprayed your keys everywhere
21:56:06 efried sean-k-mooney: ftr, this is where dansmith convinced me the vTPM needs to ride with the backup http://eavesdrop.openstack.org/irclogs/%23openstack-nova/%23openstack-nova.2019-11-20.log.html#t2019-11-20T17:00:12
21:56:09 dansmith like a dog marking its territory on a hot summer day
21:56:24 efried nice visual
21:56:25 dansmith efried: again
21:56:47 dansmith efried: I'm not saying that it should necessarily, I'm saying it's going to eff up someone one way or the other
21:57:09 dansmith people use snapshot to roll back from a bad OS upgrade, and if you don't keep their keys, they're going to get screwed
21:57:12 sean-k-mooney should it be configurable?
21:57:19 dansmith but if you do, and they spawn that image in a bunch of places, ...
21:57:37 efried dansmith: well... we could also delete the swift obj (and mod the image meta) the first time we restore it.
21:57:50 efried that works poorly for crashy hosts tho
21:58:09 dansmith efried: that's fine if you store it in the image record, despite a failed host
21:58:26 dansmith efried: but I think that's probably going to be fragile and maybe confusing why the first instance got it and none of the others did
21:58:30 dansmith but perhaps that's the best option
21:58:35 dansmith just saying, it's going to be icky
21:59:13 efried I'm saying: if my workflow is to back up stuff in case my host goes down, and I do that backup, and then my host goes down, and I restore it, and then my host goes down again, then when I restore it the second time, my vTPM will be gone.
21:59:31 dansmith oh, yes, I see what you mean
21:59:35 sean-k-mooney is this something we want the user to express a policy on e.g. hw_vtpm_snapshot=true|false
21:59:53 dansmith efried: this is also bad for the shared storage case in general actually,
22:00:02 efried sean-k-mooney: we could configure it to death, really.
22:00:07 dansmith efried: one thing about ceph backed instances is you can recover from a failed host with an evac
22:00:40 dansmith efried: but in this case you will fail to boot after evac because we don't have the keys anywhere
22:00:40 sean-k-mooney dansmith: you would loose the tpm data in that case
22:00:41 dansmith so that's another case this breaks pretty badly
22:00:46 dansmith volume-backed or ceph/nfs backed ephemeral
22:00:51 dansmith sean-k-mooney: right, that's what I just said
22:01:46 sean-k-mooney nfs might be ok if you also put the tpm datastore on nfs but ya. the impact of that will depend on what you used the tpm for i guess
22:02:09 dansmith sean-k-mooney: oh right, but not ceph
22:02:31 efried because ceph can only handle the disk itself?
22:02:36 efried similar to volume-backed
22:02:44 dansmith we don't actually mount the storage on the host with ceph
22:02:54 dansmith qemu talks to ceph itself, so we don't like have a dir mounted or anything
22:02:55 sean-k-mooney the tpm data store is not a volume so we wont create it as such in ceph
22:03:03 efried wow, neat.
22:03:22 efried dansmith: you know, in that case it's juuust possible qemu will store the vTPM file on ceph too
22:03:26 sean-k-mooney we kind of which qemu could do the same for the other backedn too
22:03:37 dansmith efried: I don't think so
22:03:38 efried I doubt it though.
22:03:39 efried yeah.
22:03:45 dansmith efried: we'd have to give it a volume id for that or something
22:04:08 dansmith the flow chart of "how your instance behaves if you have a vtpm" is starting to get pretty scary
22:04:10 efried today it makes up the directory based on the VM ID
22:04:20 efried dansmith: well, it's always the corner cases that suck.
22:04:25 efried 80/20
22:04:37 dansmith efried: well, snapshot is not a corner case IMHO :)

Earlier   Later