| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-27 | |||
| 09:34:37 | auniyal__ | (Pdb) | |
| 09:34:37 | auniyal__ | {'id': '2d867afe-0e2d-480b-82e9-e19fcd7b16d5', 'links': [{'rel': 'self', 'href': 'http://c6ac6857-c7be-439e-ba5f-43639cd09e84/v2.1/servers/2d867afe-0e2d-480b-82e9-e19fcd7b16d5'}, {'rel': 'bookmark', 'href': 'http://c6ac6857-c7be-439e-ba5f-43639cd09e84/servers/2d867afe-0e2d-480b-82e9-e19fcd7b16d5'}], 'OS-DCF:diskConfig': 'MANUAL', 'security_groups': [{'name': 'default'}], 'adminPass': 'SoXFm9gf3ksh'} | |
| 09:34:42 | gibi | in glance | |
| 09:34:46 | gibi | not gerrit /o\ | |
| 09:34:47 | bauzas | gibi: context is, auniyal__ is creating a BFV instance with some properties | |
| 09:35:02 | bauzas | gibi: and after this, snapshoting the instance | |
| 09:35:17 | bauzas | but when snapshoting, the quiesce property doesn't exist | |
| 09:35:29 | bauzas | my two guesses are : | |
| 09:35:45 | bauzas | 1/ the fixture can create a fake instance that needs to adjust its properties | |
| 09:36:12 | bauzas | 2/ the snapshot command doesn't really pull the image properties of the original image | |
| 09:37:22 | auniyal__ | in server object, there are 2 links, one "rel": self, and other "rel": bookmark ? | |
| 09:37:28 | gibi | I would verify 2/ in devstack first. maybe we have a bug on snapshot? | |
| 09:37:54 | bauzas | gibi: indeed we have | |
| 09:39:04 | auniyal__ | gibi, in here - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/ | |
| 09:39:07 | bauzas | auniyal__: remind me the snapshot bug report ? | |
| 09:39:20 | bauzas | gibi: and agreed on the need for a devstack testing | |
| 09:39:26 | auniyal__ | bug: https://bugs.launchpad.net/nova/+bug/1980720 | |
| 09:40:22 | auniyal__ | in pastebin, at line 83 - properties | |
| 09:42:07 | gibi | bauzas: I mean we have another more generic bug on snapshot where we simply not copy over image properties to snapshots | |
| 09:42:39 | gibi | maybe? | |
| 09:42:44 | auniyal__ | I belive, require_quiensce is getting set - because later my control, does come here - https://opendev.org/openstack/nova/src/commit/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/compute/api.py#L3471 | |
| 09:43:15 | auniyal__ | but then failed - saying - nova.exception.QemuGuestAgentNotEnabled: QEMU guest agent is not enabled | |
| 09:44:01 | bauzas | gibi: yeah that's my guess #2 | |
| 09:44:09 | bauzas | but that needs to be verified | |
| 09:44:40 | bauzas | eventually found the bug report... | |
| 09:44:43 | bauzas | gibi: https://bugs.launchpad.net/nova/+bug/1980720 | |
| 09:45:07 | gibi | auniyal__: can you track down from where the QemuGuestAgentNotEnabled is raised? | |
| 09:45:12 | bauzas | this is specific to quiesce, but I guess no properties are given | |
| 09:45:20 | bauzas | gibi: yeah we did | |
| 09:45:32 | bauzas | and this is when you quiesce on the snapshot | |
| 09:46:19 | gibi | bauzas: but that only complain about the error message not on the error it self. the report says there is no qemu agent running, so the error is expected | |
| 09:46:22 | bauzas | https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/virt/libvirt/driver.py#L3244 | |
| 09:46:26 | bauzas | yeah | |
| 09:46:38 | bauzas | anyway, I need to stop by now | |
| 09:46:46 | bauzas | let's discuss this this afternoon | |
| 09:47:00 | gibi | so that does not prove we have a bug in snapshot that failes to copy some image properties | |
| 09:47:03 | gibi | bauzas: ack | |
| 09:48:40 | gibi | bauzas: have a nice one | |
| 09:52:13 | opendevreview | Amit Uniyal proposed openstack/nova stable/yoga: Adds check if blk_dev_info has correct flavor.swap https://review.opendev.org/c/openstack/nova/+/859246 | |
| 09:53:06 | opendevreview | Amit Uniyal proposed openstack/nova stable/wallaby: Adds check if blk_dev_info has correct flavor.swap https://review.opendev.org/c/openstack/nova/+/859247 | |
| 10:28:57 | auniyal__ | gibi, QemuGuestAgentNotEnabled is coming from https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/virt/libvirt/driver.py#L3223 | |
| 10:30:35 | gibi | auniyal__: could you check image_meta.properties there? is it only hw_qemu_guest_agent missing or os_require_quiesce too? | |
| 10:32:47 | auniyal__ | yes, thats also missing | |
| 10:33:05 | auniyal__ | this is a object, - https://paste.opendev.org/show/bD0rlumubeiQ6Yeowaiw/ | |
| 10:33:34 | auniyal__ | in this - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/ | |
| 10:33:52 | auniyal__ | property is preset in image at line 57 | |
| 10:34:36 | auniyal__ | s/preset/present/ | |
| 10:38:34 | gibi | auniyal__: this point to the direction that we have an issue saving the image properties when we are doing the snapshot | |
| 10:43:57 | auniyal__ | gibi: in here - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/ | |
| 10:44:40 | auniyal__ | before line 88, we have server object, later from this object only we retrieve these image property | |
| 10:44:46 | auniyal__ | *must be | |
| 10:45:22 | auniyal__ | so is there a way, we can check here itself, if propeties gets attached or not, | |
| 10:45:27 | auniyal__ | before calling snapshot | |
| 10:47:51 | auniyal__ | here, https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/objects/image_meta.py#L114 | |
| 10:49:52 | gibi | auniyal__, bauzas: as far as I see we create some of the image metadata during snapshot here https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/virt/libvirt/driver.py#L2994 but there is no sign that we intended to copy image properties over to the snapshot. We did copy os_type interestingly but not the rest. The os_type was added there as a bug fix | |
| 10:49:58 | gibi | https://review.opendev.org/c/openstack/nova/+/42877 So now I'm wondering if we want to copy everything there or not | |
| 10:51:32 | gibi | auniyal__: you can try to copy the two image prop your fix need there https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/virt/libvirt/driver.py#L2994 to see if that helps | |
| 11:29:04 | auniyal__ | gibi, actually this get called https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/api/openstack/compute/servers.py#L1342 | |
| 11:31:00 | auniyal__ | at this place also - properties are not set | |
| 11:31:01 | auniyal__ | instance.image_meta.properties.get('hw_qemu_guest_agent') | |
| 13:03:37 | bauzas | gibi: good point, that was my guess #2 | |
| 13:33:40 | gibi | ahh this is volume backed snapshot. But then the instance is volume booted, so where qemu_guest_agent is coming from during the original boot from volume? | |
| 13:34:09 | opendevreview | Bence Romsics proposed openstack/nova stable/ussuri: Fix unplugging VIF when migrate/resize VM https://review.opendev.org/c/openstack/nova/+/859433 | |
| 13:34:24 | gibi | I'm confused that we talk about image properties but the instance is booted from volume so there might be no image at all involved | |
| 13:35:35 | opendevreview | Bence Romsics proposed openstack/nova stable/train: Fix unplugging VIF when migrate/resize VM https://review.opendev.org/c/openstack/nova/+/859434 | |
| 13:38:15 | rubasov | hi Nova folks: may I ask for reviews on these two backport series? both merged on master long ago, just taking the backports further: https://review.opendev.org/q/I2c195df5fcf844c0587933b5b5995bdca1a3ebed https://review.opendev.org/q/I3cb39a9ec2c260f422b3c48122b9db512cdd799b | |
| 14:27:58 | Uggla | gibi, bauzas, sorry coming back to this morning discussion. I'm afraid we can have a "deadlock" if we fail to mount a share, we may fail to umount it as well, and we could be stuck with a record. So maybe admin need something to remove this record in the DB ? | |
| 14:29:08 | gibi | Uggla: do you mean the unmount fails *because* we failed to mount the share in the first place? | |
| 14:29:37 | Uggla | I think we could have such kind of situation. | |
| 14:30:00 | Uggla | yes | |
| 14:31:25 | Uggla | and today if a share_mapping is already in the DB, there is an error --> no attempt to mount it again. | |
| 14:31:59 | Uggla | I could change that, to allow to mount it again if it is in error. | |
| 14:35:03 | Uggla | the idea is to discuss how to go out of this if it happens without deleted the db record in the db itself. | |
| 14:35:18 | Uggla | s/deleted/deleting | |
| 14:39:00 | Uggla | I would like also to discuss the behavior starting a vm if some share_mapping are in error. Should I prevent the vm from starting up or start it and do not mount the error shares and warn about it ? | |
| 14:39:46 | gibi | these are good questions | |
| 14:40:44 | Uggla | gibi, fyi also I have changed the behavior from what you saw in the previous code. | |
| 14:40:48 | gibi | so today a ShareMapping in error does not mean that the instance is in ERROR to? | |
| 14:40:51 | gibi | too | |
| 14:41:57 | Uggla | today 1- VM should be stopped, 2- Attach a share --> error --> share_mapping in status error in DB. | |
| 14:43:26 | Uggla | So here if the user starts the VM, I'm in favor of starting it and do not mount the faulty shares. | |
| 14:43:31 | gibi | ack | |
| 14:44:14 | Uggla | User can still see the reason why the shares are not mounted because the share_mapping has a error status. | |
| 14:44:25 | Uggla | *have | |
| 14:44:33 | Uggla | *an | |
| 14:45:41 | gibi | so at 2- nova tries to mount the share to the host the VM is on? | |
| 14:46:11 | gibi | and that mount can fail | |
| 14:46:51 | gibi | and 2- is sync so it does two thing i) puts the ShareMapping in ERROR ii) sends back http 500 to the user | |
| 14:48:13 | Uggla | yep except vm is off. | |
| 14:49:01 | gibi | and I guess if the VM is not on any host (never scheduled or it is shelved_offloaded) then we don't allow to attach a share | |
| 14:49:45 | Uggla | yep with this version the vm should exist and be powered off (shelve is not allowed.) | |
| 14:49:54 | gibi | cool | |
| 14:51:46 | gibi | so the only way to re-try the mount of the failes ShareMapping is to detach and then attach the share again. But you are worried what if detach fails at unmount. I think that is fine. If the system is in that bad shape then the admin needs to intervene | |
| 14:52:54 | gibi | i'm not sure what is the exact low lever scenario when the mount fails and leaves the host in a state where unmount is not possible, but I think we don't have to automate a fix for that right now | |
| 14:53:25 | gibi | if it turns out there a common case where mount then unmount fails, then probably there will be a common solution for that, and then we can add that common solution to our unmount codepath to try | |
| 14:54:02 | Uggla | exactly, but I'm afraid that the only solution for admin will be to remove the share in the DB itself. Not really convenient. | |
| 14:54:34 | gibi | no, the admin can fix the host in a way that the next detach share call that tries the unmount will succeed | |
| 14:54:44 | gibi | s/can/should be able to/ | |
| 14:55:17 | gibi | so we need to allow detach share to be called on an ShareMapping in error | |
| 14:55:23 | gibi | to re-try the unmount | |