Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-26
16:53:32 opendevreview Balazs Gibizer proposed openstack/nova stable/xena: Gracefully ERROR in _init_instance if vnic_type changed https://review.opendev.org/c/openstack/nova/+/859315
16:54:21 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Reproduce bug 1981813 in func env https://review.opendev.org/c/openstack/nova/+/859320
16:54:22 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Gracefully ERROR in _init_instance if vnic_type changed https://review.opendev.org/c/openstack/nova/+/859321
#openstack-nova - 2022-09-27
00:35:02 opendevreview melanie witt proposed openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358
09:09:49 Uggla gibi, bauzas, hello. I have a question about share_mapping deletion behavior. Assuming we cannot umount the share due to error. So we could be stuck in the state share cannot be deleted because it cannot be unmounted. What do you prefer a "force" option in the API or deleting it despite the error and warn the user that the umount was not properly done ?
09:10:17 bauzas damn
09:11:01 bauzas I'd prefer to return an error and still having the share status to be ACTIVE
09:11:50 gibi Uggla: yeah what bauzas said. The DB record is cheap to keep so I would not optimize on removing that.
09:12:53 bauzas Uggla: tbc, if we can't umount the share, then the user would need to ask again
09:14:43 Uggla ok in that case it means the op need to fix the umount issue to remove the share. Do we agree on that ?
09:15:11 bauzas honestly, I think so
09:15:32 bauzas the user can't know why nova doesn't work
09:15:41 bauzas so he could ask the operator to look at the problem
09:16:19 bauzas forcing to just delete the DB value while we would still have the mount could be a problem for a host after some time
09:17:00 bauzas and if the operator doesn't know that all the shares are still mounted, then he wouldn't see it until it could be a larger problem
09:17:57 Uggla ok fyi this is the msg given to user: {"badRequest": {"code": 400, "message": "Share id e8debdc0-447a-4376-a10a-4cd9122d7986 mount error from server 36a6c053-78c6-4409-9a44-b1e81244e61e.\\nReason: Unexpected error while running command.\\nCommand: mount\\nExit code: 1\\nStdout: \'This is stdout\'\\nStderr: \'This is stderror\'."}}
09:18:23 Uggla oops copy/paste the wrong one
09:18:37 Uggla s/mount/umount/
09:19:14 gibi if the command line or the error does not leak infra information to the user then I'm OK with this response
09:19:17 tobias-urdin should probably censor the command error and return a more generic error msg, had the same for ceph where it leaked the username in the error message returned to the user
09:19:29 tobias-urdin gibi: was faster :p
09:19:32 bauzas yeah agreed with tobias-urdin
09:19:37 gibi just tiny bit :)\
09:19:44 bauzas I'd prefer to have a better error
09:19:51 bauzas and, shouldn't be 400
09:20:07 bauzas that's not a *bad request* right?
09:20:40 gibi bauzas: good point this cannot be fixed by fixing the request
09:20:45 gibi so it is more like 500
09:21:56 bauzas Uggla: remind me something
09:22:06 bauzas Uggla: do we have share statuses ?
09:22:23 Uggla ok so do we agree that log should contains the reason (mount, stderr, stdout ...)
09:22:33 bauzas for the log, yes
09:22:39 Uggla yep we have a share_mapping status status
09:22:42 bauzas not for the error message to the user :)
09:23:20 bauzas like, say a user is asking to delete a share for an instance
09:23:39 bauzas he/she calls nova API for "deleting the share mapping"
09:23:53 bauzas he eventually gets "sorry, nay"
09:24:02 bauzas then the share mapping still exists
09:24:16 bauzas he then asks again "please delete my share mapping"
09:24:24 bauzas nova continues to tell "sorry, nay again"
09:24:56 bauzas then, the user would ask the operator to tell him/her "meh, can't delete my share mapping"
09:25:14 gibi ^^ +1
09:25:15 bauzas then the operator looks at the share mapping by the logs and see
09:25:29 bauzas 'oh man, this mapping UUID got an exception"
09:25:57 bauzas and then he/she says "oh, f*** that's why, this crazy umount didn't work"
09:26:18 bauzas that's how I see it
09:26:22 Uggla bauzas, sounds good to me and easier to manage. :)
09:27:34 bauzas the better would be to see the share mapping requests in the server actions list
09:27:38 bauzas Uggla: https://docs.openstack.org/api-ref/compute/#list-actions-for-server
09:27:50 bauzas but I don't think this is possible
09:28:16 bauzas ideally, if the operator could get the request ID of the "delete share mapping" call, that would be loving
09:28:45 bauzas and again, that's a question about whether we want to have a share mapping deletion to be synchronous or async
09:29:10 Uggla sync
09:29:29 bauzas then 500
09:29:43 Uggla ok
09:33:23 bauzas gibi: remind me, do we have a way to see the image properties an instance got ?
09:33:34 bauzas https://docs.openstack.org/api-ref/compute/?expanded=show-server-details-detail#show-server-details can't see them in the API ref
09:33:54 bauzas it will just give you the image UUID
09:34:11 gibi we copy them over to sysmeta but that is not visible on the API I guess
09:34:21 bauzas auniyal__: ^
09:34:35 auniyal__ I have this a server
09:34:36 auniyal__ (Pdb) server
09:34:37 gibi so the end user needs to use the image UUID and look up theimage in gerrit
09:34:37 auniyal__ (Pdb)
09:34:37 auniyal__ {'id': '2d867afe-0e2d-480b-82e9-e19fcd7b16d5', 'links': [{'rel': 'self', 'href': 'http://c6ac6857-c7be-439e-ba5f-43639cd09e84/v2.1/servers/2d867afe-0e2d-480b-82e9-e19fcd7b16d5'}, {'rel': 'bookmark', 'href': 'http://c6ac6857-c7be-439e-ba5f-43639cd09e84/servers/2d867afe-0e2d-480b-82e9-e19fcd7b16d5'}], 'OS-DCF:diskConfig': 'MANUAL', 'security_groups': [{'name': 'default'}], 'adminPass': 'SoXFm9gf3ksh'}
09:34:42 gibi in glance
09:34:46 gibi not gerrit /o\
09:34:47 bauzas gibi: context is, auniyal__ is creating a BFV instance with some properties
09:35:02 bauzas gibi: and after this, snapshoting the instance
09:35:17 bauzas but when snapshoting, the quiesce property doesn't exist
09:35:29 bauzas my two guesses are :
09:35:45 bauzas 1/ the fixture can create a fake instance that needs to adjust its properties
09:36:12 bauzas 2/ the snapshot command doesn't really pull the image properties of the original image
09:37:22 auniyal__ in server object, there are 2 links, one "rel": self, and other "rel": bookmark ?
09:37:28 gibi I would verify 2/ in devstack first. maybe we have a bug on snapshot?
09:37:54 bauzas gibi: indeed we have
09:39:04 auniyal__ gibi, in here - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/
09:39:07 bauzas auniyal__: remind me the snapshot bug report ?
09:39:20 bauzas gibi: and agreed on the need for a devstack testing
09:39:26 auniyal__ bug: https://bugs.launchpad.net/nova/+bug/1980720
09:40:22 auniyal__ in pastebin, at line 83 - properties
09:42:07 gibi bauzas: I mean we have another more generic bug on snapshot where we simply not copy over image properties to snapshots
09:42:39 gibi maybe?
09:42:44 auniyal__ I belive, require_quiensce is getting set - because later my control, does come here - https://opendev.org/openstack/nova/src/commit/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/compute/api.py#L3471
09:43:15 auniyal__ but then failed - saying - nova.exception.QemuGuestAgentNotEnabled: QEMU guest agent is not enabled
09:44:01 bauzas gibi: yeah that's my guess #2
09:44:09 bauzas but that needs to be verified
09:44:40 bauzas eventually found the bug report...
09:44:43 bauzas gibi: https://bugs.launchpad.net/nova/+bug/1980720
09:45:07 gibi auniyal__: can you track down from where the QemuGuestAgentNotEnabled is raised?
09:45:12 bauzas this is specific to quiesce, but I guess no properties are given
09:45:20 bauzas gibi: yeah we did
09:45:32 bauzas and this is when you quiesce on the snapshot
09:46:19 gibi bauzas: but that only complain about the error message not on the error it self. the report says there is no qemu agent running, so the error is expected
09:46:22 bauzas https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/virt/libvirt/driver.py#L3244
09:46:26 bauzas yeah
09:46:38 bauzas anyway, I need to stop by now
09:46:46 bauzas let's discuss this this afternoon
09:47:00 gibi so that does not prove we have a bug in snapshot that failes to copy some image properties
09:47:03 gibi bauzas: ack

Earlier   Later