Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-23
22:59:26 clarkb https://zuul.opendev.org/t/zuul/build/74fab5fdddee4c3d88e71e40ad6795a7/log/docker/nodepool_nodepool-launcher_1.txt#1345 is the nodepool side of things reporting the error
22:59:51 clarkb Since the code hasn't chagned as far as I can tell I'm thinking the policy may have? But I'm not seeing where the policy might be set
23:04:31 clarkb hrm 909b0b02470dc795fd3d2775ee33864b055dd678 changed the default check_str from project admin to admin
23:05:07 clarkb But I think we've been able to do this successfully more recently than when that change landed
23:10:06 clarkb Sorry here is where we log that error message https://zuul.opendev.org/t/zuul/build/74fab5fdddee4c3d88e71e40ad6795a7/log/syslog#40077 which then maps back to that nodepool error (which is far more terse as that is what the sdk gives us)
23:22:14 clarkb It seem that both the recent failures and the most recent successful devstack installs for these jobs installed the same version of nova: aad31e6ba4 Merge "Update nova-manage doc page"
23:31:06 clarkb gmann: you've been pushing the rbac stuff along do you know if anything changed for that in the last day ?
23:34:39 gmann clarkb: policy is changed to admin from project admin means any admin can access it, so it is made more broader access than restrictive. also we have the old policy supported so it should work as it is
23:35:16 gmann may be we need to check if any change in token accessing it?
23:36:08 clarkb gmann: I don't think anything changed in how we access it unless openstacksdk just made a release /me checks
23:36:24 clarkb no the sdk updated a month ago
23:36:59 gmann clarkb: here, non admin is trying to access it https://zuul.opendev.org/t/zuul/build/74fab5fdddee4c3d88e71e40ad6795a7/log/syslog#40066
23:37:19 gmann 'is_admin': False,
23:37:36 clarkb gmann: yes, but that was working yesterday
23:38:05 clarkb is it possible to be a project admin but not an admin?
23:39:35 gmann no, project admin is nothing but admin with project_id matching
23:40:18 gmann admin is just role 'admin' match so any project _id so every project admin is admin
23:41:06 gmann I think 'is_admin':False is changed somewhere in sdk or so
23:41:07 clarkb got it
23:41:23 gmann it should be true
23:41:25 clarkb I'm now trying to compare the successful run to the failed one more broadly
23:41:35 clarkb to see if there are differences
23:47:55 clarkb Both the successful and failed jobs have this problem in the nova logs. But the failed one also has nova.exception.VirtualInterfaceCreateException: Virtual Interface creation failed and eventlet timeouts
23:49:07 clarkb is it posible that rbac error is just noise? And the real issue is that nova does go ahead and try to create virtual interfaces anyway but fails?
23:52:51 clarkb gmann: ok, I think https://zuul.opendev.org/t/zuul/build/74fab5fdddee4c3d88e71e40ad6795a7/log/syslog?severity=0#84665-84692 might be the actual fatal bit (I don't understand why we get the rbac errors but I'm thinking they may juts be noise now)
23:56:13 gmann humm, not sure why rbac error that is confusing then
23:57:21 clarkb gmann: I guess it is also possible that openstacksdk is probing the nova api to determine what actions it can take?
23:57:32 clarkb but I agree it is confusing
23:58:03 gmann yeah, may be
#openstack-nova - 2022-09-24
00:02:40 clarkb gmann: most of openstack devstack testing is still on focal right now?
00:02:51 clarkb (I notice we're running on jammy so that may be one difference)
00:08:40 gmann clarkb: yes, its on Focal currently and planned to migrate to jammy in next cycle
00:12:28 clarkb ok thanks. I've got a modified job setup in ci now to grab libvirt logs to see if we can diagnoes this better with libvirt info
#openstack-nova - 2022-09-25
08:44:38 opendevreview Merged openstack/nova stable/xena: neutron: Unbind remaining ports after PortNotFound https://review.opendev.org/c/openstack/nova/+/842586
09:31:05 opendevreview Merged openstack/nova stable/yoga: nova-live-migration tests not needed for Ironic https://review.opendev.org/c/openstack/nova/+/854257
19:00:15 opendevreview J.P.Klippel proposed openstack/nova master: fix typo in architecture document https://review.opendev.org/c/openstack/nova/+/859201
#openstack-nova - 2022-09-26
04:03:02 opendevreview junbo proposed openstack/nova master: Limit bandwidth during postcopy migration. https://review.opendev.org/c/openstack/nova/+/859207
12:41:48 opendevreview Justas Poderys proposed openstack/nova-specs master: Add support for Napatech LinkVirt SmartNICs https://review.opendev.org/c/openstack/nova-specs/+/859290
12:57:07 opendevreview Justas Poderys proposed openstack/nova-specs master: Add support for Napatech LinkVirt SmartNICs https://review.opendev.org/c/openstack/nova-specs/+/859290
13:00:07 justas_napa We have submitted a spec to add support for Napatech SmartNICs in Nova. If you have any questions - please ping me here or via dm.
13:12:57 opendevreview Justas Poderys proposed openstack/nova-specs master: Add support for Napatech LinkVirt SmartNICs https://review.opendev.org/c/openstack/nova-specs/+/859290
16:09:44 gibi auniyal: responded to sean-k-mooney's comment in https://review.opendev.org/c/openstack/nova/+/791135 . Let me know if the direction is unclear to you
16:52:25 opendevreview Balazs Gibizer proposed openstack/nova stable/yoga: Reproduce bug 1981813 in func env https://review.opendev.org/c/openstack/nova/+/859312
16:52:26 opendevreview Balazs Gibizer proposed openstack/nova stable/yoga: Gracefully ERROR in _init_instance if vnic_type changed https://review.opendev.org/c/openstack/nova/+/859313
16:53:31 opendevreview Balazs Gibizer proposed openstack/nova stable/xena: Reproduce bug 1981813 in func env https://review.opendev.org/c/openstack/nova/+/859314
16:53:32 opendevreview Balazs Gibizer proposed openstack/nova stable/xena: Gracefully ERROR in _init_instance if vnic_type changed https://review.opendev.org/c/openstack/nova/+/859315
16:54:21 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Reproduce bug 1981813 in func env https://review.opendev.org/c/openstack/nova/+/859320
16:54:22 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Gracefully ERROR in _init_instance if vnic_type changed https://review.opendev.org/c/openstack/nova/+/859321
#openstack-nova - 2022-09-27
00:35:02 opendevreview melanie witt proposed openstack/nova master: Unit test exceptions raised duing live migration monitoring https://review.opendev.org/c/openstack/nova/+/859358
09:09:49 Uggla gibi, bauzas, hello. I have a question about share_mapping deletion behavior. Assuming we cannot umount the share due to error. So we could be stuck in the state share cannot be deleted because it cannot be unmounted. What do you prefer a "force" option in the API or deleting it despite the error and warn the user that the umount was not properly done ?
09:10:17 bauzas damn
09:11:01 bauzas I'd prefer to return an error and still having the share status to be ACTIVE
09:11:50 gibi Uggla: yeah what bauzas said. The DB record is cheap to keep so I would not optimize on removing that.
09:12:53 bauzas Uggla: tbc, if we can't umount the share, then the user would need to ask again
09:14:43 Uggla ok in that case it means the op need to fix the umount issue to remove the share. Do we agree on that ?
09:15:11 bauzas honestly, I think so
09:15:32 bauzas the user can't know why nova doesn't work
09:15:41 bauzas so he could ask the operator to look at the problem
09:16:19 bauzas forcing to just delete the DB value while we would still have the mount could be a problem for a host after some time
09:17:00 bauzas and if the operator doesn't know that all the shares are still mounted, then he wouldn't see it until it could be a larger problem
09:17:57 Uggla ok fyi this is the msg given to user: {"badRequest": {"code": 400, "message": "Share id e8debdc0-447a-4376-a10a-4cd9122d7986 mount error from server 36a6c053-78c6-4409-9a44-b1e81244e61e.\\nReason: Unexpected error while running command.\\nCommand: mount\\nExit code: 1\\nStdout: \'This is stdout\'\\nStderr: \'This is stderror\'."}}
09:18:23 Uggla oops copy/paste the wrong one
09:18:37 Uggla s/mount/umount/
09:19:14 gibi if the command line or the error does not leak infra information to the user then I'm OK with this response
09:19:17 tobias-urdin should probably censor the command error and return a more generic error msg, had the same for ceph where it leaked the username in the error message returned to the user
09:19:29 tobias-urdin gibi: was faster :p
09:19:32 bauzas yeah agreed with tobias-urdin
09:19:37 gibi just tiny bit :)\
09:19:44 bauzas I'd prefer to have a better error
09:19:51 bauzas and, shouldn't be 400
09:20:07 bauzas that's not a *bad request* right?
09:20:40 gibi bauzas: good point this cannot be fixed by fixing the request
09:20:45 gibi so it is more like 500
09:21:56 bauzas Uggla: remind me something
09:22:06 bauzas Uggla: do we have share statuses ?
09:22:23 Uggla ok so do we agree that log should contains the reason (mount, stderr, stdout ...)
09:22:33 bauzas for the log, yes
09:22:39 Uggla yep we have a share_mapping status status
09:22:42 bauzas not for the error message to the user :)
09:23:20 bauzas like, say a user is asking to delete a share for an instance
09:23:39 bauzas he/she calls nova API for "deleting the share mapping"
09:23:53 bauzas he eventually gets "sorry, nay"
09:24:02 bauzas then the share mapping still exists
09:24:16 bauzas he then asks again "please delete my share mapping"
09:24:24 bauzas nova continues to tell "sorry, nay again"
09:24:56 bauzas then, the user would ask the operator to tell him/her "meh, can't delete my share mapping"
09:25:14 gibi ^^ +1
09:25:15 bauzas then the operator looks at the share mapping by the logs and see
09:25:29 bauzas 'oh man, this mapping UUID got an exception"
09:25:57 bauzas and then he/she says "oh, f*** that's why, this crazy umount didn't work"
09:26:18 bauzas that's how I see it
09:26:22 Uggla bauzas, sounds good to me and easier to manage. :)
09:27:34 bauzas the better would be to see the share mapping requests in the server actions list
09:27:38 bauzas Uggla: https://docs.openstack.org/api-ref/compute/#list-actions-for-server
09:27:50 bauzas but I don't think this is possible
09:28:16 bauzas ideally, if the operator could get the request ID of the "delete share mapping" call, that would be loving
09:28:45 bauzas and again, that's a question about whether we want to have a share mapping deletion to be synchronous or async
09:29:10 Uggla sync
09:29:29 bauzas then 500
09:29:43 Uggla ok

Earlier   Later