Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-27
13:52:24 elodilles bauzas: yepp, it's a fairly long cycle (28 weeks) so milestones have longer times
14:12:07 darkhorse Hi team, Is it possible to disallow a user from all compute resources using keystone policy or any other means? The use case is I want to have multiple roles to a project. Some users have compute permissions, some have network permissions, some have volume permissions etc.
14:12:36 darkhorse I am not sure if this question is more relevant to keystone channel.
15:54:47 opendevreview Alexey Stupnikov proposed openstack/nova stable/victoria: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833436
16:01:03 opendevreview Merged openstack/nova stable/yoga: Add a workaround to skip hypervisor version check on LM https://review.opendev.org/c/openstack/nova/+/851202
16:02:42 opendevreview Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437
16:02:58 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:03:45 opendevreview Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437
16:04:03 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:04:21 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:14:33 sean-k-mooney dansmith: can you take a look at https://review.opendev.org/c/openstack/nova/+/863919/12/nova/tests/unit/compute/test_resource_tracker.py#1553 when you have time
16:15:26 dansmith sean-k-mooney: yeah, it's on my list.. ya'll drug me into a call at 6:30am and I have another starting in a few ;P
16:15:48 sean-k-mooney sorry about that
16:15:49 dansmith thanks for hitting the ones below, that's some progress for sure
16:16:25 sean-k-mooney so ill try and test teh rest when i get devstack deploy so ill loop back to them again later today if that works out
16:16:39 dansmith cool
17:10:51 sean-k-mooney cool devstack stacked first time
17:11:24 sean-k-mooney now ot actully use your patches
17:12:26 sean-k-mooney so currently i have all in one so when i checkout your branch and restart the nova serices im expecting the compute to start up and writeh the uuid to a file and update the service version right
17:16:39 sean-k-mooney yep that worked https://paste.opendev.org/show/bLMvRfuUugp9iMcn4R60/
17:17:47 sean-k-mooney im going to break this in a few differnt ways and document it in an ether pad and ill let you know if i find anything odd or broken
17:25:44 dansmith sean-k-mooney: yeah, cool
17:29:59 sean-k-mooney https://etherpad.opendev.org/p/Stable-compute-uuid-manual-testing ill be usign that to keep my notes
17:30:37 sean-k-mooney there is not really much there yet but ill kep adding to it as i go
17:59:27 sean-k-mooney dansmith: found a bug
17:59:36 sean-k-mooney getting logs now
18:00:02 dansmith in "test 3 restart with deleted file" ?
18:00:08 sean-k-mooney yep
18:00:38 dansmith does that mean deleted node id file?
18:00:53 sean-k-mooney i moved it to compute_id_old
18:01:11 sean-k-mooney the agent tried to add a duplicate row in the db which resulted in a key error
18:01:15 sean-k-mooney and it wrote a new file
18:01:22 sean-k-mooney with a diffent uuid
18:01:50 dansmith well, right,
18:02:08 dansmith because it knows it's not an upgrade (now) and it didn't find a local compute node uuid
18:02:14 dansmith what do you think it should do there?
18:02:45 sean-k-mooney well it shoudl not result in a traceback in the log for one
18:03:02 sean-k-mooney but it should see that the compute node exist in the db and not try to create a new one
18:03:12 sean-k-mooney and we should not write a new file to disk
18:03:26 dansmith but the whole point of this is that the file becomes the source of truth
18:04:00 dansmith so maybe we should see the failure to create with a keyerror as a "something's wrong abort" but I bet it happens too late for us to abort
18:04:28 sean-k-mooney well the file didnt exist correct
18:04:48 sean-k-mooney so before we try to create a new compute node record shoudl we not check if one exits the old way
18:04:55 sean-k-mooney and abort then
18:05:00 dansmith we do, specifically on upgrade only
18:05:05 dansmith which you saw work
18:05:24 dansmith but once you have gotten past upgrade, it knows you're not upgrading, and can only really assume that it's a greenfield deployment
18:05:36 dansmith if we're going to stop relying on the hostname, there's really no other way, right?
18:06:17 sean-k-mooney no if we have a compute node and the serivce is upgraed then the file shoudl exist
18:06:25 sean-k-mooney if it does not then we knwo somethign odd happened
18:07:11 sean-k-mooney basicly if compute service>=62? and not file -> error
18:07:31 sean-k-mooney for greanfiled thre wont be a compute node in the db
18:07:31 dansmith how do we know the difference between "already upgraded" and "greenfield" ?
18:07:53 dansmith but again, the only way you know the "comptue node in the db" is because you're relying on the hostname, which we should not be doing
18:08:01 sean-k-mooney or we can catch the duplicate key error form the db
18:08:16 dansmith depending on when the duplicate key happens, I'm fine aborting in that case for sure
18:08:35 sean-k-mooney we at least shoudl not write the file with the wrong uuid
18:08:43 dansmith but I think we don't create that record until far too late
18:08:43 sean-k-mooney its currently writing the one for the row that was rejected
18:08:53 sean-k-mooney proably because we set the singlton uuid
18:09:24 dansmith okay, I'm with you on not writing it if there's a keyerror, but I think that will be pretty hard to line those things up
18:09:35 dansmith so maybe delete it if we wrote it or something
18:10:01 dansmith although, hang on
18:10:22 dansmith let's say I deploy two new computes with the same hostname by accident, they both generate and write unique uuids,
18:10:30 dansmith one will fail to create their compute node because of the clash,
18:10:48 dansmith but if I just rename the offending duplicate one, then the uuid it generated is fine to use when I restart
18:10:55 sean-k-mooney yep
18:11:14 dansmith the "I deleted the id" case seems pretty edge-y to me, because there are tons of things I can randomly delete from a running system that will break stuff
18:11:25 dansmith if we can abort on key conflict at start, then I'm on board, but otherwise, I dunno
18:11:48 dansmith this is a bit like saying I deleted /etc/ssh/ssh_host_key and when I restarted it generated a new one and my clients all complain :)
18:12:05 sean-k-mooney honestly i did this as my first test because i tough it could be a trivial thing that might happen and we should prevent it
18:12:24 sean-k-mooney dansmith: i was more thinging what happens if you fail to bind mount this in a container
18:13:26 sean-k-mooney if the file is ever lost i think it defeats much of the utility of the feature if we allow the agent to start
18:13:43 dansmith okay, but if you do, there's not much harm, because the resolution is to restart with it, and no harm no foul right?
18:14:00 dansmith if you failed to bind mount this, you likely also failed to bind-mount the instance images no?
18:14:17 sean-k-mooney fair its in the state dir
18:14:18 dansmith in the pre-provisioned case, this goes in /etc/nova anyway
18:14:22 sean-k-mooney although it depend on the location
18:17:00 sean-k-mooney any way im goign to keep testing/breaking it for a while and see what else i find
18:17:09 sean-k-mooney i just added the traceback
18:17:29 dansmith ack, I say we punt on that for the moment and see what else you can find
18:17:47 dansmith maybe we can circle back with some extra stuff that makes that smarter
18:18:02 dansmith at least catching keyerror and logging something that says "okay, here's how you've screwed up..."
18:18:18 dansmith because we have that trace right now and it's not super obvious why
18:26:47 sean-k-mooney so i think the abort is broken in general
18:28:30 dansmith the abort on rename?
18:28:47 sean-k-mooney yep so i tried restarting it with the incorrect uuid
18:29:03 sean-k-mooney and it just keeps trying to create the recoed with the incorrect uuid
18:29:07 sean-k-mooney which keeps failing
18:29:10 dansmith does the uuid exist in the db though?
18:29:21 sean-k-mooney ill check but i dont think so
18:29:30 dansmith then it is doing what it should do
18:29:35 dansmith because that's the pre-provisioned case
18:29:56 dansmith if you create an object in the db with that uuid but a different host, then it should trigger the rename detection
18:30:15 sean-k-mooney so i really think this is broken as is
18:30:37 sean-k-mooney but i will try renames later
18:30:47 dansmith if you give it a uuid that doesn't exist in the database, how is it supposed to know that's not what you want it to use?
18:31:25 dansmith you've told it "this is your uuid, end of story" and if there's no object in the db with that uuid, it's going to try to create it
18:31:51 dansmith it should catch and log a better error than the trace, saying there's a name conflict or something, but otherwise it's doing the right thing I think
18:31:57 sean-k-mooney so i dont think we shoudl be ignoring the host/hypervior_hostname entirely espcially when we get the db duplicte key error

Earlier   Later