Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-27
13:18:52 bauzas elodilles: I'm not bad about it fwiw, as we would have the Summit *before* Specfreeze which is nice
13:44:38 opendevreview Jorge San Emeterio proposed openstack/nova master: DNM: Testing check pipeline on master branch https://review.opendev.org/c/openstack/nova/+/872018
13:52:24 elodilles bauzas: yepp, it's a fairly long cycle (28 weeks) so milestones have longer times
14:12:07 darkhorse Hi team, Is it possible to disallow a user from all compute resources using keystone policy or any other means? The use case is I want to have multiple roles to a project. Some users have compute permissions, some have network permissions, some have volume permissions etc.
14:12:36 darkhorse I am not sure if this question is more relevant to keystone channel.
15:54:47 opendevreview Alexey Stupnikov proposed openstack/nova stable/victoria: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833436
16:01:03 opendevreview Merged openstack/nova stable/yoga: Add a workaround to skip hypervisor version check on LM https://review.opendev.org/c/openstack/nova/+/851202
16:02:42 opendevreview Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437
16:02:58 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:03:45 opendevreview Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437
16:04:03 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:04:21 opendevreview Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438
16:14:33 sean-k-mooney dansmith: can you take a look at https://review.opendev.org/c/openstack/nova/+/863919/12/nova/tests/unit/compute/test_resource_tracker.py#1553 when you have time
16:15:26 dansmith sean-k-mooney: yeah, it's on my list.. ya'll drug me into a call at 6:30am and I have another starting in a few ;P
16:15:48 sean-k-mooney sorry about that
16:15:49 dansmith thanks for hitting the ones below, that's some progress for sure
16:16:25 sean-k-mooney so ill try and test teh rest when i get devstack deploy so ill loop back to them again later today if that works out
16:16:39 dansmith cool
17:10:51 sean-k-mooney cool devstack stacked first time
17:11:24 sean-k-mooney now ot actully use your patches
17:12:26 sean-k-mooney so currently i have all in one so when i checkout your branch and restart the nova serices im expecting the compute to start up and writeh the uuid to a file and update the service version right
17:16:39 sean-k-mooney yep that worked https://paste.opendev.org/show/bLMvRfuUugp9iMcn4R60/
17:17:47 sean-k-mooney im going to break this in a few differnt ways and document it in an ether pad and ill let you know if i find anything odd or broken
17:25:44 dansmith sean-k-mooney: yeah, cool
17:29:59 sean-k-mooney https://etherpad.opendev.org/p/Stable-compute-uuid-manual-testing ill be usign that to keep my notes
17:30:37 sean-k-mooney there is not really much there yet but ill kep adding to it as i go
17:59:27 sean-k-mooney dansmith: found a bug
17:59:36 sean-k-mooney getting logs now
18:00:02 dansmith in "test 3 restart with deleted file" ?
18:00:08 sean-k-mooney yep
18:00:38 dansmith does that mean deleted node id file?
18:00:53 sean-k-mooney i moved it to compute_id_old
18:01:11 sean-k-mooney the agent tried to add a duplicate row in the db which resulted in a key error
18:01:15 sean-k-mooney and it wrote a new file
18:01:22 sean-k-mooney with a diffent uuid
18:01:50 dansmith well, right,
18:02:08 dansmith because it knows it's not an upgrade (now) and it didn't find a local compute node uuid
18:02:14 dansmith what do you think it should do there?
18:02:45 sean-k-mooney well it shoudl not result in a traceback in the log for one
18:03:02 sean-k-mooney but it should see that the compute node exist in the db and not try to create a new one
18:03:12 sean-k-mooney and we should not write a new file to disk
18:03:26 dansmith but the whole point of this is that the file becomes the source of truth
18:04:00 dansmith so maybe we should see the failure to create with a keyerror as a "something's wrong abort" but I bet it happens too late for us to abort
18:04:28 sean-k-mooney well the file didnt exist correct
18:04:48 sean-k-mooney so before we try to create a new compute node record shoudl we not check if one exits the old way
18:04:55 sean-k-mooney and abort then
18:05:00 dansmith we do, specifically on upgrade only
18:05:05 dansmith which you saw work
18:05:24 dansmith but once you have gotten past upgrade, it knows you're not upgrading, and can only really assume that it's a greenfield deployment
18:05:36 dansmith if we're going to stop relying on the hostname, there's really no other way, right?
18:06:17 sean-k-mooney no if we have a compute node and the serivce is upgraed then the file shoudl exist
18:06:25 sean-k-mooney if it does not then we knwo somethign odd happened
18:07:11 sean-k-mooney basicly if compute service>=62? and not file -> error
18:07:31 sean-k-mooney for greanfiled thre wont be a compute node in the db
18:07:31 dansmith how do we know the difference between "already upgraded" and "greenfield" ?
18:07:53 dansmith but again, the only way you know the "comptue node in the db" is because you're relying on the hostname, which we should not be doing
18:08:01 sean-k-mooney or we can catch the duplicate key error form the db
18:08:16 dansmith depending on when the duplicate key happens, I'm fine aborting in that case for sure
18:08:35 sean-k-mooney we at least shoudl not write the file with the wrong uuid
18:08:43 dansmith but I think we don't create that record until far too late
18:08:43 sean-k-mooney its currently writing the one for the row that was rejected
18:08:53 sean-k-mooney proably because we set the singlton uuid
18:09:24 dansmith okay, I'm with you on not writing it if there's a keyerror, but I think that will be pretty hard to line those things up
18:09:35 dansmith so maybe delete it if we wrote it or something
18:10:01 dansmith although, hang on
18:10:22 dansmith let's say I deploy two new computes with the same hostname by accident, they both generate and write unique uuids,
18:10:30 dansmith one will fail to create their compute node because of the clash,
18:10:48 dansmith but if I just rename the offending duplicate one, then the uuid it generated is fine to use when I restart
18:10:55 sean-k-mooney yep
18:11:14 dansmith the "I deleted the id" case seems pretty edge-y to me, because there are tons of things I can randomly delete from a running system that will break stuff
18:11:25 dansmith if we can abort on key conflict at start, then I'm on board, but otherwise, I dunno
18:11:48 dansmith this is a bit like saying I deleted /etc/ssh/ssh_host_key and when I restarted it generated a new one and my clients all complain :)
18:12:05 sean-k-mooney honestly i did this as my first test because i tough it could be a trivial thing that might happen and we should prevent it
18:12:24 sean-k-mooney dansmith: i was more thinging what happens if you fail to bind mount this in a container
18:13:26 sean-k-mooney if the file is ever lost i think it defeats much of the utility of the feature if we allow the agent to start
18:13:43 dansmith okay, but if you do, there's not much harm, because the resolution is to restart with it, and no harm no foul right?
18:14:00 dansmith if you failed to bind mount this, you likely also failed to bind-mount the instance images no?
18:14:17 sean-k-mooney fair its in the state dir
18:14:18 dansmith in the pre-provisioned case, this goes in /etc/nova anyway
18:14:22 sean-k-mooney although it depend on the location
18:17:00 sean-k-mooney any way im goign to keep testing/breaking it for a while and see what else i find
18:17:09 sean-k-mooney i just added the traceback
18:17:29 dansmith ack, I say we punt on that for the moment and see what else you can find
18:17:47 dansmith maybe we can circle back with some extra stuff that makes that smarter
18:18:02 dansmith at least catching keyerror and logging something that says "okay, here's how you've screwed up..."
18:18:18 dansmith because we have that trace right now and it's not super obvious why
18:26:47 sean-k-mooney so i think the abort is broken in general
18:28:30 dansmith the abort on rename?
18:28:47 sean-k-mooney yep so i tried restarting it with the incorrect uuid
18:29:03 sean-k-mooney and it just keeps trying to create the recoed with the incorrect uuid
18:29:07 sean-k-mooney which keeps failing
18:29:10 dansmith does the uuid exist in the db though?
18:29:21 sean-k-mooney ill check but i dont think so
18:29:30 dansmith then it is doing what it should do
18:29:35 dansmith because that's the pre-provisioned case
18:29:56 dansmith if you create an object in the db with that uuid but a different host, then it should trigger the rename detection
18:30:15 sean-k-mooney so i really think this is broken as is
18:30:37 sean-k-mooney but i will try renames later
18:30:47 dansmith if you give it a uuid that doesn't exist in the database, how is it supposed to know that's not what you want it to use?
18:31:25 dansmith you've told it "this is your uuid, end of story" and if there's no object in the db with that uuid, it's going to try to create it

Earlier   Later