Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-01
16:11:57 gibi *why wouldi?
16:12:04 bauzas glad I'm not quoted :)
16:12:04 sean-k-mooney bauzas: right but this concept was orginally borrowed form ec2
16:12:46 sean-k-mooney bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html
16:13:50 sean-k-mooney https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change
16:14:16 bauzas the stop requirement from EC2 may help
16:15:09 bauzas anyway, I was about to draft a design modification
16:15:46 sean-k-mooney well doing it on hard reboot was an optimisation of stop update start
16:15:52 sean-k-mooney but we could do htat instead
16:15:56 sean-k-mooney or shelve
16:16:21 bauzas the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated
16:16:33 sean-k-mooney unless server update is blocking right now im not sure it helps
16:16:38 bauzas at least once the request ended
16:17:04 sean-k-mooney bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive
16:17:06 dansmith we still have to regen the config drive though
16:17:10 dansmith which we don't do on reboot right?
16:17:14 dansmith we, stop..start
16:17:32 sean-k-mooney right we dont regen the config drive out side of unshelve or corss cell migration
16:17:40 dansmith but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments
16:17:43 bauzas phew
16:17:43 sean-k-mooney and unshelve only for non rbd backed instnaces
16:17:52 dansmith it addresses the "so do I need to reboot?" question on non-configdrive
16:18:07 dansmith and eliminates the need for the rpc
16:18:23 dansmith it still requires regen to work, and still may or may not be easy or possible for some drivers
16:18:35 dansmith like ironic might have to change the provision state to re-write a disk or something
16:18:35 gibi will we regen for each start then?
16:19:08 dansmith gibi: that'd be a question yeah
16:19:21 sean-k-mooney im not sure we woudl remove the rpc
16:19:21 dansmith I really don't like the idea of a dirty flag for the other reasons
16:19:34 dansmith sean-k-mooney: oh you mean a new rpc for the update
16:19:39 dansmith yeah that could work
16:19:40 sean-k-mooney ya
16:19:56 dansmith we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track
16:19:57 sean-k-mooney i mean we have another option
16:20:08 sean-k-mooney make this entirly uer driver by adding a new instance action
16:20:18 sean-k-mooney for updating the config drive
16:20:33 dansmith oh instead of a PUT
16:20:49 bauzas I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded
16:21:04 sean-k-mooney ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start
16:21:11 bauzas because we're discussing a design change
16:21:40 bauzas dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things"
16:21:42 sean-k-mooney bauzas: i think we all realise that we likely wont make progress now
16:22:07 bauzas an instance action would be nice to me, because it would get a status
16:22:21 bauzas and the async thing would be simple for the user
16:22:35 bauzas like, I started again and my userdata didn't updated, why ?
16:22:46 bauzas => server action tells me it failed
16:23:08 sean-k-mooney well the server action can be blocking or async
16:23:25 sean-k-mooney but it has a feeback mecahnim either way
16:23:26 gibi this could be a very long running blocking call
16:23:28 bauzas yup, the whole point is that it gives a clue why the userdata wasn't uptodate
16:23:45 sean-k-mooney right it woudl be in the instance event log
16:23:59 bauzas and until the userdata says "I'm updated", don't expect to get the new things if you start
16:24:01 gibi we can also make the start to fail if it needed to regen but failed to
16:24:02 sean-k-mooney if its async or in the respocne if blocking
16:24:22 sean-k-mooney gibi: that woudl requrie trackign the dirty state
16:24:34 sean-k-mooney and also some peopel might not care
16:25:16 gibi I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted
16:25:19 gibi started
16:25:48 gibi I think the dirty flag was bad because of shadow rpc
16:25:58 sean-k-mooney im thinking about something else
16:26:14 sean-k-mooney as a user i might not know if the user data is in config drive or not
16:26:26 sean-k-mooney but i said start
16:26:38 sean-k-mooney and i might not want tha to break in this case
16:27:19 gibi I see
16:27:26 sean-k-mooney the dirty flag was to give you the behaior that it would only boot if it succeed to update
16:27:42 bauzas honestly, if I'm an user, I don't care about the internals
16:27:44 sean-k-mooney if it failed at least in my version teh vm would have gone to error
16:27:57 bauzas what I want to know is when my update succeeds and when I can use it
16:28:02 dansmith sorry, I'm getting pulled away,
16:28:06 dansmith but if it was an instance action,
16:28:18 dansmith then the rebuild either works or doesn't, I can see what happened,
16:28:29 dansmith and then I either get the old version on fail, or the new version on succeed
16:28:46 dansmith if we just set a dirty flag and then fail to regen on start, then we've killed the instance with no way to restart
16:28:48 dansmith which sucks
16:28:48 bauzas yup, I get feedback from an user perspective
16:28:56 bauzas yup, agreed
16:29:13 sean-k-mooney jhartkopf: so based on the above
16:29:15 dansmith we probably need some task_state or something to keep the user from starting it until the regen completes though
16:29:22 sean-k-mooney jhartkopf: and the fact we are all runing out of steam
16:29:28 bauzas mikal would say "spec'y spec'y"
16:29:31 dansmith or just use instance lock on the compute node to hold off on any other requests.. I guess that's enough
16:29:32 sean-k-mooney i think we need to disucss this again next cycle
16:29:53 bauzas the next PTG will be virtual
16:30:05 bauzas anyone can join, budged saved
16:30:24 sean-k-mooney ya but we can reopne the sepc and discuss it again before htat
16:30:31 bauzas oh yeah
16:30:58 sean-k-mooney i see to ways we coudl go with the instance action by the way
16:31:05 sean-k-mooney one is one for regeneratin cofnig drive
16:31:11 sean-k-mooney the other is more targed for this feature
16:31:21 sean-k-mooney an instance action to update user_data
16:31:30 sean-k-mooney that either succeed or does not
16:31:49 sean-k-mooney and have that be the only wya to update it
16:31:57 sean-k-mooney so no updating it via server put
16:32:25 dansmith yeah, the user should not see "regenerate config drive"
16:32:41 dansmith so it should be "update my user_data" at the external surface for sure
16:32:52 sean-k-mooney right
16:32:55 sean-k-mooney so if we do that
16:33:16 sean-k-mooney that dedicated instance action can "do the right thing" regardess of how the instance is booted
16:33:31 dansmith right
16:33:42 sean-k-mooney and returna 409 conflict if its not supprot by the driver or whatever

Earlier   Later