Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-01
16:09:04 gibi I reviewed these patches and missed obvious design errors. I will try better next time.
16:09:04 bauzas so, now, I'm trying to find a trade-off but I don't wanna pull the strings if I think this is risky
16:09:08 sean-k-mooney well little gain is not neeisaly fiar it makes nova more "cloud native" as the idea was that user data shoudl be more liek k8s config maps
16:09:35 sean-k-mooney alhtoguh to be fiare that woudl also imply we shoudl delete and recreate the vms
16:09:44 dansmith hah right
16:09:57 dansmith no problem updating user_data if we just shoot the instances in the head :D
16:09:58 bauzas sean-k-mooney: we never proposed userdata to be mutable
16:10:12 sean-k-mooney who is we
16:10:33 bauzas in a cloud, you just spin another instance if you dislike your existing userdata
16:10:34 sean-k-mooney this has been a long runing request for multiple cycles
16:10:44 sean-k-mooney not in aws
16:10:50 sean-k-mooney and other plathforms
16:10:52 jhartkopf dansmith: Honestly it's better to notice problems now than when it's too late
16:10:57 sean-k-mooney apprenly openstack was an outlier
16:11:15 bauzas https://docs.openstack.org/nova/latest/user/metadata.html#user-provided-data
16:11:44 gibi dansmith: I would I? I feel sorry for jhartkopf's time spent on this. And I feel bad about that I was not able to find the issues you and sean-k-mooney found
16:11:57 gibi *why wouldi?
16:12:04 bauzas glad I'm not quoted :)
16:12:04 sean-k-mooney bauzas: right but this concept was orginally borrowed form ec2
16:12:46 sean-k-mooney bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html
16:13:50 sean-k-mooney https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change
16:14:16 bauzas the stop requirement from EC2 may help
16:15:09 bauzas anyway, I was about to draft a design modification
16:15:46 sean-k-mooney well doing it on hard reboot was an optimisation of stop update start
16:15:52 sean-k-mooney but we could do htat instead
16:15:56 sean-k-mooney or shelve
16:16:21 bauzas the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated
16:16:33 sean-k-mooney unless server update is blocking right now im not sure it helps
16:16:38 bauzas at least once the request ended
16:17:04 sean-k-mooney bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive
16:17:06 dansmith we still have to regen the config drive though
16:17:10 dansmith which we don't do on reboot right?
16:17:14 dansmith we, stop..start
16:17:32 sean-k-mooney right we dont regen the config drive out side of unshelve or corss cell migration
16:17:40 dansmith but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments
16:17:43 bauzas phew
16:17:43 sean-k-mooney and unshelve only for non rbd backed instnaces
16:17:52 dansmith it addresses the "so do I need to reboot?" question on non-configdrive
16:18:07 dansmith and eliminates the need for the rpc
16:18:23 dansmith it still requires regen to work, and still may or may not be easy or possible for some drivers
16:18:35 dansmith like ironic might have to change the provision state to re-write a disk or something
16:18:35 gibi will we regen for each start then?
16:19:08 dansmith gibi: that'd be a question yeah
16:19:21 sean-k-mooney im not sure we woudl remove the rpc
16:19:21 dansmith I really don't like the idea of a dirty flag for the other reasons
16:19:34 dansmith sean-k-mooney: oh you mean a new rpc for the update
16:19:39 dansmith yeah that could work
16:19:40 sean-k-mooney ya
16:19:56 dansmith we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track
16:19:57 sean-k-mooney i mean we have another option
16:20:08 sean-k-mooney make this entirly uer driver by adding a new instance action
16:20:18 sean-k-mooney for updating the config drive
16:20:33 dansmith oh instead of a PUT
16:20:49 bauzas I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded
16:21:04 sean-k-mooney ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start
16:21:11 bauzas because we're discussing a design change
16:21:40 bauzas dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things"
16:21:42 sean-k-mooney bauzas: i think we all realise that we likely wont make progress now
16:22:07 bauzas an instance action would be nice to me, because it would get a status
16:22:21 bauzas and the async thing would be simple for the user
16:22:35 bauzas like, I started again and my userdata didn't updated, why ?
16:22:46 bauzas => server action tells me it failed
16:23:08 sean-k-mooney well the server action can be blocking or async
16:23:25 sean-k-mooney but it has a feeback mecahnim either way
16:23:26 gibi this could be a very long running blocking call
16:23:28 bauzas yup, the whole point is that it gives a clue why the userdata wasn't uptodate
16:23:45 sean-k-mooney right it woudl be in the instance event log
16:23:59 bauzas and until the userdata says "I'm updated", don't expect to get the new things if you start
16:24:01 gibi we can also make the start to fail if it needed to regen but failed to
16:24:02 sean-k-mooney if its async or in the respocne if blocking
16:24:22 sean-k-mooney gibi: that woudl requrie trackign the dirty state
16:24:34 sean-k-mooney and also some peopel might not care
16:25:16 gibi I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted
16:25:19 gibi started
16:25:48 gibi I think the dirty flag was bad because of shadow rpc
16:25:58 sean-k-mooney im thinking about something else
16:26:14 sean-k-mooney as a user i might not know if the user data is in config drive or not
16:26:26 sean-k-mooney but i said start
16:26:38 sean-k-mooney and i might not want tha to break in this case
16:27:19 gibi I see
16:27:26 sean-k-mooney the dirty flag was to give you the behaior that it would only boot if it succeed to update
16:27:42 bauzas honestly, if I'm an user, I don't care about the internals
16:27:44 sean-k-mooney if it failed at least in my version teh vm would have gone to error
16:27:57 bauzas what I want to know is when my update succeeds and when I can use it
16:28:02 dansmith sorry, I'm getting pulled away,
16:28:06 dansmith but if it was an instance action,
16:28:18 dansmith then the rebuild either works or doesn't, I can see what happened,
16:28:29 dansmith and then I either get the old version on fail, or the new version on succeed
16:28:46 dansmith if we just set a dirty flag and then fail to regen on start, then we've killed the instance with no way to restart
16:28:48 dansmith which sucks
16:28:48 bauzas yup, I get feedback from an user perspective
16:28:56 bauzas yup, agreed
16:29:13 sean-k-mooney jhartkopf: so based on the above
16:29:15 dansmith we probably need some task_state or something to keep the user from starting it until the regen completes though
16:29:22 sean-k-mooney jhartkopf: and the fact we are all runing out of steam
16:29:28 bauzas mikal would say "spec'y spec'y"
16:29:31 dansmith or just use instance lock on the compute node to hold off on any other requests.. I guess that's enough
16:29:32 sean-k-mooney i think we need to disucss this again next cycle
16:29:53 bauzas the next PTG will be virtual
16:30:05 bauzas anyone can join, budged saved

Earlier   Later