Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-01
16:04:25 dansmith sean-k-mooney: no I get that we can detect if it's supported today, I'm talking about if we merge this and next cycle ironic says "we literally can't regenerate the config drive" and hyperv says 'we can, but not in the middle of a reboot because of how we implement that"
16:04:29 dansmith then we're kinda stuck
16:05:05 sean-k-mooney ya fair
16:05:35 sean-k-mooney its internal to the driver but it might for exampel require the current reboot to be slit into stop,update config drive, start
16:05:42 sean-k-mooney for the hyperv senario
16:05:56 sean-k-mooney which we might or might be able to hide
16:06:02 dansmith libvirt's hard reboot is basically a recreate, but that doesn't mean the other drivers are
16:06:17 sean-k-mooney yep
16:06:33 dansmith let me also say that I'm sorry I brought up all these concerns late, but I *was* asked to review this and have tried to put my money where my mouth is on changes
16:06:38 dansmith but I think these are all legit concerns
16:07:17 bauzas yeah, there are no easy paths for solving this problem
16:07:19 dansmith I know gibi is plotting my murder right now, probably conspiring with jhartkopf :)
16:07:25 sean-k-mooney they are. we discussed it in the ptg as we had previous rejected the spec
16:07:29 sean-k-mooney last cycle
16:08:06 gibi I'm sorry that I drove jhartkopf's solution to a dead end.
16:08:08 bauzas well, I had concerns about the complexity it was creating for little gain
16:08:21 bauzas but I was opposed this was an easy win
16:09:04 gibi I reviewed these patches and missed obvious design errors. I will try better next time.
16:09:04 bauzas so, now, I'm trying to find a trade-off but I don't wanna pull the strings if I think this is risky
16:09:08 sean-k-mooney well little gain is not neeisaly fiar it makes nova more "cloud native" as the idea was that user data shoudl be more liek k8s config maps
16:09:35 sean-k-mooney alhtoguh to be fiare that woudl also imply we shoudl delete and recreate the vms
16:09:44 dansmith hah right
16:09:57 dansmith no problem updating user_data if we just shoot the instances in the head :D
16:09:58 bauzas sean-k-mooney: we never proposed userdata to be mutable
16:10:12 sean-k-mooney who is we
16:10:33 bauzas in a cloud, you just spin another instance if you dislike your existing userdata
16:10:34 sean-k-mooney this has been a long runing request for multiple cycles
16:10:44 sean-k-mooney not in aws
16:10:50 sean-k-mooney and other plathforms
16:10:52 jhartkopf dansmith: Honestly it's better to notice problems now than when it's too late
16:10:57 sean-k-mooney apprenly openstack was an outlier
16:11:15 bauzas https://docs.openstack.org/nova/latest/user/metadata.html#user-provided-data
16:11:44 gibi dansmith: I would I? I feel sorry for jhartkopf's time spent on this. And I feel bad about that I was not able to find the issues you and sean-k-mooney found
16:11:57 gibi *why wouldi?
16:12:04 bauzas glad I'm not quoted :)
16:12:04 sean-k-mooney bauzas: right but this concept was orginally borrowed form ec2
16:12:46 sean-k-mooney bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html
16:13:50 sean-k-mooney https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change
16:14:16 bauzas the stop requirement from EC2 may help
16:15:09 bauzas anyway, I was about to draft a design modification
16:15:46 sean-k-mooney well doing it on hard reboot was an optimisation of stop update start
16:15:52 sean-k-mooney but we could do htat instead
16:15:56 sean-k-mooney or shelve
16:16:21 bauzas the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated
16:16:33 sean-k-mooney unless server update is blocking right now im not sure it helps
16:16:38 bauzas at least once the request ended
16:17:04 sean-k-mooney bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive
16:17:06 dansmith we still have to regen the config drive though
16:17:10 dansmith which we don't do on reboot right?
16:17:14 dansmith we, stop..start
16:17:32 sean-k-mooney right we dont regen the config drive out side of unshelve or corss cell migration
16:17:40 dansmith but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments
16:17:43 bauzas phew
16:17:43 sean-k-mooney and unshelve only for non rbd backed instnaces
16:17:52 dansmith it addresses the "so do I need to reboot?" question on non-configdrive
16:18:07 dansmith and eliminates the need for the rpc
16:18:23 dansmith it still requires regen to work, and still may or may not be easy or possible for some drivers
16:18:35 dansmith like ironic might have to change the provision state to re-write a disk or something
16:18:35 gibi will we regen for each start then?
16:19:08 dansmith gibi: that'd be a question yeah
16:19:21 sean-k-mooney im not sure we woudl remove the rpc
16:19:21 dansmith I really don't like the idea of a dirty flag for the other reasons
16:19:34 dansmith sean-k-mooney: oh you mean a new rpc for the update
16:19:39 dansmith yeah that could work
16:19:40 sean-k-mooney ya
16:19:56 dansmith we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track
16:19:57 sean-k-mooney i mean we have another option
16:20:08 sean-k-mooney make this entirly uer driver by adding a new instance action
16:20:18 sean-k-mooney for updating the config drive
16:20:33 dansmith oh instead of a PUT
16:20:49 bauzas I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded
16:21:04 sean-k-mooney ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start
16:21:11 bauzas because we're discussing a design change
16:21:40 bauzas dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things"
16:21:42 sean-k-mooney bauzas: i think we all realise that we likely wont make progress now
16:22:07 bauzas an instance action would be nice to me, because it would get a status
16:22:21 bauzas and the async thing would be simple for the user
16:22:35 bauzas like, I started again and my userdata didn't updated, why ?
16:22:46 bauzas => server action tells me it failed
16:23:08 sean-k-mooney well the server action can be blocking or async
16:23:25 sean-k-mooney but it has a feeback mecahnim either way
16:23:26 gibi this could be a very long running blocking call
16:23:28 bauzas yup, the whole point is that it gives a clue why the userdata wasn't uptodate
16:23:45 sean-k-mooney right it woudl be in the instance event log
16:23:59 bauzas and until the userdata says "I'm updated", don't expect to get the new things if you start
16:24:01 gibi we can also make the start to fail if it needed to regen but failed to
16:24:02 sean-k-mooney if its async or in the respocne if blocking
16:24:22 sean-k-mooney gibi: that woudl requrie trackign the dirty state
16:24:34 sean-k-mooney and also some peopel might not care
16:25:16 gibi I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted
16:25:19 gibi started
16:25:48 gibi I think the dirty flag was bad because of shadow rpc
16:25:58 sean-k-mooney im thinking about something else
16:26:14 sean-k-mooney as a user i might not know if the user data is in config drive or not
16:26:26 sean-k-mooney but i said start
16:26:38 sean-k-mooney and i might not want tha to break in this case
16:27:19 gibi I see
16:27:26 sean-k-mooney the dirty flag was to give you the behaior that it would only boot if it succeed to update
16:27:42 bauzas honestly, if I'm an user, I don't care about the internals
16:27:44 sean-k-mooney if it failed at least in my version teh vm would have gone to error

Earlier   Later