| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-01 | |||
| 16:05:42 | sean-k-mooney | for the hyperv senario | |
| 16:05:56 | sean-k-mooney | which we might or might be able to hide | |
| 16:06:02 | dansmith | libvirt's hard reboot is basically a recreate, but that doesn't mean the other drivers are | |
| 16:06:17 | sean-k-mooney | yep | |
| 16:06:33 | dansmith | let me also say that I'm sorry I brought up all these concerns late, but I *was* asked to review this and have tried to put my money where my mouth is on changes | |
| 16:06:38 | dansmith | but I think these are all legit concerns | |
| 16:07:17 | bauzas | yeah, there are no easy paths for solving this problem | |
| 16:07:19 | dansmith | I know gibi is plotting my murder right now, probably conspiring with jhartkopf :) | |
| 16:07:25 | sean-k-mooney | they are. we discussed it in the ptg as we had previous rejected the spec | |
| 16:07:29 | sean-k-mooney | last cycle | |
| 16:08:06 | gibi | I'm sorry that I drove jhartkopf's solution to a dead end. | |
| 16:08:08 | bauzas | well, I had concerns about the complexity it was creating for little gain | |
| 16:08:21 | bauzas | but I was opposed this was an easy win | |
| 16:09:04 | gibi | I reviewed these patches and missed obvious design errors. I will try better next time. | |
| 16:09:04 | bauzas | so, now, I'm trying to find a trade-off but I don't wanna pull the strings if I think this is risky | |
| 16:09:08 | sean-k-mooney | well little gain is not neeisaly fiar it makes nova more "cloud native" as the idea was that user data shoudl be more liek k8s config maps | |
| 16:09:35 | sean-k-mooney | alhtoguh to be fiare that woudl also imply we shoudl delete and recreate the vms | |
| 16:09:44 | dansmith | hah right | |
| 16:09:57 | dansmith | no problem updating user_data if we just shoot the instances in the head :D | |
| 16:09:58 | bauzas | sean-k-mooney: we never proposed userdata to be mutable | |
| 16:10:12 | sean-k-mooney | who is we | |
| 16:10:33 | bauzas | in a cloud, you just spin another instance if you dislike your existing userdata | |
| 16:10:34 | sean-k-mooney | this has been a long runing request for multiple cycles | |
| 16:10:44 | sean-k-mooney | not in aws | |
| 16:10:50 | sean-k-mooney | and other plathforms | |
| 16:10:52 | jhartkopf | dansmith: Honestly it's better to notice problems now than when it's too late | |
| 16:10:57 | sean-k-mooney | apprenly openstack was an outlier | |
| 16:11:15 | bauzas | https://docs.openstack.org/nova/latest/user/metadata.html#user-provided-data | |
| 16:11:44 | gibi | dansmith: I would I? I feel sorry for jhartkopf's time spent on this. And I feel bad about that I was not able to find the issues you and sean-k-mooney found | |
| 16:11:57 | gibi | *why wouldi? | |
| 16:12:04 | bauzas | glad I'm not quoted :) | |
| 16:12:04 | sean-k-mooney | bauzas: right but this concept was orginally borrowed form ec2 | |
| 16:12:46 | sean-k-mooney | bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html | |
| 16:13:50 | sean-k-mooney | https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change | |
| 16:14:16 | bauzas | the stop requirement from EC2 may help | |
| 16:15:09 | bauzas | anyway, I was about to draft a design modification | |
| 16:15:46 | sean-k-mooney | well doing it on hard reboot was an optimisation of stop update start | |
| 16:15:52 | sean-k-mooney | but we could do htat instead | |
| 16:15:56 | sean-k-mooney | or shelve | |
| 16:16:21 | bauzas | the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated | |
| 16:16:33 | sean-k-mooney | unless server update is blocking right now im not sure it helps | |
| 16:16:38 | bauzas | at least once the request ended | |
| 16:17:04 | sean-k-mooney | bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive | |
| 16:17:06 | dansmith | we still have to regen the config drive though | |
| 16:17:10 | dansmith | which we don't do on reboot right? | |
| 16:17:14 | dansmith | we, stop..start | |
| 16:17:32 | sean-k-mooney | right we dont regen the config drive out side of unshelve or corss cell migration | |
| 16:17:40 | dansmith | but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments | |
| 16:17:43 | bauzas | phew | |
| 16:17:43 | sean-k-mooney | and unshelve only for non rbd backed instnaces | |
| 16:17:52 | dansmith | it addresses the "so do I need to reboot?" question on non-configdrive | |
| 16:18:07 | dansmith | and eliminates the need for the rpc | |
| 16:18:23 | dansmith | it still requires regen to work, and still may or may not be easy or possible for some drivers | |
| 16:18:35 | dansmith | like ironic might have to change the provision state to re-write a disk or something | |
| 16:18:35 | gibi | will we regen for each start then? | |
| 16:19:08 | dansmith | gibi: that'd be a question yeah | |
| 16:19:21 | sean-k-mooney | im not sure we woudl remove the rpc | |
| 16:19:21 | dansmith | I really don't like the idea of a dirty flag for the other reasons | |
| 16:19:34 | dansmith | sean-k-mooney: oh you mean a new rpc for the update | |
| 16:19:39 | dansmith | yeah that could work | |
| 16:19:40 | sean-k-mooney | ya | |
| 16:19:56 | dansmith | we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track | |
| 16:19:57 | sean-k-mooney | i mean we have another option | |
| 16:20:08 | sean-k-mooney | make this entirly uer driver by adding a new instance action | |
| 16:20:18 | sean-k-mooney | for updating the config drive | |
| 16:20:33 | dansmith | oh instead of a PUT | |
| 16:20:49 | bauzas | I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded | |
| 16:21:04 | sean-k-mooney | ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start | |
| 16:21:11 | bauzas | because we're discussing a design change | |
| 16:21:40 | bauzas | dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things" | |
| 16:21:42 | sean-k-mooney | bauzas: i think we all realise that we likely wont make progress now | |
| 16:22:07 | bauzas | an instance action would be nice to me, because it would get a status | |
| 16:22:21 | bauzas | and the async thing would be simple for the user | |
| 16:22:35 | bauzas | like, I started again and my userdata didn't updated, why ? | |
| 16:22:46 | bauzas | => server action tells me it failed | |
| 16:23:08 | sean-k-mooney | well the server action can be blocking or async | |
| 16:23:25 | sean-k-mooney | but it has a feeback mecahnim either way | |
| 16:23:26 | gibi | this could be a very long running blocking call | |
| 16:23:28 | bauzas | yup, the whole point is that it gives a clue why the userdata wasn't uptodate | |
| 16:23:45 | sean-k-mooney | right it woudl be in the instance event log | |
| 16:23:59 | bauzas | and until the userdata says "I'm updated", don't expect to get the new things if you start | |
| 16:24:01 | gibi | we can also make the start to fail if it needed to regen but failed to | |
| 16:24:02 | sean-k-mooney | if its async or in the respocne if blocking | |
| 16:24:22 | sean-k-mooney | gibi: that woudl requrie trackign the dirty state | |
| 16:24:34 | sean-k-mooney | and also some peopel might not care | |
| 16:25:16 | gibi | I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted | |
| 16:25:19 | gibi | started | |
| 16:25:48 | gibi | I think the dirty flag was bad because of shadow rpc | |
| 16:25:58 | sean-k-mooney | im thinking about something else | |
| 16:26:14 | sean-k-mooney | as a user i might not know if the user data is in config drive or not | |
| 16:26:26 | sean-k-mooney | but i said start | |
| 16:26:38 | sean-k-mooney | and i might not want tha to break in this case | |
| 16:27:19 | gibi | I see | |
| 16:27:26 | sean-k-mooney | the dirty flag was to give you the behaior that it would only boot if it succeed to update | |
| 16:27:42 | bauzas | honestly, if I'm an user, I don't care about the internals | |
| 16:27:44 | sean-k-mooney | if it failed at least in my version teh vm would have gone to error | |
| 16:27:57 | bauzas | what I want to know is when my update succeeds and when I can use it | |
| 16:28:02 | dansmith | sorry, I'm getting pulled away, | |
| 16:28:06 | dansmith | but if it was an instance action, | |
| 16:28:18 | dansmith | then the rebuild either works or doesn't, I can see what happened, | |