| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-01 | |||
| 16:03:28 | sean-k-mooney | dansmith: well that is currently blocked by the trait today | |
| 16:03:29 | dansmith | sean-k-mooney: that also means that if you got a configdrive when you created the instance, it will never have update-able user_data right? | |
| 16:03:39 | sean-k-mooney | but if we tried to do it quick that could break | |
| 16:03:54 | sean-k-mooney | dansmith: yes exactly | |
| 16:04:25 | dansmith | sean-k-mooney: no I get that we can detect if it's supported today, I'm talking about if we merge this and next cycle ironic says "we literally can't regenerate the config drive" and hyperv says 'we can, but not in the middle of a reboot because of how we implement that" | |
| 16:04:29 | dansmith | then we're kinda stuck | |
| 16:05:05 | sean-k-mooney | ya fair | |
| 16:05:35 | sean-k-mooney | its internal to the driver but it might for exampel require the current reboot to be slit into stop,update config drive, start | |
| 16:05:42 | sean-k-mooney | for the hyperv senario | |
| 16:05:56 | sean-k-mooney | which we might or might be able to hide | |
| 16:06:02 | dansmith | libvirt's hard reboot is basically a recreate, but that doesn't mean the other drivers are | |
| 16:06:17 | sean-k-mooney | yep | |
| 16:06:33 | dansmith | let me also say that I'm sorry I brought up all these concerns late, but I *was* asked to review this and have tried to put my money where my mouth is on changes | |
| 16:06:38 | dansmith | but I think these are all legit concerns | |
| 16:07:17 | bauzas | yeah, there are no easy paths for solving this problem | |
| 16:07:19 | dansmith | I know gibi is plotting my murder right now, probably conspiring with jhartkopf :) | |
| 16:07:25 | sean-k-mooney | they are. we discussed it in the ptg as we had previous rejected the spec | |
| 16:07:29 | sean-k-mooney | last cycle | |
| 16:08:06 | gibi | I'm sorry that I drove jhartkopf's solution to a dead end. | |
| 16:08:08 | bauzas | well, I had concerns about the complexity it was creating for little gain | |
| 16:08:21 | bauzas | but I was opposed this was an easy win | |
| 16:09:04 | gibi | I reviewed these patches and missed obvious design errors. I will try better next time. | |
| 16:09:04 | bauzas | so, now, I'm trying to find a trade-off but I don't wanna pull the strings if I think this is risky | |
| 16:09:08 | sean-k-mooney | well little gain is not neeisaly fiar it makes nova more "cloud native" as the idea was that user data shoudl be more liek k8s config maps | |
| 16:09:35 | sean-k-mooney | alhtoguh to be fiare that woudl also imply we shoudl delete and recreate the vms | |
| 16:09:44 | dansmith | hah right | |
| 16:09:57 | dansmith | no problem updating user_data if we just shoot the instances in the head :D | |
| 16:09:58 | bauzas | sean-k-mooney: we never proposed userdata to be mutable | |
| 16:10:12 | sean-k-mooney | who is we | |
| 16:10:33 | bauzas | in a cloud, you just spin another instance if you dislike your existing userdata | |
| 16:10:34 | sean-k-mooney | this has been a long runing request for multiple cycles | |
| 16:10:44 | sean-k-mooney | not in aws | |
| 16:10:50 | sean-k-mooney | and other plathforms | |
| 16:10:52 | jhartkopf | dansmith: Honestly it's better to notice problems now than when it's too late | |
| 16:10:57 | sean-k-mooney | apprenly openstack was an outlier | |
| 16:11:15 | bauzas | https://docs.openstack.org/nova/latest/user/metadata.html#user-provided-data | |
| 16:11:44 | gibi | dansmith: I would I? I feel sorry for jhartkopf's time spent on this. And I feel bad about that I was not able to find the issues you and sean-k-mooney found | |
| 16:11:57 | gibi | *why wouldi? | |
| 16:12:04 | bauzas | glad I'm not quoted :) | |
| 16:12:04 | sean-k-mooney | bauzas: right but this concept was orginally borrowed form ec2 | |
| 16:12:46 | sean-k-mooney | bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html | |
| 16:13:50 | sean-k-mooney | https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change | |
| 16:14:16 | bauzas | the stop requirement from EC2 may help | |
| 16:15:09 | bauzas | anyway, I was about to draft a design modification | |
| 16:15:46 | sean-k-mooney | well doing it on hard reboot was an optimisation of stop update start | |
| 16:15:52 | sean-k-mooney | but we could do htat instead | |
| 16:15:56 | sean-k-mooney | or shelve | |
| 16:16:21 | bauzas | the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated | |
| 16:16:33 | sean-k-mooney | unless server update is blocking right now im not sure it helps | |
| 16:16:38 | bauzas | at least once the request ended | |
| 16:17:04 | sean-k-mooney | bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive | |
| 16:17:06 | dansmith | we still have to regen the config drive though | |
| 16:17:10 | dansmith | which we don't do on reboot right? | |
| 16:17:14 | dansmith | we, stop..start | |
| 16:17:32 | sean-k-mooney | right we dont regen the config drive out side of unshelve or corss cell migration | |
| 16:17:40 | dansmith | but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments | |
| 16:17:43 | bauzas | phew | |
| 16:17:43 | sean-k-mooney | and unshelve only for non rbd backed instnaces | |
| 16:17:52 | dansmith | it addresses the "so do I need to reboot?" question on non-configdrive | |
| 16:18:07 | dansmith | and eliminates the need for the rpc | |
| 16:18:23 | dansmith | it still requires regen to work, and still may or may not be easy or possible for some drivers | |
| 16:18:35 | dansmith | like ironic might have to change the provision state to re-write a disk or something | |
| 16:18:35 | gibi | will we regen for each start then? | |
| 16:19:08 | dansmith | gibi: that'd be a question yeah | |
| 16:19:21 | sean-k-mooney | im not sure we woudl remove the rpc | |
| 16:19:21 | dansmith | I really don't like the idea of a dirty flag for the other reasons | |
| 16:19:34 | dansmith | sean-k-mooney: oh you mean a new rpc for the update | |
| 16:19:39 | dansmith | yeah that could work | |
| 16:19:40 | sean-k-mooney | ya | |
| 16:19:56 | dansmith | we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track | |
| 16:19:57 | sean-k-mooney | i mean we have another option | |
| 16:20:08 | sean-k-mooney | make this entirly uer driver by adding a new instance action | |
| 16:20:18 | sean-k-mooney | for updating the config drive | |
| 16:20:33 | dansmith | oh instead of a PUT | |
| 16:20:49 | bauzas | I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded | |
| 16:21:04 | sean-k-mooney | ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start | |
| 16:21:11 | bauzas | because we're discussing a design change | |
| 16:21:40 | bauzas | dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things" | |
| 16:21:42 | sean-k-mooney | bauzas: i think we all realise that we likely wont make progress now | |
| 16:22:07 | bauzas | an instance action would be nice to me, because it would get a status | |
| 16:22:21 | bauzas | and the async thing would be simple for the user | |
| 16:22:35 | bauzas | like, I started again and my userdata didn't updated, why ? | |
| 16:22:46 | bauzas | => server action tells me it failed | |
| 16:23:08 | sean-k-mooney | well the server action can be blocking or async | |
| 16:23:25 | sean-k-mooney | but it has a feeback mecahnim either way | |
| 16:23:26 | gibi | this could be a very long running blocking call | |
| 16:23:28 | bauzas | yup, the whole point is that it gives a clue why the userdata wasn't uptodate | |
| 16:23:45 | sean-k-mooney | right it woudl be in the instance event log | |
| 16:23:59 | bauzas | and until the userdata says "I'm updated", don't expect to get the new things if you start | |
| 16:24:01 | gibi | we can also make the start to fail if it needed to regen but failed to | |
| 16:24:02 | sean-k-mooney | if its async or in the respocne if blocking | |
| 16:24:22 | sean-k-mooney | gibi: that woudl requrie trackign the dirty state | |
| 16:24:34 | sean-k-mooney | and also some peopel might not care | |
| 16:25:16 | gibi | I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted | |
| 16:25:19 | gibi | started | |
| 16:25:48 | gibi | I think the dirty flag was bad because of shadow rpc | |
| 16:25:58 | sean-k-mooney | im thinking about something else | |
| 16:26:14 | sean-k-mooney | as a user i might not know if the user data is in config drive or not | |
| 16:26:26 | sean-k-mooney | but i said start | |
| 16:26:38 | sean-k-mooney | and i might not want tha to break in this case | |