| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-01 | |||
| 16:11:57 | gibi | *why wouldi? | |
| 16:12:04 | bauzas | glad I'm not quoted :) | |
| 16:12:04 | sean-k-mooney | bauzas: right but this concept was orginally borrowed form ec2 | |
| 16:12:46 | sean-k-mooney | bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html | |
| 16:13:50 | sean-k-mooney | https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change | |
| 16:14:16 | bauzas | the stop requirement from EC2 may help | |
| 16:15:09 | bauzas | anyway, I was about to draft a design modification | |
| 16:15:46 | sean-k-mooney | well doing it on hard reboot was an optimisation of stop update start | |
| 16:15:52 | sean-k-mooney | but we could do htat instead | |
| 16:15:56 | sean-k-mooney | or shelve | |
| 16:16:21 | bauzas | the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated | |
| 16:16:33 | sean-k-mooney | unless server update is blocking right now im not sure it helps | |
| 16:16:38 | bauzas | at least once the request ended | |
| 16:17:04 | sean-k-mooney | bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive | |
| 16:17:06 | dansmith | we still have to regen the config drive though | |
| 16:17:10 | dansmith | which we don't do on reboot right? | |
| 16:17:14 | dansmith | we, stop..start | |
| 16:17:32 | sean-k-mooney | right we dont regen the config drive out side of unshelve or corss cell migration | |
| 16:17:40 | dansmith | but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments | |
| 16:17:43 | bauzas | phew | |
| 16:17:43 | sean-k-mooney | and unshelve only for non rbd backed instnaces | |
| 16:17:52 | dansmith | it addresses the "so do I need to reboot?" question on non-configdrive | |
| 16:18:07 | dansmith | and eliminates the need for the rpc | |
| 16:18:23 | dansmith | it still requires regen to work, and still may or may not be easy or possible for some drivers | |
| 16:18:35 | dansmith | like ironic might have to change the provision state to re-write a disk or something | |
| 16:18:35 | gibi | will we regen for each start then? | |
| 16:19:08 | dansmith | gibi: that'd be a question yeah | |
| 16:19:21 | sean-k-mooney | im not sure we woudl remove the rpc | |
| 16:19:21 | dansmith | I really don't like the idea of a dirty flag for the other reasons | |
| 16:19:34 | dansmith | sean-k-mooney: oh you mean a new rpc for the update | |
| 16:19:39 | dansmith | yeah that could work | |
| 16:19:40 | sean-k-mooney | ya | |
| 16:19:56 | dansmith | we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track | |
| 16:19:57 | sean-k-mooney | i mean we have another option | |
| 16:20:08 | sean-k-mooney | make this entirly uer driver by adding a new instance action | |
| 16:20:18 | sean-k-mooney | for updating the config drive | |
| 16:20:33 | dansmith | oh instead of a PUT | |
| 16:20:49 | bauzas | I know we're on Zed-3 day and I'm supposed to stay but this is 6:20pm here and I wonder how much of this could be etherpaded | |
| 16:21:04 | sean-k-mooney | ya, stop, update whatever you liek, explictly tell it to update config drive if you care, start | |
| 16:21:11 | bauzas | because we're discussing a design change | |
| 16:21:40 | bauzas | dansmith: yeah, something we would say "damn nova, please do the necessary to plumb my things" | |
| 16:21:42 | sean-k-mooney | bauzas: i think we all realise that we likely wont make progress now | |
| 16:22:07 | bauzas | an instance action would be nice to me, because it would get a status | |
| 16:22:21 | bauzas | and the async thing would be simple for the user | |
| 16:22:35 | bauzas | like, I started again and my userdata didn't updated, why ? | |
| 16:22:46 | bauzas | => server action tells me it failed | |
| 16:23:08 | sean-k-mooney | well the server action can be blocking or async | |
| 16:23:25 | sean-k-mooney | but it has a feeback mecahnim either way | |
| 16:23:26 | gibi | this could be a very long running blocking call | |
| 16:23:28 | bauzas | yup, the whole point is that it gives a clue why the userdata wasn't uptodate | |
| 16:23:45 | sean-k-mooney | right it woudl be in the instance event log | |
| 16:23:59 | bauzas | and until the userdata says "I'm updated", don't expect to get the new things if you start | |
| 16:24:01 | gibi | we can also make the start to fail if it needed to regen but failed to | |
| 16:24:02 | sean-k-mooney | if its async or in the respocne if blocking | |
| 16:24:22 | sean-k-mooney | gibi: that woudl requrie trackign the dirty state | |
| 16:24:34 | sean-k-mooney | and also some peopel might not care | |
| 16:25:16 | gibi | I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted | |
| 16:25:19 | gibi | started | |
| 16:25:48 | gibi | I think the dirty flag was bad because of shadow rpc | |
| 16:25:58 | sean-k-mooney | im thinking about something else | |
| 16:26:14 | sean-k-mooney | as a user i might not know if the user data is in config drive or not | |
| 16:26:26 | sean-k-mooney | but i said start | |
| 16:26:38 | sean-k-mooney | and i might not want tha to break in this case | |
| 16:27:19 | gibi | I see | |
| 16:27:26 | sean-k-mooney | the dirty flag was to give you the behaior that it would only boot if it succeed to update | |
| 16:27:42 | bauzas | honestly, if I'm an user, I don't care about the internals | |
| 16:27:44 | sean-k-mooney | if it failed at least in my version teh vm would have gone to error | |
| 16:27:57 | bauzas | what I want to know is when my update succeeds and when I can use it | |
| 16:28:02 | dansmith | sorry, I'm getting pulled away, | |
| 16:28:06 | dansmith | but if it was an instance action, | |
| 16:28:18 | dansmith | then the rebuild either works or doesn't, I can see what happened, | |
| 16:28:29 | dansmith | and then I either get the old version on fail, or the new version on succeed | |
| 16:28:46 | dansmith | if we just set a dirty flag and then fail to regen on start, then we've killed the instance with no way to restart | |
| 16:28:48 | dansmith | which sucks | |
| 16:28:48 | bauzas | yup, I get feedback from an user perspective | |
| 16:28:56 | bauzas | yup, agreed | |
| 16:29:13 | sean-k-mooney | jhartkopf: so based on the above | |
| 16:29:15 | dansmith | we probably need some task_state or something to keep the user from starting it until the regen completes though | |
| 16:29:22 | sean-k-mooney | jhartkopf: and the fact we are all runing out of steam | |
| 16:29:28 | bauzas | mikal would say "spec'y spec'y" | |
| 16:29:31 | dansmith | or just use instance lock on the compute node to hold off on any other requests.. I guess that's enough | |
| 16:29:32 | sean-k-mooney | i think we need to disucss this again next cycle | |
| 16:29:53 | bauzas | the next PTG will be virtual | |
| 16:30:05 | bauzas | anyone can join, budged saved | |
| 16:30:24 | sean-k-mooney | ya but we can reopne the sepc and discuss it again before htat | |
| 16:30:31 | bauzas | oh yeah | |
| 16:30:58 | sean-k-mooney | i see to ways we coudl go with the instance action by the way | |
| 16:31:05 | sean-k-mooney | one is one for regeneratin cofnig drive | |
| 16:31:11 | sean-k-mooney | the other is more targed for this feature | |
| 16:31:21 | sean-k-mooney | an instance action to update user_data | |
| 16:31:30 | sean-k-mooney | that either succeed or does not | |
| 16:31:49 | sean-k-mooney | and have that be the only wya to update it | |
| 16:31:57 | sean-k-mooney | so no updating it via server put | |
| 16:32:25 | dansmith | yeah, the user should not see "regenerate config drive" | |
| 16:32:41 | dansmith | so it should be "update my user_data" at the external surface for sure | |
| 16:32:52 | sean-k-mooney | right | |
| 16:32:55 | sean-k-mooney | so if we do that | |
| 16:33:16 | sean-k-mooney | that dedicated instance action can "do the right thing" regardess of how the instance is booted | |
| 16:33:31 | dansmith | right | |
| 16:33:42 | sean-k-mooney | and returna 409 conflict if its not supprot by the driver or whatever | |