| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-01 | |||
| 15:58:03 | sean-k-mooney | jhartkopf: i would prefer not to do that if we can avoid it but its an option | |
| 15:58:34 | sean-k-mooney | i think making this feature depend on if config drive is used or not is worse then only supproting it for some virt dirvrs | |
| 15:58:37 | gibi | stephenfin: there is two FUPs at the top. I will move them to be on the mergeable part of the series | |
| 15:58:46 | sean-k-mooney | sicne the config drive can change based on what host you land on | |
| 15:59:00 | sean-k-mooney | there is no way to knwo if it will work or not | |
| 15:59:33 | dansmith | sean-k-mooney: hmm, the mandatory config drive thing is a compute-node config right? | |
| 15:59:42 | sean-k-mooney | dansmith: yes | |
| 15:59:45 | dansmith | yeah, that sucks | |
| 15:59:58 | dansmith | so instances at the edge which use configdrive become un-updatable | |
| 16:00:10 | sean-k-mooney | yep | |
| 16:00:11 | gibi | user could use the capability trait to land on a host that supports regen | |
| 16:00:12 | dansmith | so another thing, | |
| 16:00:18 | sean-k-mooney | or ones in isolated networks | |
| 16:00:26 | dansmith | we really should support reboot with user_data for any instance, | |
| 16:00:36 | dansmith | because that's the consistent way to get it updated, regardless of what type it is | |
| 16:00:45 | dansmith | that way you get updated, you get rebooted, etc | |
| 16:01:04 | dansmith | gibi: the user doesn't know this is a problem until it's too late is the point | |
| 16:01:12 | sean-k-mooney | dansmith: that more or less what we were tryign to do and why i was insitieng on config drive support | |
| 16:01:18 | sean-k-mooney | i guess the disconenct is | |
| 16:01:22 | dansmith | gibi: they don't choose a host based on whether they want to update the user_data a month from now :) | |
| 16:01:27 | sean-k-mooney | i tought it was ok to start with the libvirt dirver | |
| 16:01:37 | sean-k-mooney | since we coudl add other drivers later | |
| 16:02:13 | dansmith | sean-k-mooney: I don't think I said I'm opposed to that, what I'm opposed to is just never having that implemented for the other drivers | |
| 16:02:20 | sean-k-mooney | gibi: dansmith also i fixed a bug where migration could cause you to gain/loose a config drive by making it sticky | |
| 16:02:50 | sean-k-mooney | so if the vm is first created with a config driver because of the option i made that sitcky | |
| 16:02:52 | dansmith | and given how intertwined this is with the virt driver, and that it was already done wrong once, I would hate to implement this one way and find out that vmware or hyperv or ironic can't honor the request to rebuild during a reboot at the right time | |
| 16:03:28 | sean-k-mooney | dansmith: well that is currently blocked by the trait today | |
| 16:03:29 | dansmith | sean-k-mooney: that also means that if you got a configdrive when you created the instance, it will never have update-able user_data right? | |
| 16:03:39 | sean-k-mooney | but if we tried to do it quick that could break | |
| 16:03:54 | sean-k-mooney | dansmith: yes exactly | |
| 16:04:25 | dansmith | sean-k-mooney: no I get that we can detect if it's supported today, I'm talking about if we merge this and next cycle ironic says "we literally can't regenerate the config drive" and hyperv says 'we can, but not in the middle of a reboot because of how we implement that" | |
| 16:04:29 | dansmith | then we're kinda stuck | |
| 16:05:05 | sean-k-mooney | ya fair | |
| 16:05:35 | sean-k-mooney | its internal to the driver but it might for exampel require the current reboot to be slit into stop,update config drive, start | |
| 16:05:42 | sean-k-mooney | for the hyperv senario | |
| 16:05:56 | sean-k-mooney | which we might or might be able to hide | |
| 16:06:02 | dansmith | libvirt's hard reboot is basically a recreate, but that doesn't mean the other drivers are | |
| 16:06:17 | sean-k-mooney | yep | |
| 16:06:33 | dansmith | let me also say that I'm sorry I brought up all these concerns late, but I *was* asked to review this and have tried to put my money where my mouth is on changes | |
| 16:06:38 | dansmith | but I think these are all legit concerns | |
| 16:07:17 | bauzas | yeah, there are no easy paths for solving this problem | |
| 16:07:19 | dansmith | I know gibi is plotting my murder right now, probably conspiring with jhartkopf :) | |
| 16:07:25 | sean-k-mooney | they are. we discussed it in the ptg as we had previous rejected the spec | |
| 16:07:29 | sean-k-mooney | last cycle | |
| 16:08:06 | gibi | I'm sorry that I drove jhartkopf's solution to a dead end. | |
| 16:08:08 | bauzas | well, I had concerns about the complexity it was creating for little gain | |
| 16:08:21 | bauzas | but I was opposed this was an easy win | |
| 16:09:04 | gibi | I reviewed these patches and missed obvious design errors. I will try better next time. | |
| 16:09:04 | bauzas | so, now, I'm trying to find a trade-off but I don't wanna pull the strings if I think this is risky | |
| 16:09:08 | sean-k-mooney | well little gain is not neeisaly fiar it makes nova more "cloud native" as the idea was that user data shoudl be more liek k8s config maps | |
| 16:09:35 | sean-k-mooney | alhtoguh to be fiare that woudl also imply we shoudl delete and recreate the vms | |
| 16:09:44 | dansmith | hah right | |
| 16:09:57 | dansmith | no problem updating user_data if we just shoot the instances in the head :D | |
| 16:09:58 | bauzas | sean-k-mooney: we never proposed userdata to be mutable | |
| 16:10:12 | sean-k-mooney | who is we | |
| 16:10:33 | bauzas | in a cloud, you just spin another instance if you dislike your existing userdata | |
| 16:10:34 | sean-k-mooney | this has been a long runing request for multiple cycles | |
| 16:10:44 | sean-k-mooney | not in aws | |
| 16:10:50 | sean-k-mooney | and other plathforms | |
| 16:10:52 | jhartkopf | dansmith: Honestly it's better to notice problems now than when it's too late | |
| 16:10:57 | sean-k-mooney | apprenly openstack was an outlier | |
| 16:11:15 | bauzas | https://docs.openstack.org/nova/latest/user/metadata.html#user-provided-data | |
| 16:11:44 | gibi | dansmith: I would I? I feel sorry for jhartkopf's time spent on this. And I feel bad about that I was not able to find the issues you and sean-k-mooney found | |
| 16:11:57 | gibi | *why wouldi? | |
| 16:12:04 | bauzas | glad I'm not quoted :) | |
| 16:12:04 | sean-k-mooney | bauzas: right but this concept was orginally borrowed form ec2 | |
| 16:12:46 | sean-k-mooney | bauzas: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html | |
| 16:13:50 | sean-k-mooney | https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/user-data.html#user-data-view-change | |
| 16:14:16 | bauzas | the stop requirement from EC2 may help | |
| 16:15:09 | bauzas | anyway, I was about to draft a design modification | |
| 16:15:46 | sean-k-mooney | well doing it on hard reboot was an optimisation of stop update start | |
| 16:15:52 | sean-k-mooney | but we could do htat instead | |
| 16:15:56 | sean-k-mooney | or shelve | |
| 16:16:21 | bauzas | the fact is, if we only allow the instance to be stopped, we would only allow the instance to be started once the userdata is updated | |
| 16:16:33 | sean-k-mooney | unless server update is blocking right now im not sure it helps | |
| 16:16:38 | bauzas | at least once the request ended | |
| 16:17:04 | sean-k-mooney | bauzas: well we woudl need an new rpc to the compute to rebuild teh cofnig drive | |
| 16:17:06 | dansmith | we still have to regen the config drive though | |
| 16:17:10 | dansmith | which we don't do on reboot right? | |
| 16:17:14 | dansmith | we, stop..start | |
| 16:17:32 | sean-k-mooney | right we dont regen the config drive out side of unshelve or corss cell migration | |
| 16:17:40 | dansmith | but yeah, saying you can only update user_data if the instance is stopped is probably a lot better for all the arguments | |
| 16:17:43 | bauzas | phew | |
| 16:17:43 | sean-k-mooney | and unshelve only for non rbd backed instnaces | |
| 16:17:52 | dansmith | it addresses the "so do I need to reboot?" question on non-configdrive | |
| 16:18:07 | dansmith | and eliminates the need for the rpc | |
| 16:18:23 | dansmith | it still requires regen to work, and still may or may not be easy or possible for some drivers | |
| 16:18:35 | dansmith | like ironic might have to change the provision state to re-write a disk or something | |
| 16:18:35 | gibi | will we regen for each start then? | |
| 16:19:08 | dansmith | gibi: that'd be a question yeah | |
| 16:19:21 | sean-k-mooney | im not sure we woudl remove the rpc | |
| 16:19:21 | dansmith | I really don't like the idea of a dirty flag for the other reasons | |
| 16:19:34 | dansmith | sean-k-mooney: oh you mean a new rpc for the update | |
| 16:19:39 | dansmith | yeah that could work | |
| 16:19:40 | sean-k-mooney | ya | |
| 16:19:56 | dansmith | we could just nuke the existing drive in the libvirt case, but something like ironic would have to keep track | |
| 16:19:57 | sean-k-mooney | i mean we have another option | |
| 16:20:08 | sean-k-mooney | make this entirly uer driver by adding a new instance action | |
| 16:20:18 | sean-k-mooney | for updating the config drive | |
| 16:20:33 | dansmith | oh instead of a PUT | |