| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-01 | |||
| 15:38:33 | bauzas | and then I don't understand, my f*** script doesn't run as expected when I restart the guest | |
| 15:38:44 | bauzas | shall I open a ticket ? | |
| 15:39:08 | dansmith | oh it's in hyperv, just in vmops.py | |
| 15:39:24 | bauzas | oh, simple, your cloud was running some driver on that host that was having trouble with updating the configdrive | |
| 15:39:35 | bauzas | either because it was a legacy driver | |
| 15:39:44 | bauzas | or because something (like a perm error) expected | |
| 15:40:10 | bauzas | honestly, this spec was just about touching an internal Nova DB record | |
| 15:40:21 | bauzas | now, we're pulling way more than that | |
| 15:40:47 | dansmith | yeah, it's becoming more about configdrive than anything else | |
| 15:41:27 | sean-k-mooney | bauzas: in two converstaion but form my persepcitve i dont think we really have discoverd anythign bar the rpc change. i reveiw the spec with the understandin it was scoped to libvirt | |
| 15:42:44 | bauzas | yet again, can't we just assume to update instance's metadata only if not configdrive-driven, and leave people who care about configdrives deal with the complexity one day or another? | |
| 15:42:45 | dansmith | I dunno, I can see the argument for this being scoped to just one driver, but for something this fundamental it seems pretty unfortunate to say it's only for libvirt | |
| 15:43:10 | bauzas | I'm pretty sure we wouldn't have this back-to-back | |
| 15:44:31 | dansmith | so we reject the api call if configdrive=True always? | |
| 15:44:57 | bauzas | that's my question | |
| 15:45:10 | dansmith | doesn't that get closer to your concern on the spec that people think they have updated their instance but didn't reboot it and don't realize that cloud-init doesn't update their stuff? | |
| 15:45:29 | dansmith | I mean, that's just a misunderstanding of cloud-init, but I thought that was your concern | |
| 15:45:36 | dansmith | anyway, I dunno | |
| 15:45:37 | bauzas | if the update failed, they know the instance userdata wasn't updated | |
| 15:46:53 | bauzas | they would know this was because configdriver if they asked by boot params | |
| 15:47:15 | bauzas | but they wouldn't know if the clould defaulted to force configdrives or if the image was asking for it | |
| 15:47:22 | dansmith | right, it's not always their choice | |
| 15:47:45 | bauzas | anyway, I'm just offering a trade-off because I'm not really happy with the current state of the series | |
| 15:48:22 | bauzas | also, is the owner of the patch around ? | |
| 15:48:29 | bauzas | or are we offering our help for free ? | |
| 15:48:47 | dansmith | who is asking for this btw? | |
| 15:48:59 | dansmith | aside from the author | |
| 15:49:02 | bauzas | a private clould company | |
| 15:49:06 | bauzas | Inovex | |
| 15:49:07 | jhartkopf | I'm here | |
| 15:49:17 | bauzas | cool, glad to hear you jhartkopf | |
| 15:49:24 | bauzas | jhartkopf: lemme summarize | |
| 15:49:25 | dansmith | jhartkopf: do you care about the configdrive case? | |
| 15:49:37 | bauzas | we're having a problem with https://review.opendev.org/c/openstack/nova/+/816157/16 | |
| 15:49:45 | bauzas | and we don't know how to move forward | |
| 15:50:02 | jhartkopf | yep I quickly read your discussion | |
| 15:50:03 | bauzas | there is a proposed solution, but this is premature work and we're very late in the cycle | |
| 15:50:46 | bauzas | we're also creating a dependency to the virt drivers and unfortunately, Hyper-V, VMWare and Ironic could be impacted if we merge | |
| 15:51:43 | bauzas | lastly, I have concnerns about ensuring that the userdata API query is tied to the successfulness of the configdrive regen op | |
| 15:51:58 | bauzas | which would require the API call to be synchronous on a configdrive regen | |
| 15:52:21 | bauzas | so, my question is | |
| 15:52:31 | bauzas | do you really care of configdrives? | |
| 15:53:08 | dansmith | bauzas: we can't make the api call synchronous for config drive regen | |
| 15:53:23 | dansmith | that could take a long time, be dependent on the IO load on the compute, etc | |
| 15:53:25 | sean-k-mooney | i dont think that was an option | |
| 15:53:27 | bauzas | sean-k-mooney: I think it would be fair to say that this feature wouldn't be available to the other drivers while apparently they already support configdrive regen | |
| 15:53:40 | bauzas | eek, not regen, but creation | |
| 15:53:46 | sean-k-mooney | bauzas: yes that is correct | |
| 15:53:58 | bauzas | dansmith: right, I was pointing out loudly the problem | |
| 15:54:15 | dansmith | ah okay | |
| 15:54:16 | bauzas | this is a design problem | |
| 15:54:27 | bauzas | configdrives take time to regenerate | |
| 15:54:38 | bauzas | on the other hand, users have an API for updating userdata | |
| 15:55:02 | bauzas | they just assume they just have to restart their guests in order to have the latest update | |
| 15:55:21 | jhartkopf | bauzas: Initially we did not even want to touch config drives at all. We'd only like to update user data in the metadata service really. | |
| 15:55:32 | bauzas | I know | |
| 15:55:32 | sean-k-mooney | right but the api prevents you updating the user data if the instance has a conf driver outside of hard_reboot | |
| 15:55:34 | stephenfin | gibi: Sorry for the delay. I'd run through most of those PCI patches yesterday but didn't leave reviews until I got to the end of the series. Done now (y) | |
| 15:55:52 | gibi | stephenfin: much appreciated. I will look | |
| 15:55:56 | stephenfin | I went as far as sean-k-mooney did. Could go further but I see -1's from you | |
| 15:56:32 | gibi | stephenfin: yes, I left them -1 there as there is a bug in the scheduling part | |
| 15:56:45 | gibi | stephenfin: I think it is totally OK to merge it up until https://review.opendev.org/c/openstack/nova/+/850468/20 | |
| 15:57:11 | jhartkopf | What's up with the plan you guys proposed yesterday to split the patch and only allow updating instances without config drive for now? | |
| 15:57:25 | gibi | stephenfin: the rest slips to AA | |
| 15:57:38 | stephenfin | sounds reasonable to me | |
| 15:58:03 | sean-k-mooney | jhartkopf: i would prefer not to do that if we can avoid it but its an option | |
| 15:58:34 | sean-k-mooney | i think making this feature depend on if config drive is used or not is worse then only supproting it for some virt dirvrs | |
| 15:58:37 | gibi | stephenfin: there is two FUPs at the top. I will move them to be on the mergeable part of the series | |
| 15:58:46 | sean-k-mooney | sicne the config drive can change based on what host you land on | |
| 15:59:00 | sean-k-mooney | there is no way to knwo if it will work or not | |
| 15:59:33 | dansmith | sean-k-mooney: hmm, the mandatory config drive thing is a compute-node config right? | |
| 15:59:42 | sean-k-mooney | dansmith: yes | |
| 15:59:45 | dansmith | yeah, that sucks | |
| 15:59:58 | dansmith | so instances at the edge which use configdrive become un-updatable | |
| 16:00:10 | sean-k-mooney | yep | |
| 16:00:11 | gibi | user could use the capability trait to land on a host that supports regen | |
| 16:00:12 | dansmith | so another thing, | |
| 16:00:18 | sean-k-mooney | or ones in isolated networks | |
| 16:00:26 | dansmith | we really should support reboot with user_data for any instance, | |
| 16:00:36 | dansmith | because that's the consistent way to get it updated, regardless of what type it is | |
| 16:00:45 | dansmith | that way you get updated, you get rebooted, etc | |
| 16:01:04 | dansmith | gibi: the user doesn't know this is a problem until it's too late is the point | |
| 16:01:12 | sean-k-mooney | dansmith: that more or less what we were tryign to do and why i was insitieng on config drive support | |
| 16:01:18 | sean-k-mooney | i guess the disconenct is | |
| 16:01:22 | dansmith | gibi: they don't choose a host based on whether they want to update the user_data a month from now :) | |
| 16:01:27 | sean-k-mooney | i tought it was ok to start with the libvirt dirver | |
| 16:01:37 | sean-k-mooney | since we coudl add other drivers later | |
| 16:02:13 | dansmith | sean-k-mooney: I don't think I said I'm opposed to that, what I'm opposed to is just never having that implemented for the other drivers | |
| 16:02:20 | sean-k-mooney | gibi: dansmith also i fixed a bug where migration could cause you to gain/loose a config drive by making it sticky | |
| 16:02:50 | sean-k-mooney | so if the vm is first created with a config driver because of the option i made that sitcky | |
| 16:02:52 | dansmith | and given how intertwined this is with the virt driver, and that it was already done wrong once, I would hate to implement this one way and find out that vmware or hyperv or ironic can't honor the request to rebuild during a reboot at the right time | |
| 16:03:28 | sean-k-mooney | dansmith: well that is currently blocked by the trait today | |
| 16:03:29 | dansmith | sean-k-mooney: that also means that if you got a configdrive when you created the instance, it will never have update-able user_data right? | |
| 16:03:39 | sean-k-mooney | but if we tried to do it quick that could break | |
| 16:03:54 | sean-k-mooney | dansmith: yes exactly | |
| 16:04:25 | dansmith | sean-k-mooney: no I get that we can detect if it's supported today, I'm talking about if we merge this and next cycle ironic says "we literally can't regenerate the config drive" and hyperv says 'we can, but not in the middle of a reboot because of how we implement that" | |
| 16:04:29 | dansmith | then we're kinda stuck | |
| 16:05:05 | sean-k-mooney | ya fair | |
| 16:05:35 | sean-k-mooney | its internal to the driver but it might for exampel require the current reboot to be slit into stop,update config drive, start | |
| 16:05:42 | sean-k-mooney | for the hyperv senario | |