Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-01
16:23:45 sean-k-mooney right it woudl be in the instance event log
16:23:59 bauzas and until the userdata says "I'm updated", don't expect to get the new things if you start
16:24:01 gibi we can also make the start to fail if it needed to regen but failed to
16:24:02 sean-k-mooney if its async or in the respocne if blocking
16:24:22 sean-k-mooney gibi: that woudl requrie trackign the dirty state
16:24:34 sean-k-mooney and also some peopel might not care
16:25:16 gibi I don't see how somebody explicitly updated the user_data and then don't case if that update is ending up in the VM being sarted
16:25:19 gibi started
16:25:48 gibi I think the dirty flag was bad because of shadow rpc
16:25:58 sean-k-mooney im thinking about something else
16:26:14 sean-k-mooney as a user i might not know if the user data is in config drive or not
16:26:26 sean-k-mooney but i said start
16:26:38 sean-k-mooney and i might not want tha to break in this case
16:27:19 gibi I see
16:27:26 sean-k-mooney the dirty flag was to give you the behaior that it would only boot if it succeed to update
16:27:42 bauzas honestly, if I'm an user, I don't care about the internals
16:27:44 sean-k-mooney if it failed at least in my version teh vm would have gone to error
16:27:57 bauzas what I want to know is when my update succeeds and when I can use it
16:28:02 dansmith sorry, I'm getting pulled away,
16:28:06 dansmith but if it was an instance action,
16:28:18 dansmith then the rebuild either works or doesn't, I can see what happened,
16:28:29 dansmith and then I either get the old version on fail, or the new version on succeed
16:28:46 dansmith if we just set a dirty flag and then fail to regen on start, then we've killed the instance with no way to restart
16:28:48 dansmith which sucks
16:28:48 bauzas yup, I get feedback from an user perspective
16:28:56 bauzas yup, agreed
16:29:13 sean-k-mooney jhartkopf: so based on the above
16:29:15 dansmith we probably need some task_state or something to keep the user from starting it until the regen completes though
16:29:22 sean-k-mooney jhartkopf: and the fact we are all runing out of steam
16:29:28 bauzas mikal would say "spec'y spec'y"
16:29:31 dansmith or just use instance lock on the compute node to hold off on any other requests.. I guess that's enough
16:29:32 sean-k-mooney i think we need to disucss this again next cycle
16:29:53 bauzas the next PTG will be virtual
16:30:05 bauzas anyone can join, budged saved
16:30:24 sean-k-mooney ya but we can reopne the sepc and discuss it again before htat
16:30:31 bauzas oh yeah
16:30:58 sean-k-mooney i see to ways we coudl go with the instance action by the way
16:31:05 sean-k-mooney one is one for regeneratin cofnig drive
16:31:11 sean-k-mooney the other is more targed for this feature
16:31:21 sean-k-mooney an instance action to update user_data
16:31:30 sean-k-mooney that either succeed or does not
16:31:49 sean-k-mooney and have that be the only wya to update it
16:31:57 sean-k-mooney so no updating it via server put
16:32:25 dansmith yeah, the user should not see "regenerate config drive"
16:32:41 dansmith so it should be "update my user_data" at the external surface for sure
16:32:52 sean-k-mooney right
16:32:55 sean-k-mooney so if we do that
16:33:16 sean-k-mooney that dedicated instance action can "do the right thing" regardess of how the instance is booted
16:33:31 dansmith right
16:33:42 sean-k-mooney and returna 409 conflict if its not supprot by the driver or whatever
16:33:55 jhartkopf sean-k-mooney: Oh that'd be bad news. Was hoping to get at least the stuff without config drives merged this cycle.
16:34:12 sean-k-mooney but it can be atomic
16:34:40 JayF o/, I have a handful of ironic driver patches I'm still trying to get backported thru. They are all clean backports and passing tests.
16:34:42 JayF https://review.opendev.org/c/openstack/nova/+/854257 (no live migration tests on ironic driver changes -> yoga) https://review.opendev.org/c/openstack/nova/+/821351 (ignore plug_viles on ironic driver -> ussuri) https://review.opendev.org/c/openstack/nova/+/853546 (minimize window for a resource provider to be lost -> train)
16:34:42 bauzas sean-k-mooney: dansmith: yeah, the user-surfaced API is the userdata, not the configdrive
16:35:01 bauzas so the instance action has to be related to it
16:35:39 sean-k-mooney JayF: did we land the skip of the migration job yet
16:35:47 JayF sean-k-mooney: that's the yoga backport above
16:35:52 JayF sean-k-mooney: the version -> master is in
16:35:55 sean-k-mooney an cool
16:36:03 sean-k-mooney i was onder if the master on had landed
16:36:05 bauzas jhartkopf: the problem is that we can't introduce this userdata update by splitting without configdrive as we just agreed on a brand new API
16:36:24 JayF oh, I've already checked all these -- previous branches have merged, these should be good to go. e.g. I'd approve these if they were ironic stable changes
16:36:32 bauzas for consistency reasons, we have to hide the configdrive details from the user request
16:36:43 sean-k-mooney bauzas: that and my +2 on the spec was contingent on it not depending on if you use configdrive or not
16:36:46 JayF and elod helped me clean up the only piece around nova stable policy that's more strict than Ironic :D
16:37:33 bauzas sean-k-mooney: that kind of design problems happening late in the cycle are common
16:38:14 bauzas jhartkopf: this isn't at all about your implementation, this is rather about the fact we discovered some design problems we weren't considering at the spec review time
16:38:16 gibi we should be better at doing this kind of design discussion before the FF day
16:38:37 gibi but I don't know how
16:38:43 bauzas honestly, this is the only blueprint that fails this for this cycle AFAICT
16:38:51 bauzas gibi: we had precedents
16:39:07 bauzas some spec approved and a design question raising late
16:39:08 jhartkopf bauzas: I understand that, but it's really unfortunate though
16:39:52 bauzas jhartkopf: I don't disagree with you and I understand how much this can be frustrating
16:40:01 artom bauzas, wait, so ephemral disk, PCI, and manial will all make it?
16:40:12 artom *Manilla
16:40:20 gibi bauzas: I don't think we can detect all the design issues at spec time. But still I think we should be better than doing it at the day of FF
16:40:21 sean-k-mooney not sure about manilla
16:40:22 bauzas jhartkopf: I just hope you can understand that our API resources are important to us and we really care about the UX
16:40:39 sean-k-mooney ephmeral also likely not
16:40:51 bauzas artom: no, I was mentioning a design issue being raised
16:40:53 artom Ah
16:41:38 bauzas ephemerals and PCI are on their way but they're big and need a lot of back and forths
16:41:48 gibi the only easy answer for me know is that I should have asked dansmith earlier to review the series
16:41:53 bauzas but they hadn't required a redesign
16:42:11 bauzas gibi: well, if you read the spec, I had concerns about the restart
16:42:37 dansmith gibi: this is why I'm against yearly releases.. because this stuff, very unfortunately, only comes to a head when there's a deadline...
16:42:40 gibi so the I should have asked you to review the series too
16:42:41 bauzas we just didn't went further thinking about configdrive problems
16:42:59 gibi dansmith: I totally agree about having more meaningful deadline would help
16:43:07 sean-k-mooney dansmith: because thats the forcing function that get a lot of use to review a tthe same time
16:43:14 sean-k-mooney dansmith: in the absence of an inperson PTG
16:43:19 dansmith and I only reviewed it because it was in front of something else that I know we needed to land, which was up against the deadline
16:43:27 bauzas and I was on PTO for 3.5 weeks
16:43:30 bauzas so blame me too
16:43:33 dansmith WE ALL SUCK
16:43:36 dansmith there ^ :P
16:43:40 sean-k-mooney :)
16:43:59 gibi I don't want to blame, I want to find something I or we can do different next time

Earlier   Later