Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-12
10:07:27 bauzas before we flip the default and remove the old bits, we propose you to check that you opted in before
10:07:45 noonedeadpunk aha, yes, exactly what I was unsure about. As currently we run upgrade check basically after cell_v2 discover_hosts at the very end
10:08:18 noonedeadpunk which didn't make much sense to me
10:08:43 noonedeadpunk is there any way to tell if online_data_migrations are needed at all? Like by some exit code of smth?
10:09:12 bauzas noonedeadpunk: see the upgrade docs, this explains it way better than me now https://docs.openstack.org/nova/latest/admin/upgrades.html#rolling-upgrade-process
10:10:06 bauzas noonedeadpunk: wait, I said something wrong
10:10:29 bauzas noonedeadpunk: https://docs.openstack.org/nova/latest/cli/nova-status.html#upgrade
10:11:25 bauzas you need to do DB schema expands/contracts before with the migrations
10:11:38 bauzas then, you run the status check
10:11:51 bauzas and only then you need to start the compute services
10:12:37 bauzas neither the online data migrations
10:15:11 bauzas noonedeadpunk: as a summary, say that during a Bobcat release we may add some data online migrations
10:15:33 noonedeadpunk ok, so status check before service restart but after online migrations?
10:15:39 bauzas we would provide the nova-manage command to perform such data migrations
10:16:11 bauzas accordingly, we would propose a nova-status check addition that would check whether the data is fully migrated
10:16:38 bauzas so when upgrading to C, you would run this nova-status check that would fail if you forgot to run the online data migrations command
10:16:56 bauzas same goes with Placement
10:17:26 bauzas say that by Bobcat, we start using a new Placement API call
10:17:38 bauzas we would need to continue supporting the old Placement calls
10:18:13 bauzas but if we start using that sole version of Placement, then we would need to add a nova-status check that would fail if you forgot to update Placement
10:19:36 bauzas I just hope I haven't lost you :)
10:25:06 wangrong dansmith: sean-k-mooney: Thank you for your feedback in generic-vdpa spec. Regarding some of the issues you raised, I have explained them in my reply. Could you please take a look at the submission again. Thank you for your assistance. https://review.opendev.org/c/openstack/nova-specs/+/879338
10:25:52 sean-k-mooney i saw i skimmed it breifly this morning
10:25:56 gibi sean-k-mooney: https://bugs.launchpad.net/nova/+bug/2008716 this seems to be a real bug in our migration code but you had comments before reagarding different dbs being targeted. I don't see how that is possible given the error message states the that column is missing not the table
10:26:15 noonedeadpunk bauzas: sorry, was side-pinged :(
10:26:19 sean-k-mooney wangrong: in general i think we will still need to change the design singnifcalty form what you orginaly propsosed
10:27:07 noonedeadpunk that makes total sense.
10:27:44 noonedeadpunk So yes, once I've upgraded packages on Bobcat, I should do migrations, and then check upgrade to ensure it's safe to restart now
10:35:48 noonedeadpunk ok, now will try to make some patches out of this discussion
10:37:29 wangrong sean-k-mooney: ok, could you offer more detail info about the problem of current design in the review comment
10:41:52 sean-k-mooney yes i need to also fully re read your comments
10:42:35 sean-k-mooney wangrong: basically i think we need two feature before we can implement the feature you are proposing
10:43:05 sean-k-mooney allow seting the hw_vif_model per port and allow setting hw_disk_bus per cinder volume
10:43:29 sean-k-mooney if we have both feature we can add hw_vif_modle=vdpa and hw_disk_bus=vdpa
10:43:36 sean-k-mooney to enabel generic vdpa support
10:46:23 wangrong sean-k-mooney: yes, I think so, if we plan to use hw_xxx.
11:07:12 opendevreview Oleksandr Klymenko proposed openstack/placement master: CADF audit support for Placement API https://review.opendev.org/c/openstack/placement/+/880145
11:38:40 noonedeadpunk No, wait. I'm confused again. What's written in docs, is that I should run online_data_migrations _after_ service restart, and upgrade check _before_ restart. So upgrade check should be run before migrations
11:46:40 bauzas noonedeadpunk: because the upgrade check would verify something else but the online data migrations for the same cycle
11:51:20 noonedeadpunk yeah, ok, we just got things mixed up a bit
11:51:47 noonedeadpunk we were running online_data_migrations early, before first conductor is restarted, but upgrade_check when all services are
11:51:54 noonedeadpunk So i got really confused
11:52:16 sean-k-mooney so you should be able to do db sync before an restarts
11:52:23 sean-k-mooney im not sure about online migration
11:52:59 noonedeadpunk yeah, db_syncs are done before, but were followed with online migrations...
11:53:01 sean-k-mooney part of me expect thost to happen after the db models are updated i.e. after all the contoelr services are restarted
11:53:35 noonedeadpunk What I'm not sure about, if upgrade_check is going to pass if only conductor/api is available, but no computes. I assume it should not care as long as cell mappings are in place.
11:53:38 sean-k-mooney so im not sure that will break things but it leave a window where more data could be created because fo the old db model that is in use
11:54:00 noonedeadpunk Well. That is true.
11:54:15 sean-k-mooney so i think online migration shoudl be after the restarts
11:54:16 noonedeadpunk But docs you've shown say it should be done after everything
11:54:30 sean-k-mooney ya
11:54:39 sean-k-mooney i would generally put them right at the end
11:54:43 noonedeadpunk As partially data will be migrated when requested even without migrations
11:54:44 sean-k-mooney when all services are fully upgraded
11:55:01 noonedeadpunk yup, great, thanks for help as usual!
11:55:38 sean-k-mooney so the thing is as long as you are not trying to skip several release
11:55:45 sean-k-mooney we will have the fall back code
11:55:57 sean-k-mooney so the online migrations coudl be done even in a seperate maintaince window
11:55:59 noonedeadpunk well, it seemed working nicely old way as well for years, so I was quite tentative to touch anything there :D
11:56:15 sean-k-mooney but tis better to do them as part of the normal upgrade after all the services are updated
11:56:52 sean-k-mooney where you would have a probelm is if there is a db contraction
11:57:00 sean-k-mooney alotugh i dont think that woule be part of the online migrations
11:57:14 sean-k-mooney i.e. if we removed a column
11:57:32 noonedeadpunk Yeah, that what I was thinking about
11:57:57 noonedeadpunk that if db schema changes - it should not be done post, but I assume it's done during db_sync
11:58:14 noonedeadpunk which is done really at early stage
11:58:15 sean-k-mooney ya that a good point
11:58:26 sean-k-mooney schema chagnes are seperate
11:58:39 sean-k-mooney the online migation are just data migrations
11:59:07 noonedeadpunk ok, then will see how CI will fail dramatically with my changes :D
12:04:49 bauzas noonedeadpunk: as it was written in the docs I passed, db sync are only changing the schemas
12:05:42 bauzas noonedeadpunk: then either we change the data internally (either by a lazy call, or when restarting the service), or we ask operators to run online_data_migration
12:06:26 bauzas but online_data_migration can run during all the cycle, until you want to upgrade to a new release
12:07:34 bauzas then, once you want to upgrade, you can before run nova-status upgrade check for verifying that you're done with the online data migrations
12:08:38 noonedeadpunk with N-1 online data migrations, right?
12:10:07 noonedeadpunk so running upgrade check on N will verify that N-1 migrations are done
12:10:22 noonedeadpunk or well, N-2 with slurp, but that's different topic :D
12:13:40 bauzas yup
12:14:08 noonedeadpunk ok, awesome, thanks for your time and patience :)
13:35:58 dansmith noonedeadpunk: yeah, can't run online migrations until you're past the point where conductors, api, scheduler are upgraded (or stopped before upgrade
13:36:12 dansmith noonedeadpunk: the idea is that online migrations should be run *after* everything is done and back up
13:36:34 dansmith services will migrate what they need on-demand, and the online migrations are just there to push things that don't get migrated on demand
13:38:53 noonedeadpunk yup, thanks for confirming that!
13:41:10 noonedeadpunk will go bug cinder folks with the same question then :-)
13:45:37 dansmith ack, I didn't mean to repeat, just wasn't sure there was a clear summary of that convo, but sounds like you got it :)
13:46:56 noonedeadpunk yeah, I've pushed change as a result - I should have posted it https://review.opendev.org/c/openstack/openstack-ansible-os_nova/+/880147
13:57:41 dansmith noonedeadpunk: I think running the status check after online migrations also makes sense as there are cases where it will tell you that there are still pending migrations to do if they're not complete
13:57:51 dansmith but regardless, the meat of that change sounds right yeah
14:52:29 bauzas dansmith: thanks for having explained it better than me :)
14:52:53 bauzas your summary is far easier :)
15:12:49 dansmith bauzas: when you're done with the current call you're on I want to chat about some resource tracker stuff
15:54:38 noonedeadpunk between not that long ago (on PTG?) we've discussed enable_new_services option. And it should be applied not for computes at the end of the day, but somewhere on conductor/scheduler/api - not sure where exactly as I have them combined at same place (with same nova.conf)
15:55:34 noonedeadpunk so while it's defined in nova.conf.compute - it's not compute option
15:55:51 noonedeadpunk (https://opendev.org/openstack/nova/src/branch/master/nova/conf/compute.py#L1456-L1474)
15:57:08 dansmith conductor
15:57:11 dansmith but we could also fix that
16:00:09 bauzas dansmith: I'm here
16:00:29 dansmith bauzas: so... the RT has this notion of "disabled compute nodes"
16:00:51 dansmith which seems to cover things actually disabled (via ironic) as well as "no such compute node on this host"

Earlier   Later