| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-12 | |||
| 09:45:29 | sean-k-mooney | it discovers by connecting to all the cell dbs | |
| 09:45:34 | noonedeadpunk | ok-ok, yes, I recalled :) | |
| 09:46:38 | sean-k-mooney | https://github.com/openstack/nova/blob/master/nova/cmd/manage.py#L1043 | |
| 09:46:40 | sean-k-mooney | calls https://github.com/openstack/nova/blob/ae42400b7663bc58d5562de99e976c95131b77a9/nova/objects/host_mapping.py#L247-L285 | |
| 09:47:02 | noonedeadpunk | and in order for compute to be discovered by discover_hosts it must already be in service list? | |
| 09:47:12 | opendevreview | Merged openstack/placement master: Db: Drop redundant indexes for columns already having unique constraint https://review.opendev.org/c/openstack/placement/+/856770 | |
| 09:47:14 | sean-k-mooney | correct | |
| 09:47:26 | sean-k-mooney | so the compute agent connects to the conducotr on startup | |
| 09:47:47 | sean-k-mooney | and then it ask the conductor to register a service in the db | |
| 09:48:01 | sean-k-mooney | "the cell db" | |
| 09:48:23 | sean-k-mooney | the conducotor may or may not have access to the api db | |
| 09:48:32 | sean-k-mooney | so we cant assume it can add the cell mapping | |
| 09:50:11 | noonedeadpunk | yeah, ok, thanks for the explanation | |
| 09:50:33 | noonedeadpunk | As I was mixing up discovery flow | |
| 09:50:45 | bauzas | noonedeadpunk: maybe this would help you :p https://docs.openstack.org/nova/latest/admin/cells.html#faqs | |
| 09:51:31 | bauzas | but yeah, discover_hosts is idempotent so you can call it anytime you want | |
| 09:52:14 | bauzas | this is explained at the end of https://docs.openstack.org/nova/latest/admin/cells.html#configuring-a-new-deployment | |
| 09:52:38 | noonedeadpunk | well, I was unsure about logic, as I could not fully understand the reason why we're waiting for compute to be present in service list when nova_discover_hosts_in_cells_interval=-1 but run cell_v2 discover_hosts only after that | |
| 09:53:05 | bauzas | noonedeadpunk: we have two different databases now as a reminder | |
| 09:53:17 | bauzas | so we need to make sure we synchronize them | |
| 09:53:24 | noonedeadpunk | As for some reason I thought that discover_hosts also will add to service list, but now I know it;s not how it works) | |
| 09:54:30 | sean-k-mooney | bauzas: technially we have 3. we alway have api + cell 0 and cell one in any real deployment :) | |
| 09:57:21 | bauzas | sean-k-mooney: we have two different table schemas, but yeah :) | |
| 09:57:36 | bauzas | cell0 is logically identically to any cell | |
| 09:57:40 | bauzas | identical* | |
| 09:57:54 | bauzas | only the usage of this cell is different | |
| 09:58:05 | bauzas | but we're bikeshedding :) | |
| 09:59:58 | noonedeadpunk | Another thing I am unsure about - db online_data_migrations and nova-status upgrade check and when they can/should be run. | |
| 10:00:12 | noonedeadpunk | As far as I got - db online_data_migrations should be done only during upgrades | |
| 10:00:34 | noonedeadpunk | But is it recommended to run `nova-status upgrade check` before that? | |
| 10:01:26 | noonedeadpunk | As it feels like upgrade check should be run once all computes are upgraded, while online_data_migrations requires only conductor to be done? | |
| 10:05:16 | bauzas | noonedeadpunk: this is correct | |
| 10:05:44 | bauzas | you should run nova-status upgrade check *before* the db upgrade as a pre-flight check | |
| 10:06:24 | bauzas | nova-status is a controller script, checking a few APIs and DBs | |
| 10:06:33 | bauzas | to see whether you're safe to upgrade | |
| 10:07:02 | bauzas | like, say we introduced a thing but wasn't the default | |
| 10:07:27 | bauzas | before we flip the default and remove the old bits, we propose you to check that you opted in before | |
| 10:07:45 | noonedeadpunk | aha, yes, exactly what I was unsure about. As currently we run upgrade check basically after cell_v2 discover_hosts at the very end | |
| 10:08:18 | noonedeadpunk | which didn't make much sense to me | |
| 10:08:43 | noonedeadpunk | is there any way to tell if online_data_migrations are needed at all? Like by some exit code of smth? | |
| 10:09:12 | bauzas | noonedeadpunk: see the upgrade docs, this explains it way better than me now https://docs.openstack.org/nova/latest/admin/upgrades.html#rolling-upgrade-process | |
| 10:10:06 | bauzas | noonedeadpunk: wait, I said something wrong | |
| 10:10:29 | bauzas | noonedeadpunk: https://docs.openstack.org/nova/latest/cli/nova-status.html#upgrade | |
| 10:11:25 | bauzas | you need to do DB schema expands/contracts before with the migrations | |
| 10:11:38 | bauzas | then, you run the status check | |
| 10:11:51 | bauzas | and only then you need to start the compute services | |
| 10:12:37 | bauzas | neither the online data migrations | |
| 10:15:11 | bauzas | noonedeadpunk: as a summary, say that during a Bobcat release we may add some data online migrations | |
| 10:15:33 | noonedeadpunk | ok, so status check before service restart but after online migrations? | |
| 10:15:39 | bauzas | we would provide the nova-manage command to perform such data migrations | |
| 10:16:11 | bauzas | accordingly, we would propose a nova-status check addition that would check whether the data is fully migrated | |
| 10:16:38 | bauzas | so when upgrading to C, you would run this nova-status check that would fail if you forgot to run the online data migrations command | |
| 10:16:56 | bauzas | same goes with Placement | |
| 10:17:26 | bauzas | say that by Bobcat, we start using a new Placement API call | |
| 10:17:38 | bauzas | we would need to continue supporting the old Placement calls | |
| 10:18:13 | bauzas | but if we start using that sole version of Placement, then we would need to add a nova-status check that would fail if you forgot to update Placement | |
| 10:19:36 | bauzas | I just hope I haven't lost you :) | |
| 10:25:06 | wangrong | dansmith: sean-k-mooney: Thank you for your feedback in generic-vdpa spec. Regarding some of the issues you raised, I have explained them in my reply. Could you please take a look at the submission again. Thank you for your assistance. https://review.opendev.org/c/openstack/nova-specs/+/879338 | |
| 10:25:52 | sean-k-mooney | i saw i skimmed it breifly this morning | |
| 10:25:56 | gibi | sean-k-mooney: https://bugs.launchpad.net/nova/+bug/2008716 this seems to be a real bug in our migration code but you had comments before reagarding different dbs being targeted. I don't see how that is possible given the error message states the that column is missing not the table | |
| 10:26:15 | noonedeadpunk | bauzas: sorry, was side-pinged :( | |
| 10:26:19 | sean-k-mooney | wangrong: in general i think we will still need to change the design singnifcalty form what you orginaly propsosed | |
| 10:27:07 | noonedeadpunk | that makes total sense. | |
| 10:27:44 | noonedeadpunk | So yes, once I've upgraded packages on Bobcat, I should do migrations, and then check upgrade to ensure it's safe to restart now | |
| 10:35:48 | noonedeadpunk | ok, now will try to make some patches out of this discussion | |
| 10:37:29 | wangrong | sean-k-mooney: ok, could you offer more detail info about the problem of current design in the review comment | |
| 10:41:52 | sean-k-mooney | yes i need to also fully re read your comments | |
| 10:42:35 | sean-k-mooney | wangrong: basically i think we need two feature before we can implement the feature you are proposing | |
| 10:43:05 | sean-k-mooney | allow seting the hw_vif_model per port and allow setting hw_disk_bus per cinder volume | |
| 10:43:29 | sean-k-mooney | if we have both feature we can add hw_vif_modle=vdpa and hw_disk_bus=vdpa | |
| 10:43:36 | sean-k-mooney | to enabel generic vdpa support | |
| 10:46:23 | wangrong | sean-k-mooney: yes, I think so, if we plan to use hw_xxx. | |
| 11:07:12 | opendevreview | Oleksandr Klymenko proposed openstack/placement master: CADF audit support for Placement API https://review.opendev.org/c/openstack/placement/+/880145 | |
| 11:38:40 | noonedeadpunk | No, wait. I'm confused again. What's written in docs, is that I should run online_data_migrations _after_ service restart, and upgrade check _before_ restart. So upgrade check should be run before migrations | |
| 11:46:40 | bauzas | noonedeadpunk: because the upgrade check would verify something else but the online data migrations for the same cycle | |
| 11:51:20 | noonedeadpunk | yeah, ok, we just got things mixed up a bit | |
| 11:51:47 | noonedeadpunk | we were running online_data_migrations early, before first conductor is restarted, but upgrade_check when all services are | |
| 11:51:54 | noonedeadpunk | So i got really confused | |
| 11:52:16 | sean-k-mooney | so you should be able to do db sync before an restarts | |
| 11:52:23 | sean-k-mooney | im not sure about online migration | |
| 11:52:59 | noonedeadpunk | yeah, db_syncs are done before, but were followed with online migrations... | |
| 11:53:01 | sean-k-mooney | part of me expect thost to happen after the db models are updated i.e. after all the contoelr services are restarted | |
| 11:53:35 | noonedeadpunk | What I'm not sure about, if upgrade_check is going to pass if only conductor/api is available, but no computes. I assume it should not care as long as cell mappings are in place. | |
| 11:53:38 | sean-k-mooney | so im not sure that will break things but it leave a window where more data could be created because fo the old db model that is in use | |
| 11:54:00 | noonedeadpunk | Well. That is true. | |
| 11:54:15 | sean-k-mooney | so i think online migration shoudl be after the restarts | |
| 11:54:16 | noonedeadpunk | But docs you've shown say it should be done after everything | |
| 11:54:30 | sean-k-mooney | ya | |
| 11:54:39 | sean-k-mooney | i would generally put them right at the end | |
| 11:54:43 | noonedeadpunk | As partially data will be migrated when requested even without migrations | |
| 11:54:44 | sean-k-mooney | when all services are fully upgraded | |
| 11:55:01 | noonedeadpunk | yup, great, thanks for help as usual! | |
| 11:55:38 | sean-k-mooney | so the thing is as long as you are not trying to skip several release | |
| 11:55:45 | sean-k-mooney | we will have the fall back code | |
| 11:55:57 | sean-k-mooney | so the online migrations coudl be done even in a seperate maintaince window | |
| 11:55:59 | noonedeadpunk | well, it seemed working nicely old way as well for years, so I was quite tentative to touch anything there :D | |
| 11:56:15 | sean-k-mooney | but tis better to do them as part of the normal upgrade after all the services are updated | |
| 11:56:52 | sean-k-mooney | where you would have a probelm is if there is a db contraction | |
| 11:57:00 | sean-k-mooney | alotugh i dont think that woule be part of the online migrations | |
| 11:57:14 | sean-k-mooney | i.e. if we removed a column | |