Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-12
09:34:52 sean-k-mooney that should be the first set of patches in the serise so that can be merged to allow scaphanda work to start in parallel
09:35:08 sean-k-mooney while we wait for manialla to add the lock api
09:36:32 manuvakery1 Hi, I have a 2 node cluster and I am running a rally scenario against it to create and delete servers. When i set the server count to 2 I see the response time as follows
09:36:32 manuvakery1 nova.boot_servers : 11.941 sec (Avg)
09:36:32 manuvakery1 nova.delete_servers: 14.675 sec (Avg)
09:36:32 manuvakery1 make the server count to 5 added a huge difference in delete server
09:36:32 manuvakery1 nova.boot_servers: 18.803 sec (Avg)
09:36:33 manuvakery1 nova.delete_servers: 42.988 sec (Avg)
09:36:33 manuvakery1 is this expected?
09:37:14 manuvakery1 the compute nodes are not under any load
09:37:49 Uggla sean-k-mooney, hum, not really.
09:40:45 sean-k-mooney ok im not sure when we disucssed that
09:41:05 kashyap bauzas: Do we have a conclusion on this? -- https://review.opendev.org/c/openstack/nova/+/879021
09:41:28 sean-k-mooney but if we want to merge any code before the manial api is implmented the libvirt dirver chage for the xml generation is the lowest risk
09:42:09 bauzas sean-k-mooney: I haven't went more than the 4th change, but the Manila API should be the last change in the series
09:42:10 noonedeadpunk hey folks! We're reviewing our nova role, as we haven't touched in for a while, and now a bit /o\ about logic we have there. so would be glad if you can share your prespective on some things :)
09:42:34 sean-k-mooney bauzas: yes the api should be last and the xml change sin the drive rhsould be first
09:42:51 sean-k-mooney driver then db then rpc finaly api
09:42:55 bauzas sean-k-mooney: for the libvirt driver, meh to me, since I think we should be quite okay for the 4 first patches (DB and objects)
09:43:14 bauzas noonedeadpunk: shoot
09:43:16 noonedeadpunk So first question - when `discover_hosts_in_cells_interval` is set to -1, I assume you need to run `nova-manage cell_v2 discover_hosts` before compute will appear in `openstack compute service list`?
09:43:42 noonedeadpunk As I can recall that compute itself does report back on startup as well
09:43:47 sean-k-mooney noonedeadpunk: no they will show up in the list
09:43:54 sean-k-mooney but they wont be in the cell mappings
09:43:54 noonedeadpunk and this option is for scheduler
09:43:58 bauzas noonedeadpunk: no, but they won't be scheduled
09:44:09 bauzas since they won't have a cell mapping
09:44:16 noonedeadpunk aha, gotcha
09:44:16 bauzas haha
09:44:47 sean-k-mooney discover_hosts tell use which rabbit to use to talk to the compute and what cell db its in
09:44:52 noonedeadpunk ok, then we have this set correctly :D
09:45:14 noonedeadpunk but it discovers... through API?
09:45:21 sean-k-mooney no
09:45:21 noonedeadpunk not through conductor then?
09:45:28 noonedeadpunk ah, right
09:45:29 sean-k-mooney it discovers by connecting to all the cell dbs
09:45:34 noonedeadpunk ok-ok, yes, I recalled :)
09:46:38 sean-k-mooney https://github.com/openstack/nova/blob/master/nova/cmd/manage.py#L1043
09:46:40 sean-k-mooney calls https://github.com/openstack/nova/blob/ae42400b7663bc58d5562de99e976c95131b77a9/nova/objects/host_mapping.py#L247-L285
09:47:02 noonedeadpunk and in order for compute to be discovered by discover_hosts it must already be in service list?
09:47:12 opendevreview Merged openstack/placement master: Db: Drop redundant indexes for columns already having unique constraint https://review.opendev.org/c/openstack/placement/+/856770
09:47:14 sean-k-mooney correct
09:47:26 sean-k-mooney so the compute agent connects to the conducotr on startup
09:47:47 sean-k-mooney and then it ask the conductor to register a service in the db
09:48:01 sean-k-mooney "the cell db"
09:48:23 sean-k-mooney the conducotor may or may not have access to the api db
09:48:32 sean-k-mooney so we cant assume it can add the cell mapping
09:50:11 noonedeadpunk yeah, ok, thanks for the explanation
09:50:33 noonedeadpunk As I was mixing up discovery flow
09:50:45 bauzas noonedeadpunk: maybe this would help you :p https://docs.openstack.org/nova/latest/admin/cells.html#faqs
09:51:31 bauzas but yeah, discover_hosts is idempotent so you can call it anytime you want
09:52:14 bauzas this is explained at the end of https://docs.openstack.org/nova/latest/admin/cells.html#configuring-a-new-deployment
09:52:38 noonedeadpunk well, I was unsure about logic, as I could not fully understand the reason why we're waiting for compute to be present in service list when nova_discover_hosts_in_cells_interval=-1 but run cell_v2 discover_hosts only after that
09:53:05 bauzas noonedeadpunk: we have two different databases now as a reminder
09:53:17 bauzas so we need to make sure we synchronize them
09:53:24 noonedeadpunk As for some reason I thought that discover_hosts also will add to service list, but now I know it;s not how it works)
09:54:30 sean-k-mooney bauzas: technially we have 3. we alway have api + cell 0 and cell one in any real deployment :)
09:57:21 bauzas sean-k-mooney: we have two different table schemas, but yeah :)
09:57:36 bauzas cell0 is logically identically to any cell
09:57:40 bauzas identical*
09:57:54 bauzas only the usage of this cell is different
09:58:05 bauzas but we're bikeshedding :)
09:59:58 noonedeadpunk Another thing I am unsure about - db online_data_migrations and nova-status upgrade check and when they can/should be run.
10:00:12 noonedeadpunk As far as I got - db online_data_migrations should be done only during upgrades
10:00:34 noonedeadpunk But is it recommended to run `nova-status upgrade check` before that?
10:01:26 noonedeadpunk As it feels like upgrade check should be run once all computes are upgraded, while online_data_migrations requires only conductor to be done?
10:05:16 bauzas noonedeadpunk: this is correct
10:05:44 bauzas you should run nova-status upgrade check *before* the db upgrade as a pre-flight check
10:06:24 bauzas nova-status is a controller script, checking a few APIs and DBs
10:06:33 bauzas to see whether you're safe to upgrade
10:07:02 bauzas like, say we introduced a thing but wasn't the default
10:07:27 bauzas before we flip the default and remove the old bits, we propose you to check that you opted in before
10:07:45 noonedeadpunk aha, yes, exactly what I was unsure about. As currently we run upgrade check basically after cell_v2 discover_hosts at the very end
10:08:18 noonedeadpunk which didn't make much sense to me
10:08:43 noonedeadpunk is there any way to tell if online_data_migrations are needed at all? Like by some exit code of smth?
10:09:12 bauzas noonedeadpunk: see the upgrade docs, this explains it way better than me now https://docs.openstack.org/nova/latest/admin/upgrades.html#rolling-upgrade-process
10:10:06 bauzas noonedeadpunk: wait, I said something wrong
10:10:29 bauzas noonedeadpunk: https://docs.openstack.org/nova/latest/cli/nova-status.html#upgrade
10:11:25 bauzas you need to do DB schema expands/contracts before with the migrations
10:11:38 bauzas then, you run the status check
10:11:51 bauzas and only then you need to start the compute services
10:12:37 bauzas neither the online data migrations
10:15:11 bauzas noonedeadpunk: as a summary, say that during a Bobcat release we may add some data online migrations
10:15:33 noonedeadpunk ok, so status check before service restart but after online migrations?
10:15:39 bauzas we would provide the nova-manage command to perform such data migrations
10:16:11 bauzas accordingly, we would propose a nova-status check addition that would check whether the data is fully migrated
10:16:38 bauzas so when upgrading to C, you would run this nova-status check that would fail if you forgot to run the online data migrations command
10:16:56 bauzas same goes with Placement
10:17:26 bauzas say that by Bobcat, we start using a new Placement API call
10:17:38 bauzas we would need to continue supporting the old Placement calls
10:18:13 bauzas but if we start using that sole version of Placement, then we would need to add a nova-status check that would fail if you forgot to update Placement
10:19:36 bauzas I just hope I haven't lost you :)
10:25:06 wangrong dansmith: sean-k-mooney: Thank you for your feedback in generic-vdpa spec. Regarding some of the issues you raised, I have explained them in my reply. Could you please take a look at the submission again. Thank you for your assistance. https://review.opendev.org/c/openstack/nova-specs/+/879338
10:25:52 sean-k-mooney i saw i skimmed it breifly this morning
10:25:56 gibi sean-k-mooney: https://bugs.launchpad.net/nova/+bug/2008716 this seems to be a real bug in our migration code but you had comments before reagarding different dbs being targeted. I don't see how that is possible given the error message states the that column is missing not the table
10:26:15 noonedeadpunk bauzas: sorry, was side-pinged :(
10:26:19 sean-k-mooney wangrong: in general i think we will still need to change the design singnifcalty form what you orginaly propsosed
10:27:07 noonedeadpunk that makes total sense.
10:27:44 noonedeadpunk So yes, once I've upgraded packages on Bobcat, I should do migrations, and then check upgrade to ensure it's safe to restart now
10:35:48 noonedeadpunk ok, now will try to make some patches out of this discussion

Earlier   Later