| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-11-29 | |||
| 23:50:18 | mriedem | cellsv1 upgrade testing was usually literally alaski or johnthetubaguy saying something broke at rax | |
| 23:50:31 | mgagne | ok, that's fine, mitaka migration is behind us. but I guess same migration will fail again with newton if I try to run it in api cell. | |
| 23:50:42 | mriedem | that code was dropped in newton | |
| 23:50:50 | mriedem | because we have that schema migration blocker in newton | |
| 23:50:56 | mriedem | https://github.com/openstack/nova/blob/stable/newton/nova/db/sqlalchemy/migrate_repo/versions/330_enforce_mitaka_online_migrations.py | |
| 23:51:11 | mriedem | but ^ assumes the child cell db | |
| 23:51:33 | mriedem | so nova-manage db sync | |
| 23:51:38 | mriedem | not nova-manage api_db sync | |
| 23:51:43 | mgagne | that's a fun one: "until all records have been migrated". Ok, let's run that migration then! oh way, code is gone. what now? /sad panda | |
| 23:51:47 | mriedem | so you should be fine | |
| 23:52:17 | mgagne | but then, you will say there is an upgrade readiness check now you can run | |
| 23:52:27 | mgagne | which I'm fine with. ;) | |
| 23:52:30 | mriedem | nova-status was added in ocata | |
| 23:53:20 | mriedem | and checked for things like making sure placement was deployed and cellsv2 mappings existed | |
| 23:53:22 | mgagne | yea, I had a lot of "fun" when it complained about the online migration and old code was gone. had to find a copy in a different environment and rsync that thing. | |
| 23:54:16 | mriedem | i thought people stood up a separate env to run the db sync on the new code before upgrading the old code that's actually running? | |
| 23:54:52 | mgagne | I guess I'm not in that ideal world yet :P | |
| 23:55:33 | mriedem | you can also run the online data migrations from the old mitaka code before upgrading to newton, and run them after upgrading to newton if yo uwant | |
| 23:55:36 | mgagne | usually I stop all services, upgrade package, run db sync, start service. | |
| 23:56:19 | mgagne | but if I forgot to run online migration and db sync fails, I'm screwed because I already upgraded the packages. but my bad for not checking if all online migration ran properly before the upgrade. | |
| 23:56:54 | mgagne | sure but I wasn't prepared for that maneuver | |
| 23:58:02 | mriedem | mgagne: even with that blocker migration script in newton, we didn't delete the cold that allows you to run the online data migrations from mitaka https://github.com/openstack/nova/blob/stable/newton/nova/cmd/manage.py#L787 | |
| 23:58:08 | mriedem | so, we didn't hose you there | |
| 23:58:18 | mriedem | we just said, you can't continue until you do your homework from mitaka | |
| 23:58:21 | mgagne | very much appreciated =) | |
| 23:59:31 | mriedem | pretty sure that's standard operating procedure | |
| 23:59:41 | mriedem | in queens we still have online data migration code from newton | |
| 23:59:57 | mriedem | we only remove it if we're at least n+1 and someone gets around to caring | |
| #openstack-nova - 2017-11-30 | |||
| 00:00:18 | mriedem | e.g. https://review.openstack.org/#/c/517158/ | |
| 00:01:07 | mriedem | mgagne: btw, that's another reason people can't/shoudn't literally skip through releases for upgrades | |
| 00:02:06 | mgagne | mriedem: I can't afford to not skip versions ;) | |
| 00:02:19 | mriedem | fast forwarding through versions is fine | |
| 00:02:23 | mriedem | but you have to run the data migratoins | |
| 00:02:26 | mriedem | per release | |
| 00:02:45 | mriedem | same issues with dropping config options after n+1 | |
| 00:03:09 | mgagne | yes, but with cells, I'm not sure if I will be able to skip anymore, too many unknown for now | |
| 00:03:37 | mgagne | mriedem: configs are fine, we have funky stuff in puppet to support multiple versions | |
| 00:03:54 | mriedem | like aliases? | |
| 00:06:01 | mgagne | very funky stuff: https://gist.github.com/mgagne/7146424416eda597563c4018ce50cf97 | |
| 00:06:08 | mgagne | copied as-is so you can see our mess | |
| 00:06:33 | mgagne | this class is included in our main nova.pp which does the main configuration | |
| 00:08:14 | mgagne | so I just make it so puppet-nova for newton works with mitaka. and I do the same with other services/modules | |
| 00:08:49 | tssurya_ | mriedem : actually belmiro is currently on newton, moving to ocata (which would be only for a short duration), but main goal is pike. | |
| 00:09:44 | tssurya_ | mridem, dansmith : http://eavesdrop.openstack.org/irclogs/%23openstack-nova/%23openstack-nova.2017-10-24.log.html#t2017-10-24T13:12:39 , the conversation you guys had regarding placement, | |
| 00:10:06 | mriedem | tssurya_: thanks, put that into https://etherpad.openstack.org/p/cellsv1-to-v2-migration | |
| 00:10:17 | tssurya_ | mriedem : sure ! | |
| 00:14:16 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Enable cold migration with target host(2/2) https://review.openstack.org/408964 | |
| 00:14:17 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add multi-cell negative test for cold migration with target host https://review.openstack.org/524027 | |
| 00:16:09 | mgagne | mriedem: ok so if you run Cellsv1, you should run placement per cell otherwise nova-scheduler in cell *could* pickup hosts from a different cell? | |
| 00:16:56 | dansmith | it will | |
| 00:17:01 | dansmith | and will have to filter them out | |
| 00:17:14 | mgagne | "and will have to filter them out" how? | |
| 00:17:28 | dansmith | by the ones it has host state for | |
| 00:17:28 | mriedem | the scheduler does a db query to the compute_nodes table per cell | |
| 00:17:54 | mriedem | 1. scheduler asks placement for resource providers (compute nodes) for a given request (flavor) | |
| 00:18:05 | mriedem | 2. scheduler queries the cell db for the compute nodes by the list of uuids from placement | |
| 00:18:17 | mriedem | 3. scheduler converts those compute nodes to HostState objects and those go through the enabled filters | |
| 00:18:57 | mgagne | does it mean UUID returned by 2) from placement would be filtered by the ones found in compute nodes in cell db? | |
| 00:19:35 | mriedem | yeah https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L628 | |
| 00:19:53 | mriedem | starts here https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L645 | |
| 00:20:50 | mriedem | although, that's pike+ | |
| 00:20:51 | mriedem | https://github.com/openstack/nova/commit/d1de5385233ce4379b17a7404557c6724dc37cd4 | |
| 00:23:23 | mgagne | I'm more concerned about cellsv1, I feel like linked commit is for cellsv2 where scheduling happens in api/top cell? Or am I mistaken? | |
| 00:23:49 | mriedem | with cells v2 yeah there is just a flat scheduler | |
| 00:23:55 | mriedem | which is multi-cell aware starting in pike | |
| 00:25:11 | mriedem | my brain is getting fried and i can already hear my wife getting mad because i'm in my office after 6pm, | |
| 00:25:30 | mriedem | but i'm not sure if https://review.openstack.org/#/c/436648/ makes a significant difference in the decision to run per-cell placement or not in ocata | |
| 00:25:50 | mriedem | because the nova-scheduler in each cell in ocata would only query computes from the cell db it's configured for | |
| 00:26:06 | mriedem | so placement might say there are 1000 computes that could satisfy a request, but only 100 of those might be in the cell db that that cell knows about | |
| 00:26:16 | mriedem | which is this https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/host_manager.py#L579 | |
| 00:26:46 | mriedem | so i think global placement is still ok in ocata (still not sure why belmiro wanted per-cell placement, would have to read the notes again) | |
| 00:28:03 | mgagne | alright, I guess it addresses my concern. would need to test for sure but I get the idea | |
| 00:29:01 | mriedem | i put more notes about this in https://etherpad.openstack.org/p/cellsv1-to-v2-migration | |
| 00:29:09 | mriedem | would need dansmith to confirm my thinking, but i think global placement would be fine | |
| 00:29:58 | dansmith | we get the host list from placement in ocata IIRC, so I think its' the same as pike | |
| 00:30:01 | mriedem | given that plus the catalog issue, i think global placement would be the way to go unless i'm missing something | |
| 00:30:10 | dansmith | but I am also rather fried for the day | |
| 00:30:14 | mriedem | right, we get the host list from placement in ocata, | |
| 00:30:23 | mriedem | and then get the compute nodes from the uuids placement returned https://etherpad.openstack.org/p/cellsv1-to-v2-migration | |
| 00:30:25 | mriedem | oops | |
| 00:30:28 | mriedem | https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/host_manager.py#L579 | |
| 00:30:36 | mriedem | which in ocata are just for the cell that the scheduler has access to | |
| 00:30:43 | mriedem | but, that's fine, it's just an IN query | |
| 00:30:50 | mriedem | yo'ud get the subset that the child cell db knows about | |
| 00:31:13 | mriedem | the inefficiency would be on the placement side | |
| 00:31:18 | dansmith | it's still going to return all the nodes in the system from the placement call | |
| 00:31:22 | dansmith | and make a massive SQL query | |
| 00:31:39 | mriedem | that's no different in pike though | |
| 00:31:45 | dansmith | that's what I said | |
| 00:31:58 | mriedem | heh, yeah | |
| 00:32:14 | mriedem | well, shit, let's just add a cell mapping to placement, i'm sure cdent and edleafe and jaypipes would be on board with that | |
| 00:32:54 | mriedem | with a grenade, we're both going to die right? | |
| 00:32:58 | mriedem | in close proximity | |
| 00:33:05 | dansmith | if I let go | |
| 00:33:07 | dansmith | that's the idea | |
| 00:33:26 | mgagne | can those details be added to the etherpad? like you should be running global placement. it will return all UUIDs of all compute nodes but scheduler will filter them out. And there are plans to optimize that. | |
| 00:33:46 | mriedem | mgagne: i added something along those lines in the pros/cons sections for global vs local placement | |
| 00:33:51 | mriedem | might not be coherent at this point | |
| 00:33:59 | mgagne | ok, I didn't fully understand what it meant ^^' | |
| 00:34:01 | mriedem | dansmith: i just like the idea of saying we should do that and watching heads explode | |