| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-11-29 | |||
| 23:26:09 | mriedem | mgagne: i don't think you've been able to skip the online data migrations since kilo | |
| 23:26:18 | mriedem | or shouldn't have | |
| 23:26:29 | mgagne | well, I upgraded from kilo to mitaka just fine =) | |
| 23:26:31 | mriedem | we've put in blocker schema migrations if you haven't completed the online data migrations | |
| 23:26:41 | mriedem | sure, maybe you just didn't need to migrate any of that stuff | |
| 23:27:20 | mgagne | well, there is a glitch with cellsv1 where service rpc version is checked in database and since you aren't running conductor or scheduler in API, migration fails. I just commented that part. | |
| 23:27:48 | mriedem | i might have been thinking of this https://github.com/openstack/nova/blob/stable/newton/nova/db/sqlalchemy/migrate_repo/versions/330_enforce_mitaka_online_migrations.py | |
| 23:27:58 | mriedem | newton schema migration requires that you have done the mitaka online data migrations | |
| 23:28:11 | mgagne | yes, I'm not there yet :P | |
| 23:29:08 | mriedem | "there is a glitch with cellsv1 where service rpc version is checked in database and since you aren't running conductor or scheduler in API, migration fails. I just commented that part." | |
| 23:29:11 | mriedem | where does that happen? | |
| 23:29:29 | melwitt | mriedem: just to clarify, if someone installs nova ocata, they will be prevented from installing the latest os-brick because of upper-constraints? or no? trying to understand the difference between them updating to a os-brick .z release vs just pulling the latest os-brick | |
| 23:29:46 | mriedem | that does remind me of something i think i've talked with dansmith about before which i don't think we're handling yet, maybe we are, but when the api checks the minimum compute service version, i don't know if it's iterating all cells | |
| 23:30:09 | mriedem | melwitt: totally depends on how nova is installed | |
| 23:30:20 | mgagne | will need to check | |
| 23:30:23 | mriedem | if you're pip installing into a tox venv from the nova git repo, then sure :) | |
| 23:30:28 | mriedem | but i doubt most people are doing that | |
| 23:30:40 | melwitt | mriedem: :\ okay. I was thinking whether going through all the .z release stuff helps much or not | |
| 23:30:49 | mriedem | fedora rpm specs null out the requirements.txt so that the rpms are just installed from the versions in the rpm spec fedora says is good | |
| 23:31:19 | dansmith | mriedem: it is iterating | |
| 23:31:41 | melwitt | I see | |
| 23:31:44 | mriedem | get_minimum_version_all_cells is | |
| 23:32:29 | mriedem | dansmith: looks like we only use that in one place, server create for device tagging, | |
| 23:32:45 | mriedem | but, | |
| 23:32:56 | dansmith | hmm, actually | |
| 23:33:00 | mriedem | anything else acting on an instance would be scoped to the cell that instance is in, so min service version checks scoped to that cell would be fine | |
| 23:33:01 | dansmith | now that you mention it, | |
| 23:33:18 | mriedem | i'm thinking about your conductor migration allocation min service versoin check | |
| 23:34:00 | mriedem | when we do that, the context should be targeted to the cell that the instance was pulled from | |
| 23:34:14 | dansmith | the problem is we cache it in places | |
| 23:34:19 | mriedem | https://github.com/openstack/nova/blob/master/nova/conductor/tasks/migrate.py#L148 | |
| 23:35:01 | mriedem | oh hmm | |
| 23:35:06 | dansmith | https://github.com/openstack/nova/blob/master/nova/objects/service.py#L396 | |
| 23:35:14 | mriedem | so min compute in cell1 might be 22 and 23 in cell2 | |
| 23:35:21 | mriedem | and if we cached 23, we could f up cell1 | |
| 23:35:25 | mriedem | ? | |
| 23:35:31 | dansmith | I was going to say we iterate it where we need it and otherwise it's scoped per instance, but.... | |
| 23:36:02 | mriedem | yeah we cache in the api, compute, conductor and scheduler services | |
| 23:36:16 | mriedem | compute is probably fine right? it's scoped to a cell anyway | |
| 23:36:25 | mriedem | but api/scheduler/conductor could be a problem | |
| 23:36:46 | dansmith | right "inside the cell things" are fine | |
| 23:37:59 | dansmith | man I wish we had the CellMapping in the context so we could make the cache just use that to key | |
| 23:38:02 | mgagne | so I can't find the exact cause but I think it was looking for service entries in the API cell database and didn't exclude deleted rows. | |
| 23:38:39 | dansmith | mgagne: I have no doubt that none of this works the way it should for cellsv1 | |
| 23:39:17 | dansmith | mixed-version computes I mean | |
| 23:39:17 | mgagne | or could be with the online migration | |
| 23:39:36 | mriedem | services entries are in the cell dbs | |
| 23:39:38 | mriedem | not the api db | |
| 23:40:11 | mgagne | here https://github.com/openstack/nova/blob/mitaka-eol/nova/objects/pci_device.py#L118-L127 | |
| 23:40:24 | mgagne | if no conductor service is found, the value returned fails the check | |
| 23:40:53 | mgagne | well, maybe I shouldn't have run the online migration in the API cell? ¯\_(ツ)_/¯ | |
| 23:41:44 | mriedem | hm | |
| 23:42:37 | mriedem | dansmith: we create the api/scheduler/conductor services in the cell0 db don't we? | |
| 23:42:47 | dansmith | yup | |
| 23:42:50 | dansmith | I mean, in devstack | |
| 23:42:54 | mriedem | right | |
| 23:43:26 | mriedem | yeah http://logs.openstack.org/87/523187/2/check/legacy-tempest-dsvm-neutron-full/6b78222/logs/etc/nova/nova.conf.txt.gz | |
| 23:43:29 | mriedem | [database] connection = mysql+pymysql://root:secretmysql@127.0.0.1/nova_cell0?charset=utf8 | |
| 23:43:40 | dansmith | dude, passwords! | |
| 23:46:09 | mriedem | mgagne: i wonder if you had this fix before you upgraded https://review.openstack.org/#/q/Ic96a5eb3728f97a3c35d2c5121e6fdcd4fd1c70b | |
| 23:46:16 | mriedem | https://review.openstack.org/#/c/438632/ | |
| 23:46:48 | mgagne | yes | |
| 23:47:07 | mgagne | but if no entry is found, I think it returns 0 ou None and it fails | |
| 23:47:59 | mriedem | yeah you're right https://github.com/openstack/nova/blob/master/nova/objects/service.py#L431-L434 | |
| 23:48:03 | mgagne | but I don't know if I had to run the migration in api cell or not | |
| 23:48:54 | mgagne | if i shouldn't, well I think the migration script should have told me: hey, this is an api cell, you shouldn't do that. But I understand that cellsv1 isn't fully tested so yea, what can you do =) | |
| 23:49:26 | mriedem | https://github.com/openstack/nova/commit/50355c4595e08f293f610da32247e405b20c1c5b | |
| 23:49:44 | mriedem | yeah my guess is there was no consideration for cells v1 when that was written in mitaka | |
| 23:49:51 | mriedem | and we don't have grenade (upgrade) ci jobs for cellsv1 | |
| 23:50:18 | mriedem | cellsv1 upgrade testing was usually literally alaski or johnthetubaguy saying something broke at rax | |
| 23:50:31 | mgagne | ok, that's fine, mitaka migration is behind us. but I guess same migration will fail again with newton if I try to run it in api cell. | |
| 23:50:42 | mriedem | that code was dropped in newton | |
| 23:50:50 | mriedem | because we have that schema migration blocker in newton | |
| 23:50:56 | mriedem | https://github.com/openstack/nova/blob/stable/newton/nova/db/sqlalchemy/migrate_repo/versions/330_enforce_mitaka_online_migrations.py | |
| 23:51:11 | mriedem | but ^ assumes the child cell db | |
| 23:51:33 | mriedem | so nova-manage db sync | |
| 23:51:38 | mriedem | not nova-manage api_db sync | |
| 23:51:43 | mgagne | that's a fun one: "until all records have been migrated". Ok, let's run that migration then! oh way, code is gone. what now? /sad panda | |
| 23:51:47 | mriedem | so you should be fine | |
| 23:52:17 | mgagne | but then, you will say there is an upgrade readiness check now you can run | |
| 23:52:27 | mgagne | which I'm fine with. ;) | |
| 23:52:30 | mriedem | nova-status was added in ocata | |
| 23:53:20 | mriedem | and checked for things like making sure placement was deployed and cellsv2 mappings existed | |
| 23:53:22 | mgagne | yea, I had a lot of "fun" when it complained about the online migration and old code was gone. had to find a copy in a different environment and rsync that thing. | |
| 23:54:16 | mriedem | i thought people stood up a separate env to run the db sync on the new code before upgrading the old code that's actually running? | |
| 23:54:52 | mgagne | I guess I'm not in that ideal world yet :P | |
| 23:55:33 | mriedem | you can also run the online data migrations from the old mitaka code before upgrading to newton, and run them after upgrading to newton if yo uwant | |
| 23:55:36 | mgagne | usually I stop all services, upgrade package, run db sync, start service. | |
| 23:56:19 | mgagne | but if I forgot to run online migration and db sync fails, I'm screwed because I already upgraded the packages. but my bad for not checking if all online migration ran properly before the upgrade. | |
| 23:56:54 | mgagne | sure but I wasn't prepared for that maneuver | |
| 23:58:02 | mriedem | mgagne: even with that blocker migration script in newton, we didn't delete the cold that allows you to run the online data migrations from mitaka https://github.com/openstack/nova/blob/stable/newton/nova/cmd/manage.py#L787 | |
| 23:58:08 | mriedem | so, we didn't hose you there | |
| 23:58:18 | mriedem | we just said, you can't continue until you do your homework from mitaka | |
| 23:58:21 | mgagne | very much appreciated =) | |
| 23:59:31 | mriedem | pretty sure that's standard operating procedure | |
| 23:59:41 | mriedem | in queens we still have online data migration code from newton | |
| 23:59:57 | mriedem | we only remove it if we're at least n+1 and someone gets around to caring | |
| #openstack-nova - 2017-11-30 | |||
| 00:00:18 | mriedem | e.g. https://review.openstack.org/#/c/517158/ | |
| 00:01:07 | mriedem | mgagne: btw, that's another reason people can't/shoudn't literally skip through releases for upgrades | |
| 00:02:06 | mgagne | mriedem: I can't afford to not skip versions ;) | |