Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-28
16:44:38 jaypipes dansmith: yeah, it's a long conversation on previous revisions... I'm going to change it to only compare the total field value and ignore the other things like reserved/min_unit/allocation_ratio, etc
16:45:14 dansmith jaypipes: oh, sorry I can re-read.. but, is this not to be used for "know when I need to report new inventory" sort of thing?
16:45:18 jaypipes dansmith: the reason it's like that is because we still don't have any consistency and agreement on who (virt driver or resource tracker) owns various fields about the inventory like alloc ratio etc
16:45:28 efried dansmith The history is here: https://review.openstack.org/#/c/470575/2/nova/compute/provider_tree.py@98
16:46:08 dansmith jaypipes: ah, to avoid us reporting new just because the virt driver thinks the allocation_ratio should be different?
16:46:21 jaypipes yup
16:47:05 dansmith hmm
16:47:16 dansmith the virt driver needs to report total and reserved
16:47:42 dansmith but allocation_ratio is clearly the purview of the compute manager
16:47:56 dansmith min_unit is probably virt driver I guess
16:48:15 dansmith taking the scaleio example of 8GB minimum slice or whatever
16:51:06 mriedem stvnoyes: yeah the grenade live migration job is just totally busted http://tinyurl.com/y9ca6zox
16:51:53 mriedem something changed in august http://tinyurl.com/yby7m5p8
16:52:27 stvnoyes that's pretty bad
16:53:45 efried johnthetubaguy jaypipes Where does ironic (or ftm any virt driver) set up its resource providers today?
16:53:59 jaypipes efried: it doesn't. the RT does.
16:54:06 jaypipes efried: because there's only one RP.
16:54:32 mriedem stvnoyes: looking at http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/console.html#_2017-09-26_14_28_03_894125
16:54:37 efried jaypipes In ironic, only one RP per compute host? Not one RP per ironic node?
16:54:41 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:50 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:55 mriedem oops
16:55:02 jaypipes efried: we want to get to where there is a sort of mediation/negotiation between the RT and the virt driver and the generic device manager as each of them "processes" the ProviderTree and sets inventory records and traits based on their own information.
16:55:03 mriedem tempest.api.compute.admin.test_live_migration.LiveMigrationTest.test_live_block_migration [10.052345s] ... FAILED
16:55:17 mriedem so it works with the 2.25 microversion, where we pass block_migration=auto
16:55:21 mriedem but fails before that
16:55:22 jaypipes efried: in ironic, each baremetal Ironic node is its own RP in placementy.
16:55:48 mriedem so i bet something with the auto changes broke how we do live block migration between n-1 hosts
16:56:25 stvnoyes what were the auto changes?
16:56:26 efried jaypipes Aha. Then bug 1714248 exists because... the compute host is creating its own separate RP for some reason?
16:56:28 openstack bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] https://launchpad.net/bugs/1714248
16:56:53 jaypipes efried: not sure, lemme read
16:57:18 efried oh, or because the RT is trying to create a new RP for the node it's taking over...
16:57:29 mriedem stvnoyes: https://docs.openstack.org/nova/pike/reference/api-microversion-history.html#maximum-in-mitaka
16:57:29 jaypipes efried: note that nova-compute was *never* intended for "HA setups"...
16:57:39 mriedem ^ predates the spike in failures though, which started around 8/18
16:57:48 efried ...And that's out of the control of ironic-specific code...
16:58:28 efried which is why my comment #5 is a non-starter
16:58:42 efried until, as you say, there's some coordination between RT and virt to manage the RPs.
16:59:00 stvnoyes mriedem: brb- grabbing lunch
17:00:53 mriedem stvnoyes: ok. it also fails on the pike compute...
17:04:22 johnthetubaguy efried: I have to head out now and cook my dinner, but I think this maps nicely to how we setup traits, I am leaning towards ironic doing things in placement
17:05:02 efried johnthetubaguy Okay. I've got a passel of draft comments on your spec, which should be ready for your perusal next time you're on.
17:05:13 johnthetubaguy cool, thanks
17:06:12 dansmith jaypipes: in the db migration, why aren't we referencing the parent provider by id instead of uuid?
17:06:36 dansmith jaypipes: and, isn't recording the root just going to limit us later when we need to restructure a tree?
17:06:39 efried johnthetubaguy (Looking back over 'em, removing the ones we've talked about in here, the only things left are typos ):
17:07:39 jaypipes dansmith: a good question on the parent provider UUID thing. not entirely sure why I did that.
17:09:22 dansmith jaypipes: I can't tell you how elated I feel -1ing a db schema change of yours for performance reasons
17:09:34 jaypipes dansmith: recording the root is an optimization to avoid needing to do hierarchical queries. and we don't restructure the root, only potentially children within the tree (in other words, root_provider_id won't change.
17:09:41 jaypipes dansmith: :)
17:10:05 dansmith jaypipes: right, it won't for compute nodes, but it could for other types of resources
17:10:22 dansmith jaypipes: like you move a disk shelf from one NAS device to another
17:11:08 jaypipes dansmith: possibly, sure, but those kinds of moves are few and far between in comparison to the 10 or 100 times as many read requests for tree data
17:11:44 dansmith hmm
17:13:33 dansmith jaypipes: well, I commented for later
17:13:45 jaypipes okey dokey
17:14:24 dansmith jaypipes: I haven't gotten to how we tell placement about our parent, but presumably we're forbidden from trying to reparent a first-level provider?
17:15:05 jaypipes dansmith: haven't gotten to that either.
17:15:54 dansmith jaypipes: also, in case you're wondering
17:16:11 dansmith jaypipes: yes, it's amazingly beautiful out here on the deck... 71F and clear skies
17:16:20 jaypipes dansmith: lol
17:26:57 openstackgerrit Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275
17:32:52 mriedem stvnoyes: aha
17:32:58 mriedem it's a ci job configuration issue
17:33:06 mriedem the nodes are configured differently for block migration
17:36:05 stvnoyes mriedem: that's good to hear. much better than a code/upgrade issue.
17:36:46 mriedem i'm pretty sure i've had to fix this before...
17:44:37 eandersson Is versioned notifications properly implemented in Mitaka? We don't see any versioned notifications being sent.
17:44:40 openstackgerrit Dan Smith proposed openstack/nova master: Use improved instance_list module in compute API https://review.openstack.org/505418
17:44:41 openstackgerrit Dan Smith proposed openstack/nova master: Fix minor input items from previous patches https://review.openstack.org/506416
17:44:41 openstackgerrit Dan Smith proposed openstack/nova master: Fix CellDatabases fixture swallowing exceptions https://review.openstack.org/506312
17:56:40 mriedem eandersson: probably not at that point
17:57:11 mriedem introduced in newton https://specs.openstack.org/openstack/nova-specs/specs/newton/implemented/versioned-notification-transformation-newton.html
17:57:37 mriedem the framework code was in mitaka
17:58:03 eandersson I see - thanks mriedem
18:11:13 mriedem mtreinish: do you know anything about how tempest is upgraded in grenade?
18:12:13 mtreinish mriedem: it's not, the code and config should be the same between versions
18:12:20 mtreinish everything should be master for tempest
18:12:20 sean-k-mooney mriedem: i taught tempest was not upgreaed in grenade because it was not ment to be version specific
18:12:33 mriedem mtreinish: the config is different between pike and queens
18:12:43 mtreinish mriedem: links?
18:12:47 mriedem at least in this multinode grenade job
18:12:48 mriedem http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/old/tempest_conf.txt.gz
18:12:51 mriedem ^ pike
18:12:58 mriedem note: block_migration_for_live_migration = True
18:13:06 mriedem http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/new/tempest_conf.txt.gz
18:13:08 mriedem queens ^
18:13:11 sean-k-mooney mriedem: the config is likely different yes but it should be useing the same tempest version
18:13:13 mriedem block_migration_for_live_migration = False
18:13:34 mriedem the config is supposed to be copied from old to new https://github.com/openstack-dev/grenade/blob/master/upgrade-tempest#L82
18:14:07 mriedem looks like that is happening here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/grenade.sh.txt.gz#_2017-09-26_14_23_46_489
18:14:11 mtreinish mriedem: yeah I don't know why that's being switched or where that's happening
18:15:43 sean-k-mooney is devstack regenerating the config on the second run and overriting it?
18:15:48 mriedem old tempest.conf gets it set here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/grenade.sh.txt.gz#_2017-09-26_13_58_48_782
18:16:09 mtreinish sean-k-mooney: config should just be copied from old devstack
18:16:14 mtreinish that's how grenade is supposed to work
18:16:28 mtreinish unless it's explicitly changed in the upgrade script
18:16:45 mtreinish which is not something we approve lightly because that means a manual upgrade step
18:17:22 sean-k-mooney mtreinish: yes but devstack generate a new tempest config by default when run so if grenage is reusing that logic it may be chagining it just a guess. you know more about this then i

Earlier   Later