Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-28
16:27:45 cfriesen dansmith: got it, thanks. wonder where I got flipped around.
16:28:12 dansmith cfriesen: it's unfortunately very convoluted for little gain :(
16:30:42 cfriesen so how does "compute_node = self._get_compute_info(context, self.host)" make any sense then, if there could be multiple compute nodes per host?
16:30:59 dansmith cfriesen: is that in the resource tracker?
16:31:44 cfriesen nova/compute/manager.py
16:31:56 dansmith cfriesen: oh, look at the function
16:31:58 dansmith cfriesen: it returns the first
16:32:06 dansmith "for old compat"
16:32:35 dansmith and since ironic doesn't support live migration...YET
16:38:14 openstackgerrit Jay Pipes proposed openstack/nova master: placement: set/check if inventory change in tree https://review.openstack.org/470575
16:38:15 openstackgerrit Jay Pipes proposed openstack/nova master: placement: add nested resource providers https://review.openstack.org/377138
16:38:15 openstackgerrit Jay Pipes proposed openstack/nova master: placement: integrate ProviderTree to report client https://review.openstack.org/415921
16:38:16 openstackgerrit Jay Pipes proposed openstack/nova master: placement: adds REST API for nested providers https://review.openstack.org/384807
16:38:16 openstackgerrit Jay Pipes proposed openstack/nova master: placement: allow filter providers in tree https://review.openstack.org/377215
16:38:17 openstackgerrit Jay Pipes proposed openstack/nova master: placement: update client to set parent provider https://review.openstack.org/385693
16:43:04 mriedem stvnoyes: the multinode live migration grenade job failed on your live migration new style attach patch, but it failed here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/subnode-2/screen-n-cpu.txt.gz?level=TRACE#_Sep_26_14_28_11_637958
16:43:25 mriedem and looking at logstash, that happens on a ton of patches, so it's probably just a 100% failure in that job right now
16:43:42 dansmith jaypipes: I'm missing the reasoning for the complicated "compare two dicts" method in that first patch
16:44:38 jaypipes dansmith: yeah, it's a long conversation on previous revisions... I'm going to change it to only compare the total field value and ignore the other things like reserved/min_unit/allocation_ratio, etc
16:45:14 dansmith jaypipes: oh, sorry I can re-read.. but, is this not to be used for "know when I need to report new inventory" sort of thing?
16:45:18 jaypipes dansmith: the reason it's like that is because we still don't have any consistency and agreement on who (virt driver or resource tracker) owns various fields about the inventory like alloc ratio etc
16:45:28 efried dansmith The history is here: https://review.openstack.org/#/c/470575/2/nova/compute/provider_tree.py@98
16:46:08 dansmith jaypipes: ah, to avoid us reporting new just because the virt driver thinks the allocation_ratio should be different?
16:46:21 jaypipes yup
16:47:05 dansmith hmm
16:47:16 dansmith the virt driver needs to report total and reserved
16:47:42 dansmith but allocation_ratio is clearly the purview of the compute manager
16:47:56 dansmith min_unit is probably virt driver I guess
16:48:15 dansmith taking the scaleio example of 8GB minimum slice or whatever
16:51:06 mriedem stvnoyes: yeah the grenade live migration job is just totally busted http://tinyurl.com/y9ca6zox
16:51:53 mriedem something changed in august http://tinyurl.com/yby7m5p8
16:52:27 stvnoyes that's pretty bad
16:53:45 efried johnthetubaguy jaypipes Where does ironic (or ftm any virt driver) set up its resource providers today?
16:53:59 jaypipes efried: it doesn't. the RT does.
16:54:06 jaypipes efried: because there's only one RP.
16:54:32 mriedem stvnoyes: looking at http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/console.html#_2017-09-26_14_28_03_894125
16:54:37 efried jaypipes In ironic, only one RP per compute host? Not one RP per ironic node?
16:54:41 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:50 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:55 mriedem oops
16:55:02 jaypipes efried: we want to get to where there is a sort of mediation/negotiation between the RT and the virt driver and the generic device manager as each of them "processes" the ProviderTree and sets inventory records and traits based on their own information.
16:55:03 mriedem tempest.api.compute.admin.test_live_migration.LiveMigrationTest.test_live_block_migration [10.052345s] ... FAILED
16:55:17 mriedem so it works with the 2.25 microversion, where we pass block_migration=auto
16:55:21 mriedem but fails before that
16:55:22 jaypipes efried: in ironic, each baremetal Ironic node is its own RP in placementy.
16:55:48 mriedem so i bet something with the auto changes broke how we do live block migration between n-1 hosts
16:56:25 stvnoyes what were the auto changes?
16:56:26 efried jaypipes Aha. Then bug 1714248 exists because... the compute host is creating its own separate RP for some reason?
16:56:28 openstack bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] https://launchpad.net/bugs/1714248
16:56:53 jaypipes efried: not sure, lemme read
16:57:18 efried oh, or because the RT is trying to create a new RP for the node it's taking over...
16:57:29 mriedem stvnoyes: https://docs.openstack.org/nova/pike/reference/api-microversion-history.html#maximum-in-mitaka
16:57:29 jaypipes efried: note that nova-compute was *never* intended for "HA setups"...
16:57:39 mriedem ^ predates the spike in failures though, which started around 8/18
16:57:48 efried ...And that's out of the control of ironic-specific code...
16:58:28 efried which is why my comment #5 is a non-starter
16:58:42 efried until, as you say, there's some coordination between RT and virt to manage the RPs.
16:59:00 stvnoyes mriedem: brb- grabbing lunch
17:00:53 mriedem stvnoyes: ok. it also fails on the pike compute...
17:04:22 johnthetubaguy efried: I have to head out now and cook my dinner, but I think this maps nicely to how we setup traits, I am leaning towards ironic doing things in placement
17:05:02 efried johnthetubaguy Okay. I've got a passel of draft comments on your spec, which should be ready for your perusal next time you're on.
17:05:13 johnthetubaguy cool, thanks
17:06:12 dansmith jaypipes: in the db migration, why aren't we referencing the parent provider by id instead of uuid?
17:06:36 dansmith jaypipes: and, isn't recording the root just going to limit us later when we need to restructure a tree?
17:06:39 efried johnthetubaguy (Looking back over 'em, removing the ones we've talked about in here, the only things left are typos ):
17:07:39 jaypipes dansmith: a good question on the parent provider UUID thing. not entirely sure why I did that.
17:09:22 dansmith jaypipes: I can't tell you how elated I feel -1ing a db schema change of yours for performance reasons
17:09:34 jaypipes dansmith: recording the root is an optimization to avoid needing to do hierarchical queries. and we don't restructure the root, only potentially children within the tree (in other words, root_provider_id won't change.
17:09:41 jaypipes dansmith: :)
17:10:05 dansmith jaypipes: right, it won't for compute nodes, but it could for other types of resources
17:10:22 dansmith jaypipes: like you move a disk shelf from one NAS device to another
17:11:08 jaypipes dansmith: possibly, sure, but those kinds of moves are few and far between in comparison to the 10 or 100 times as many read requests for tree data
17:11:44 dansmith hmm
17:13:33 dansmith jaypipes: well, I commented for later
17:13:45 jaypipes okey dokey
17:14:24 dansmith jaypipes: I haven't gotten to how we tell placement about our parent, but presumably we're forbidden from trying to reparent a first-level provider?
17:15:05 jaypipes dansmith: haven't gotten to that either.
17:15:54 dansmith jaypipes: also, in case you're wondering
17:16:11 dansmith jaypipes: yes, it's amazingly beautiful out here on the deck... 71F and clear skies
17:16:20 jaypipes dansmith: lol
17:26:57 openstackgerrit Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275
17:32:52 mriedem stvnoyes: aha
17:32:58 mriedem it's a ci job configuration issue
17:33:06 mriedem the nodes are configured differently for block migration
17:36:05 stvnoyes mriedem: that's good to hear. much better than a code/upgrade issue.
17:36:46 mriedem i'm pretty sure i've had to fix this before...
17:44:37 eandersson Is versioned notifications properly implemented in Mitaka? We don't see any versioned notifications being sent.
17:44:40 openstackgerrit Dan Smith proposed openstack/nova master: Use improved instance_list module in compute API https://review.openstack.org/505418
17:44:41 openstackgerrit Dan Smith proposed openstack/nova master: Fix minor input items from previous patches https://review.openstack.org/506416
17:44:41 openstackgerrit Dan Smith proposed openstack/nova master: Fix CellDatabases fixture swallowing exceptions https://review.openstack.org/506312
17:56:40 mriedem eandersson: probably not at that point
17:57:11 mriedem introduced in newton https://specs.openstack.org/openstack/nova-specs/specs/newton/implemented/versioned-notification-transformation-newton.html
17:57:37 mriedem the framework code was in mitaka
17:58:03 eandersson I see - thanks mriedem
18:11:13 mriedem mtreinish: do you know anything about how tempest is upgraded in grenade?
18:12:13 mtreinish mriedem: it's not, the code and config should be the same between versions
18:12:20 mtreinish everything should be master for tempest
18:12:20 sean-k-mooney mriedem: i taught tempest was not upgreaed in grenade because it was not ment to be version specific
18:12:33 mriedem mtreinish: the config is different between pike and queens
18:12:43 mtreinish mriedem: links?

Earlier   Later