Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-28
16:09:18 efried "going into queens" or "going away in queens" ?
16:09:25 johnthetubaguy going away in queens
16:09:37 johnthetubaguy oops, missed the important word there
16:09:47 cfriesen efried: there's an interesting short email chain here: https://openstack.nimeyo.com/91493/openstack-dev-ironic-nova-indivisible-resource-providers
16:10:26 efried johnthetubaguy I assume that's just an ironic statement
16:10:45 efried I honestly know nothing about ironic, learning as I go in the context of this chat.
16:10:52 johnthetubaguy that's what the comment in the Nova code was saying, based on plan A for Nova I believe?
16:11:25 johnthetubaguy https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L758
16:11:50 johnthetubaguy ironic resources are indivisable
16:12:01 johnthetubaguy custom resource classes model that nicely
16:12:46 efried Right, so the point will be to have a custom resource class representing a "kind" of node, such that any node associated with that resource class is effectively "identical" to any other.
16:13:14 johnthetubaguy well, "identical" from the point of view of the flavor
16:13:30 johnthetubaguy you might have several generations mapped to a single resource class, if you want
16:13:32 efried Right. And you're modeling the compute process as the resource *provider* of the resources of the various classes it can see.
16:13:50 efried (What do we call the thing the compute process runs on? The "hypervisor"?)
16:14:05 johnthetubaguy no, the resource provider is the ironic node uuid in Ironic
16:14:36 efried Where "node" isn't the baremetal machine, it's the thing that manages 'em, right?
16:14:42 johnthetubaguy nova-compute process manages multiple ironic nodes (that list varies depending on the hash ring, hence the bug)
16:14:52 johnthetubaguy so the host is nova-compute
16:14:55 efried ayee
16:15:03 johnthetubaguy the node is the ironic node
16:15:03 efried Terminology fail.
16:15:21 johnthetubaguy the node is the resource provider
16:15:32 cdent nova-compute hosts multiple ironic nodes, each of which have an inventory of customr resource class with count of 1
16:15:36 cdent (my memory is starting to refresh)
16:15:43 johnthetubaguy +1 cdent
16:16:37 efried But some of those ironic nodes can be using the same resource class, if they're functionally "identical"
16:16:38 efried ?
16:16:47 cdent yes
16:17:04 efried So it would seem like, as part of the failover process, the host that's taking over ought to delete the RP for the failed host.
16:17:13 johnthetubaguy yeah, each is offering 1: CUSTOM_GOLD
16:17:26 efried Seems like that ought to happen regardless, otherwise the scheduler will think it's still there and may try to schedule stuff to it.
16:17:35 johnthetubaguy so for that bug
16:17:40 efried yeah, for that bug
16:17:46 johnthetubaguy the dead thing would normally be responsible for the delete of it
16:17:59 efried That... doesn't make sense. It's dead. It can't call the placement API.
16:18:10 johnthetubaguy right, that's the problem
16:18:25 efried At some level, the new guy must know he's taking over for the old guy.
16:18:39 efried So why can't he delete the old guy's RP?
16:18:49 johnthetubaguy not sure if it sees that edge
16:18:59 johnthetubaguy it just sees a new node right now
16:19:09 johnthetubaguy but not 100% sure on that
16:19:10 efried Well.
16:19:23 efried Is there a split-brain possibility here?
16:19:44 cfriesen johnthetubaguy: for virtual machines (like with a vmware backend) nova-compute runs on the "node", and the hypervisor runs on the "host", right? why did ironic reverse the naming?
16:21:20 dansmith cfriesen: no, the reverse
16:21:30 dansmith cfriesen: but vmware doesn't actually run like that
16:21:34 dansmith (anymore)
16:21:41 dansmith well,
16:22:04 dansmith it doesn't expose a 1:N host:node structure I mean
16:22:05 dansmith it still runs external I think, but...
16:23:14 dansmith cfriesen: nova-compute always runs on the "host" which is where the service record is (nova-compute being a service) and the "node" is the actual hypervisor
16:23:36 dansmith ironic runs one nova-compute for multiple "hypervisor" nodes, which can't by definition run the service on them
16:24:05 dansmith ironic is the _reason_ we have that (blasted) model so it's setting the standard, not doing something reverse of what the model was designed for
16:27:45 cfriesen dansmith: got it, thanks. wonder where I got flipped around.
16:28:12 dansmith cfriesen: it's unfortunately very convoluted for little gain :(
16:30:42 cfriesen so how does "compute_node = self._get_compute_info(context, self.host)" make any sense then, if there could be multiple compute nodes per host?
16:30:59 dansmith cfriesen: is that in the resource tracker?
16:31:44 cfriesen nova/compute/manager.py
16:31:56 dansmith cfriesen: oh, look at the function
16:31:58 dansmith cfriesen: it returns the first
16:32:06 dansmith "for old compat"
16:32:35 dansmith and since ironic doesn't support live migration...YET
16:38:14 openstackgerrit Jay Pipes proposed openstack/nova master: placement: set/check if inventory change in tree https://review.openstack.org/470575
16:38:15 openstackgerrit Jay Pipes proposed openstack/nova master: placement: add nested resource providers https://review.openstack.org/377138
16:38:15 openstackgerrit Jay Pipes proposed openstack/nova master: placement: integrate ProviderTree to report client https://review.openstack.org/415921
16:38:16 openstackgerrit Jay Pipes proposed openstack/nova master: placement: adds REST API for nested providers https://review.openstack.org/384807
16:38:16 openstackgerrit Jay Pipes proposed openstack/nova master: placement: allow filter providers in tree https://review.openstack.org/377215
16:38:17 openstackgerrit Jay Pipes proposed openstack/nova master: placement: update client to set parent provider https://review.openstack.org/385693
16:43:04 mriedem stvnoyes: the multinode live migration grenade job failed on your live migration new style attach patch, but it failed here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/subnode-2/screen-n-cpu.txt.gz?level=TRACE#_Sep_26_14_28_11_637958
16:43:25 mriedem and looking at logstash, that happens on a ton of patches, so it's probably just a 100% failure in that job right now
16:43:42 dansmith jaypipes: I'm missing the reasoning for the complicated "compare two dicts" method in that first patch
16:44:38 jaypipes dansmith: yeah, it's a long conversation on previous revisions... I'm going to change it to only compare the total field value and ignore the other things like reserved/min_unit/allocation_ratio, etc
16:45:14 dansmith jaypipes: oh, sorry I can re-read.. but, is this not to be used for "know when I need to report new inventory" sort of thing?
16:45:18 jaypipes dansmith: the reason it's like that is because we still don't have any consistency and agreement on who (virt driver or resource tracker) owns various fields about the inventory like alloc ratio etc
16:45:28 efried dansmith The history is here: https://review.openstack.org/#/c/470575/2/nova/compute/provider_tree.py@98
16:46:08 dansmith jaypipes: ah, to avoid us reporting new just because the virt driver thinks the allocation_ratio should be different?
16:46:21 jaypipes yup
16:47:05 dansmith hmm
16:47:16 dansmith the virt driver needs to report total and reserved
16:47:42 dansmith but allocation_ratio is clearly the purview of the compute manager
16:47:56 dansmith min_unit is probably virt driver I guess
16:48:15 dansmith taking the scaleio example of 8GB minimum slice or whatever
16:51:06 mriedem stvnoyes: yeah the grenade live migration job is just totally busted http://tinyurl.com/y9ca6zox
16:51:53 mriedem something changed in august http://tinyurl.com/yby7m5p8
16:52:27 stvnoyes that's pretty bad
16:53:45 efried johnthetubaguy jaypipes Where does ironic (or ftm any virt driver) set up its resource providers today?
16:53:59 jaypipes efried: it doesn't. the RT does.
16:54:06 jaypipes efried: because there's only one RP.
16:54:32 mriedem stvnoyes: looking at http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/console.html#_2017-09-26_14_28_03_894125
16:54:37 efried jaypipes In ironic, only one RP per compute host? Not one RP per ironic node?
16:54:41 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:50 mriedem tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok
16:54:55 mriedem oops
16:55:02 jaypipes efried: we want to get to where there is a sort of mediation/negotiation between the RT and the virt driver and the generic device manager as each of them "processes" the ProviderTree and sets inventory records and traits based on their own information.
16:55:03 mriedem tempest.api.compute.admin.test_live_migration.LiveMigrationTest.test_live_block_migration [10.052345s] ... FAILED
16:55:17 mriedem so it works with the 2.25 microversion, where we pass block_migration=auto
16:55:21 mriedem but fails before that
16:55:22 jaypipes efried: in ironic, each baremetal Ironic node is its own RP in placementy.
16:55:48 mriedem so i bet something with the auto changes broke how we do live block migration between n-1 hosts

Earlier   Later