| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-28 | |||
| 16:03:37 | efried | Though the name conflict bug thing is a tangent-to-a-tangent... | |
| 16:03:52 | johnthetubaguy | so for ironic, resource class matching, and requesting no VCPU,MEM, etc, is the current way forward I thought? | |
| 16:04:18 | cfriesen | johnthetubaguy: this spec? https://specs.openstack.org/openstack/ironic-specs/specs/not-implemented/node-resource-class.html | |
| 16:05:21 | cfriesen | or maybe this one is more accurate: https://blueprints.launchpad.net/nova/+spec/custom-resource-classes-pike | |
| 16:05:40 | johnthetubaguy | the later might be closer, but yeah, that support has all merged | |
| 16:05:53 | cdent | cfriesen: sorry, in yet another meeting, so lost track of the discussion, will come back soon | |
| 16:06:48 | cfriesen | cdent: no worries, just idly curious | |
| 16:06:55 | johnthetubaguy | cfriesen: both actually are done I think | |
| 16:07:10 | dansmith | cdent: either your new job comes with lots of extra meetings, or you enjoy and announce them more often | |
| 16:08:13 | cdent | dansmith: these are all upstream meetings, I guess I’m just conscious lately of the extent to which they are distracting me from chatting with you | |
| 16:08:23 | dansmith | cdent: ah :) | |
| 16:08:28 | efried | johnthetubaguy iiuc, you want to get rid of tracking CPU, memory, etc.; tag each of your nodes with some identifier; and then have your flavor just use that identifier? | |
| 16:09:01 | johnthetubaguy | efried: so I thought the CPU memory, etc tracking was all going in queens, regardless? | |
| 16:09:18 | efried | "going into queens" or "going away in queens" ? | |
| 16:09:25 | johnthetubaguy | going away in queens | |
| 16:09:37 | johnthetubaguy | oops, missed the important word there | |
| 16:09:47 | cfriesen | efried: there's an interesting short email chain here: https://openstack.nimeyo.com/91493/openstack-dev-ironic-nova-indivisible-resource-providers | |
| 16:10:26 | efried | johnthetubaguy I assume that's just an ironic statement | |
| 16:10:45 | efried | I honestly know nothing about ironic, learning as I go in the context of this chat. | |
| 16:10:52 | johnthetubaguy | that's what the comment in the Nova code was saying, based on plan A for Nova I believe? | |
| 16:11:25 | johnthetubaguy | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L758 | |
| 16:11:50 | johnthetubaguy | ironic resources are indivisable | |
| 16:12:01 | johnthetubaguy | custom resource classes model that nicely | |
| 16:12:46 | efried | Right, so the point will be to have a custom resource class representing a "kind" of node, such that any node associated with that resource class is effectively "identical" to any other. | |
| 16:13:14 | johnthetubaguy | well, "identical" from the point of view of the flavor | |
| 16:13:30 | johnthetubaguy | you might have several generations mapped to a single resource class, if you want | |
| 16:13:32 | efried | Right. And you're modeling the compute process as the resource *provider* of the resources of the various classes it can see. | |
| 16:13:50 | efried | (What do we call the thing the compute process runs on? The "hypervisor"?) | |
| 16:14:05 | johnthetubaguy | no, the resource provider is the ironic node uuid in Ironic | |
| 16:14:36 | efried | Where "node" isn't the baremetal machine, it's the thing that manages 'em, right? | |
| 16:14:42 | johnthetubaguy | nova-compute process manages multiple ironic nodes (that list varies depending on the hash ring, hence the bug) | |
| 16:14:52 | johnthetubaguy | so the host is nova-compute | |
| 16:14:55 | efried | ayee | |
| 16:15:03 | efried | Terminology fail. | |
| 16:15:03 | johnthetubaguy | the node is the ironic node | |
| 16:15:21 | johnthetubaguy | the node is the resource provider | |
| 16:15:32 | cdent | nova-compute hosts multiple ironic nodes, each of which have an inventory of customr resource class with count of 1 | |
| 16:15:36 | cdent | (my memory is starting to refresh) | |
| 16:15:43 | johnthetubaguy | +1 cdent | |
| 16:16:37 | efried | But some of those ironic nodes can be using the same resource class, if they're functionally "identical" | |
| 16:16:38 | efried | ? | |
| 16:16:47 | cdent | yes | |
| 16:17:04 | efried | So it would seem like, as part of the failover process, the host that's taking over ought to delete the RP for the failed host. | |
| 16:17:13 | johnthetubaguy | yeah, each is offering 1: CUSTOM_GOLD | |
| 16:17:26 | efried | Seems like that ought to happen regardless, otherwise the scheduler will think it's still there and may try to schedule stuff to it. | |
| 16:17:35 | johnthetubaguy | so for that bug | |
| 16:17:40 | efried | yeah, for that bug | |
| 16:17:46 | johnthetubaguy | the dead thing would normally be responsible for the delete of it | |
| 16:17:59 | efried | That... doesn't make sense. It's dead. It can't call the placement API. | |
| 16:18:10 | johnthetubaguy | right, that's the problem | |
| 16:18:25 | efried | At some level, the new guy must know he's taking over for the old guy. | |
| 16:18:39 | efried | So why can't he delete the old guy's RP? | |
| 16:18:49 | johnthetubaguy | not sure if it sees that edge | |
| 16:18:59 | johnthetubaguy | it just sees a new node right now | |
| 16:19:09 | johnthetubaguy | but not 100% sure on that | |
| 16:19:10 | efried | Well. | |
| 16:19:23 | efried | Is there a split-brain possibility here? | |
| 16:19:44 | cfriesen | johnthetubaguy: for virtual machines (like with a vmware backend) nova-compute runs on the "node", and the hypervisor runs on the "host", right? why did ironic reverse the naming? | |
| 16:21:20 | dansmith | cfriesen: no, the reverse | |
| 16:21:30 | dansmith | cfriesen: but vmware doesn't actually run like that | |
| 16:21:34 | dansmith | (anymore) | |
| 16:21:41 | dansmith | well, | |
| 16:22:04 | dansmith | it doesn't expose a 1:N host:node structure I mean | |
| 16:22:05 | dansmith | it still runs external I think, but... | |
| 16:23:14 | dansmith | cfriesen: nova-compute always runs on the "host" which is where the service record is (nova-compute being a service) and the "node" is the actual hypervisor | |
| 16:23:36 | dansmith | ironic runs one nova-compute for multiple "hypervisor" nodes, which can't by definition run the service on them | |
| 16:24:05 | dansmith | ironic is the _reason_ we have that (blasted) model so it's setting the standard, not doing something reverse of what the model was designed for | |
| 16:27:45 | cfriesen | dansmith: got it, thanks. wonder where I got flipped around. | |
| 16:28:12 | dansmith | cfriesen: it's unfortunately very convoluted for little gain :( | |
| 16:30:42 | cfriesen | so how does "compute_node = self._get_compute_info(context, self.host)" make any sense then, if there could be multiple compute nodes per host? | |
| 16:30:59 | dansmith | cfriesen: is that in the resource tracker? | |
| 16:31:44 | cfriesen | nova/compute/manager.py | |
| 16:31:56 | dansmith | cfriesen: oh, look at the function | |
| 16:31:58 | dansmith | cfriesen: it returns the first | |
| 16:32:06 | dansmith | "for old compat" | |
| 16:32:35 | dansmith | and since ironic doesn't support live migration...YET | |
| 16:38:14 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: set/check if inventory change in tree https://review.openstack.org/470575 | |
| 16:38:15 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: integrate ProviderTree to report client https://review.openstack.org/415921 | |
| 16:38:15 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: add nested resource providers https://review.openstack.org/377138 | |
| 16:38:16 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: allow filter providers in tree https://review.openstack.org/377215 | |
| 16:38:16 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: adds REST API for nested providers https://review.openstack.org/384807 | |
| 16:38:17 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: update client to set parent provider https://review.openstack.org/385693 | |
| 16:43:04 | mriedem | stvnoyes: the multinode live migration grenade job failed on your live migration new style attach patch, but it failed here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/subnode-2/screen-n-cpu.txt.gz?level=TRACE#_Sep_26_14_28_11_637958 | |
| 16:43:25 | mriedem | and looking at logstash, that happens on a ton of patches, so it's probably just a 100% failure in that job right now | |
| 16:43:42 | dansmith | jaypipes: I'm missing the reasoning for the complicated "compare two dicts" method in that first patch | |
| 16:44:38 | jaypipes | dansmith: yeah, it's a long conversation on previous revisions... I'm going to change it to only compare the total field value and ignore the other things like reserved/min_unit/allocation_ratio, etc | |
| 16:45:14 | dansmith | jaypipes: oh, sorry I can re-read.. but, is this not to be used for "know when I need to report new inventory" sort of thing? | |
| 16:45:18 | jaypipes | dansmith: the reason it's like that is because we still don't have any consistency and agreement on who (virt driver or resource tracker) owns various fields about the inventory like alloc ratio etc | |
| 16:45:28 | efried | dansmith The history is here: https://review.openstack.org/#/c/470575/2/nova/compute/provider_tree.py@98 | |
| 16:46:08 | dansmith | jaypipes: ah, to avoid us reporting new just because the virt driver thinks the allocation_ratio should be different? | |
| 16:46:21 | jaypipes | yup | |
| 16:47:05 | dansmith | hmm | |
| 16:47:16 | dansmith | the virt driver needs to report total and reserved | |
| 16:47:42 | dansmith | but allocation_ratio is clearly the purview of the compute manager | |
| 16:47:56 | dansmith | min_unit is probably virt driver I guess | |
| 16:48:15 | dansmith | taking the scaleio example of 8GB minimum slice or whatever | |
| 16:51:06 | mriedem | stvnoyes: yeah the grenade live migration job is just totally busted http://tinyurl.com/y9ca6zox | |
| 16:51:53 | mriedem | something changed in august http://tinyurl.com/yby7m5p8 | |
| 16:52:27 | stvnoyes | that's pretty bad | |
| 16:53:45 | efried | johnthetubaguy jaypipes Where does ironic (or ftm any virt driver) set up its resource providers today? | |