| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-18 | |||
| 16:06:58 | efried | I'll use "compute node" for the sake of the bug, but eventually we should have a word. | |
| 16:07:01 | cdent | the other one is the “sharing” rp, so maybe the main one is the “non-sharing” (that’s a bit weak though) | |
| 16:09:01 | cfriesen | cdent: we could have multiple non-shared RPs per compute node, no? I think I remember hearing that some Intel folks were talking about creating a separate non-shared RP to provide L3 cache. | |
| 16:09:44 | cfriesen | if so then it's really the "nova compute node RP" or something like that. :) | |
| 16:09:44 | openstackgerrit | John Garbutt proposed openstack/nova master: Keep updating allocations for Ironic https://review.openstack.org/513085 | |
| 16:09:54 | johnthetubaguy | dtantsur: that is my first stab at it ^ | |
| 16:11:28 | cdent | cfriesen: non-sharing-root-provider is more complete then | |
| 16:11:41 | cdent | was trying to leave that complexity out for mow | |
| 16:12:37 | cfriesen | cdent: although in the case of L3 cache it would be properly modelled as a per-NUMA-node resource and would need nested providers. :) | |
| 16:13:17 | cdent | let’s hope so | |
| 16:15:08 | edleafe | "greedy rp"? (i.e., not sharing) | |
| 16:16:00 | mriedem | efried: dansmith: i went through https://review.openstack.org/#/c/510244/ for the granular allocation candidates request syntax, | |
| 16:16:17 | mriedem | only sticking point right now is this means you can only specify one set of resource classes/traits which are spread across providers, right? | |
| 16:16:25 | mriedem | because all numbered groups are applied to the same RP | |
| 16:16:30 | efried | mriedem Correct. | |
| 16:16:48 | efried | mriedem Why would you need more than... oh. | |
| 16:16:54 | mriedem | i'm trying to think if that screws us later | |
| 16:17:09 | efried | mriedem Actually, no, I think that's aight. | |
| 16:17:23 | mriedem | i'm mostly worried about the shared storage scenario, | |
| 16:17:27 | mriedem | which i know we're punting, but | |
| 16:17:34 | dansmith | wait, what? | |
| 16:17:42 | dansmith | the point of this is you specifying traits/classes that go together | |
| 16:17:45 | efried | mriedem Well, the spec deliberately leaves out shared RP semantics. | |
| 16:18:10 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/nova-specs master: Network bandwidth resource provider https://review.openstack.org/502306 | |
| 16:18:20 | efried | mriedem But we do have a couple of pretty big holes even without this, when it comes to traits + aggregates. | |
| 16:18:24 | efried | I'm writing one up now :) | |
| 16:19:05 | dansmith | ah, I see | |
| 16:19:06 | mriedem | so with shared storage, i think we wanted the scheduler to ask for vcpu/memory_mb and disk_gb, and we could get 2 providers back, | |
| 16:19:18 | mriedem | one compute node RP for the vcpu/memory_mb, and one shared storage provider for the disk_gb | |
| 16:19:22 | mriedem | scheduler would claim on both | |
| 16:19:37 | dansmith | with this, | |
| 16:19:38 | mriedem | maybe that still works ok with this new thing | |
| 16:19:47 | mriedem | you request those in an unnumbered resources group | |
| 16:19:52 | mriedem | so the resources can be spread across providers | |
| 16:19:53 | dansmith | you could ask for resources2:DISK=1,required2=SHARED | |
| 16:20:06 | dansmith | to make sure you get a shared disk, right? | |
| 16:20:37 | mriedem | in previous thinking about shared storage i didn't think the flavor had to be specific as to the type of storage | |
| 16:20:57 | dansmith | it doesn't have to be | |
| 16:21:21 | mriedem | so like i said, i think we don't lose the ability to do what we can do today with an unnumbered request, | |
| 16:21:30 | mriedem | you just can't have more than one unnumbered request | |
| 16:21:38 | mriedem | but i can't think of any reason why that might be bad | |
| 16:21:52 | efried | Correct. Except "what we can do today with an unnumbered request" is poorly defined and broken. | |
| 16:22:08 | dansmith | um | |
| 16:22:25 | dansmith | is there some special behavior assigned to resources= that isn't there for resourcesN=? | |
| 16:22:31 | dansmith | I must have missed that if so | |
| 16:22:35 | mriedem | also, now that i'm thinking about this, what is preventing placementing from returning 2 compute node providers, one that satisfies the vcpu requirement and one that satisfies the memory_mb requirement, scheduler claiming on both and the build failing? | |
| 16:22:39 | efried | Sorry, by "today" I actually mean "not today at all". I mean when both traits and aggregates are considered. | |
| 16:23:03 | mriedem | dansmith: my comment here https://review.openstack.org/#/c/510244/6/specs/queens/approved/granular-resource-requests.rst@230 | |
| 16:23:51 | dansmith | mriedem: ah, I see, I had read that first bullet differently I guess | |
| 16:24:03 | dansmith | I read this as "if you only have one grouping" | |
| 16:24:54 | efried | "same tree" is the key there, mriedem. That's how you don't get VCPU and MEMORY_MB from separate computes. | |
| 16:25:40 | mriedem | ok so compute node today is a tree of 1 node | |
| 16:25:45 | mriedem | the alpha and omega | |
| 16:25:52 | efried | correct | |
| 16:26:01 | mriedem | whew :) | |
| 16:26:09 | efried | I would think we want to enforce that one allocation can only ever get resources from one tree plus zero or more shared-via-aggregate | |
| 16:26:27 | mriedem | yeah i think so | |
| 16:26:31 | mriedem | otherwise it gets crazy | |
| 16:26:50 | efried | For the unnumbered group, the resources from one class are always from the same RP; but resources from different classes can be spread throughout the tree + aggregates | |
| 16:27:08 | mriedem | i guess we never answered the question from the other day about whether or not a compute node provider can report both local disk_gb and disk_gb via a shared-with-aggregate storage pool | |
| 16:27:12 | efried | For the numbered group, all resources are always from the same RP (one node within a tree)... but I don't know how aggregates come into play there. | |
| 16:27:48 | efried | mriedem Just so. That's the bug I'm writing up now. As currently implemented, we ignore the aggregate if the compute node has DISK_GB. | |
| 16:28:03 | efried | and I don't think that's the right answer, generally/long-term. | |
| 16:28:30 | mriedem | ah, yeah, i'd think we'd pull from the aggregate inventory | |
| 16:28:45 | mriedem | since you're assuming that's a pool the operator wants you to use | |
| 16:28:48 | efried | well, we should return candidates for both. | |
| 16:28:54 | mriedem | maybe that becomes a traits tihng | |
| 16:29:10 | efried | mriedem we're busted there too | |
| 16:29:28 | efried | Because my compute RP can have e.g. RAID trait, and my shared RP can have e.g. SSD trait. | |
| 16:29:40 | efried | But placement has no way to know that those should stick together. | |
| 16:29:56 | efried | So if I ask for RAID+SSD, I'll get candidates, but I shouldn't. | |
| 16:30:29 | efried | that was in fact the bug I was trying to express (by writing a test for it) when I discovered the previous. | |
| 16:30:54 | mriedem | dansmith: so no major issues with you for this unnumbered resource limitation thing? | |
| 16:31:06 | mriedem | as noted, it's no worse than what we have today | |
| 16:31:13 | dansmith | mriedem: no, I didn't have it in my head, but Idon't think it changes anything | |
| 16:31:13 | dansmith | right | |
| 16:31:21 | cfriesen | mriedem: efried: why would the compute node provider report the shared storage amounts? shouldn't it just report that it has access to a particular shared storage? | |
| 16:31:37 | dansmith | I don't like the asymmetry, but I think we probably need to keep it to avoid breakage in the meantime | |
| 16:31:38 | efried | cfriesen Nono, the compute node *has* local storage. | |
| 16:31:49 | dansmith | it's really the resourcesN being different that I hadn't grokked | |
| 16:32:28 | efried | cfriesen So it's got a local disk, and it's also attached to a SAN or whatever. The former is reported in the compute node RP; the latter via the shared RP. | |
| 16:32:41 | dansmith | what efried said | |
| 16:32:51 | dansmith | that's not possible today, but should be eventually | |
| 16:33:17 | cfriesen | efired: yeah, that makes sense. mriedem's comment made it sound like the compute node RP was reporting on sizes of available shared storage | |
| 16:34:21 | mriedem | don't compute nodes that are getting storage from NFS today report disk_gb for the entire NFS cluster? | |
| 16:34:33 | cfriesen | mreidem: yeah, and that's been a bug for a long time I think. | |
| 16:34:44 | mriedem | so it makes it look like you have potentially 100 computes with 1TB of storage each, but not really | |
| 16:34:56 | efried | https://bugs.launchpad.net/nova/+bug/1724613 < there's the first one. | |
| 16:34:57 | openstack | Launchpad bug 1724613 in OpenStack Compute (nova) "AllocationCandidates.get_by_filters ignores shared RPs when the RC exists in both places" [Undecided,New] | |
| 16:35:26 | cfriesen | incidentally, who updates the stats for the shared storage RP? is there an auditor somewhere? | |
| 16:35:32 | mriedem | yeah so i wasn't sure how we are goign to fix the NFS disk reporting thing with shared providers, because the compute service is reporting that disk_gb | |
| 16:35:41 | mriedem | cfriesen: was supposed to be external to nova | |
| 16:35:47 | mriedem | like how neutron reports IP allocation pools | |
| 16:36:05 | efried | mriedem Multiple compute services report the same inventory, but to the same RP, because the shared RP has a UUID that they can all agree on. | |
| 16:36:25 | efried | I think that's what jaypipes update_inventory_if_needed thingy is for (or whatever it's called) | |
| 16:36:27 | mriedem | efried: i don't think that's quite accurate | |
| 16:36:48 | mriedem | each compute node get_inventory is going to report what it thinks it's local disk is, | |
| 16:36:52 | mriedem | but it can't tell if it's shared or not | |
| 16:37:08 | efried | We talking Q or later? | |
| 16:37:14 | mriedem | this is why we have the 'is_shared_storage' ssh stuff during migration | |