| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-18 | |||
| 16:16:25 | mriedem | because all numbered groups are applied to the same RP | |
| 16:16:30 | efried | mriedem Correct. | |
| 16:16:48 | efried | mriedem Why would you need more than... oh. | |
| 16:16:54 | mriedem | i'm trying to think if that screws us later | |
| 16:17:09 | efried | mriedem Actually, no, I think that's aight. | |
| 16:17:23 | mriedem | i'm mostly worried about the shared storage scenario, | |
| 16:17:27 | mriedem | which i know we're punting, but | |
| 16:17:34 | dansmith | wait, what? | |
| 16:17:42 | dansmith | the point of this is you specifying traits/classes that go together | |
| 16:17:45 | efried | mriedem Well, the spec deliberately leaves out shared RP semantics. | |
| 16:18:10 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/nova-specs master: Network bandwidth resource provider https://review.openstack.org/502306 | |
| 16:18:20 | efried | mriedem But we do have a couple of pretty big holes even without this, when it comes to traits + aggregates. | |
| 16:18:24 | efried | I'm writing one up now :) | |
| 16:19:05 | dansmith | ah, I see | |
| 16:19:06 | mriedem | so with shared storage, i think we wanted the scheduler to ask for vcpu/memory_mb and disk_gb, and we could get 2 providers back, | |
| 16:19:18 | mriedem | one compute node RP for the vcpu/memory_mb, and one shared storage provider for the disk_gb | |
| 16:19:22 | mriedem | scheduler would claim on both | |
| 16:19:37 | dansmith | with this, | |
| 16:19:38 | mriedem | maybe that still works ok with this new thing | |
| 16:19:47 | mriedem | you request those in an unnumbered resources group | |
| 16:19:52 | mriedem | so the resources can be spread across providers | |
| 16:19:53 | dansmith | you could ask for resources2:DISK=1,required2=SHARED | |
| 16:20:06 | dansmith | to make sure you get a shared disk, right? | |
| 16:20:37 | mriedem | in previous thinking about shared storage i didn't think the flavor had to be specific as to the type of storage | |
| 16:20:57 | dansmith | it doesn't have to be | |
| 16:21:21 | mriedem | so like i said, i think we don't lose the ability to do what we can do today with an unnumbered request, | |
| 16:21:30 | mriedem | you just can't have more than one unnumbered request | |
| 16:21:38 | mriedem | but i can't think of any reason why that might be bad | |
| 16:21:52 | efried | Correct. Except "what we can do today with an unnumbered request" is poorly defined and broken. | |
| 16:22:08 | dansmith | um | |
| 16:22:25 | dansmith | is there some special behavior assigned to resources= that isn't there for resourcesN=? | |
| 16:22:31 | dansmith | I must have missed that if so | |
| 16:22:35 | mriedem | also, now that i'm thinking about this, what is preventing placementing from returning 2 compute node providers, one that satisfies the vcpu requirement and one that satisfies the memory_mb requirement, scheduler claiming on both and the build failing? | |
| 16:22:39 | efried | Sorry, by "today" I actually mean "not today at all". I mean when both traits and aggregates are considered. | |
| 16:23:03 | mriedem | dansmith: my comment here https://review.openstack.org/#/c/510244/6/specs/queens/approved/granular-resource-requests.rst@230 | |
| 16:23:51 | dansmith | mriedem: ah, I see, I had read that first bullet differently I guess | |
| 16:24:03 | dansmith | I read this as "if you only have one grouping" | |
| 16:24:54 | efried | "same tree" is the key there, mriedem. That's how you don't get VCPU and MEMORY_MB from separate computes. | |
| 16:25:40 | mriedem | ok so compute node today is a tree of 1 node | |
| 16:25:45 | mriedem | the alpha and omega | |
| 16:25:52 | efried | correct | |
| 16:26:01 | mriedem | whew :) | |
| 16:26:09 | efried | I would think we want to enforce that one allocation can only ever get resources from one tree plus zero or more shared-via-aggregate | |
| 16:26:27 | mriedem | yeah i think so | |
| 16:26:31 | mriedem | otherwise it gets crazy | |
| 16:26:50 | efried | For the unnumbered group, the resources from one class are always from the same RP; but resources from different classes can be spread throughout the tree + aggregates | |
| 16:27:08 | mriedem | i guess we never answered the question from the other day about whether or not a compute node provider can report both local disk_gb and disk_gb via a shared-with-aggregate storage pool | |
| 16:27:12 | efried | For the numbered group, all resources are always from the same RP (one node within a tree)... but I don't know how aggregates come into play there. | |
| 16:27:48 | efried | mriedem Just so. That's the bug I'm writing up now. As currently implemented, we ignore the aggregate if the compute node has DISK_GB. | |
| 16:28:03 | efried | and I don't think that's the right answer, generally/long-term. | |
| 16:28:30 | mriedem | ah, yeah, i'd think we'd pull from the aggregate inventory | |
| 16:28:45 | mriedem | since you're assuming that's a pool the operator wants you to use | |
| 16:28:48 | efried | well, we should return candidates for both. | |
| 16:28:54 | mriedem | maybe that becomes a traits tihng | |
| 16:29:10 | efried | mriedem we're busted there too | |
| 16:29:28 | efried | Because my compute RP can have e.g. RAID trait, and my shared RP can have e.g. SSD trait. | |
| 16:29:40 | efried | But placement has no way to know that those should stick together. | |
| 16:29:56 | efried | So if I ask for RAID+SSD, I'll get candidates, but I shouldn't. | |
| 16:30:29 | efried | that was in fact the bug I was trying to express (by writing a test for it) when I discovered the previous. | |
| 16:30:54 | mriedem | dansmith: so no major issues with you for this unnumbered resource limitation thing? | |
| 16:31:06 | mriedem | as noted, it's no worse than what we have today | |
| 16:31:13 | dansmith | mriedem: no, I didn't have it in my head, but Idon't think it changes anything | |
| 16:31:13 | dansmith | right | |
| 16:31:21 | cfriesen | mriedem: efried: why would the compute node provider report the shared storage amounts? shouldn't it just report that it has access to a particular shared storage? | |
| 16:31:37 | dansmith | I don't like the asymmetry, but I think we probably need to keep it to avoid breakage in the meantime | |
| 16:31:38 | efried | cfriesen Nono, the compute node *has* local storage. | |
| 16:31:49 | dansmith | it's really the resourcesN being different that I hadn't grokked | |
| 16:32:28 | efried | cfriesen So it's got a local disk, and it's also attached to a SAN or whatever. The former is reported in the compute node RP; the latter via the shared RP. | |
| 16:32:41 | dansmith | what efried said | |
| 16:32:51 | dansmith | that's not possible today, but should be eventually | |
| 16:33:17 | cfriesen | efired: yeah, that makes sense. mriedem's comment made it sound like the compute node RP was reporting on sizes of available shared storage | |
| 16:34:21 | mriedem | don't compute nodes that are getting storage from NFS today report disk_gb for the entire NFS cluster? | |
| 16:34:33 | cfriesen | mreidem: yeah, and that's been a bug for a long time I think. | |
| 16:34:44 | mriedem | so it makes it look like you have potentially 100 computes with 1TB of storage each, but not really | |
| 16:34:56 | efried | https://bugs.launchpad.net/nova/+bug/1724613 < there's the first one. | |
| 16:34:57 | openstack | Launchpad bug 1724613 in OpenStack Compute (nova) "AllocationCandidates.get_by_filters ignores shared RPs when the RC exists in both places" [Undecided,New] | |
| 16:35:26 | cfriesen | incidentally, who updates the stats for the shared storage RP? is there an auditor somewhere? | |
| 16:35:32 | mriedem | yeah so i wasn't sure how we are goign to fix the NFS disk reporting thing with shared providers, because the compute service is reporting that disk_gb | |
| 16:35:41 | mriedem | cfriesen: was supposed to be external to nova | |
| 16:35:47 | mriedem | like how neutron reports IP allocation pools | |
| 16:36:05 | efried | mriedem Multiple compute services report the same inventory, but to the same RP, because the shared RP has a UUID that they can all agree on. | |
| 16:36:25 | efried | I think that's what jaypipes update_inventory_if_needed thingy is for (or whatever it's called) | |
| 16:36:27 | mriedem | efried: i don't think that's quite accurate | |
| 16:36:48 | mriedem | each compute node get_inventory is going to report what it thinks it's local disk is, | |
| 16:36:52 | mriedem | but it can't tell if it's shared or not | |
| 16:37:08 | efried | We talking Q or later? | |
| 16:37:14 | mriedem | this is why we have the 'is_shared_storage' ssh stuff during migration | |
| 16:37:33 | mriedem | well, shared storage is not Q, so later, | |
| 16:37:39 | mriedem | but this is why it's not Q, among other reasons | |
| 16:38:11 | efried | Yeah, so each compute node happily reports all the storage, but the conductor (or whatever is doing the rollup) can see that those inventories are coming from the same RP, so it can report the total just once rather than adding it up. | |
| 16:38:31 | dansmith | efried: mriedem right, computes will need to not report shared storage | |
| 16:38:33 | efried | But wait, the compute node shouldn't be reporting inventory in the compute node RP for storage that's in an aggregate. | |
| 16:38:34 | dansmith | not just report the same | |
| 16:38:44 | efried | yeah, that % | |
| 16:38:45 | dansmith | computes will report their storage if they have some, else none | |
| 16:38:45 | openstackgerrit | Merged openstack/nova-specs master: Granular Resource Request Syntax https://review.openstack.org/510244 | |
| 16:38:48 | mriedem | efried: the omputes aren't aware of hte aggregate | |
| 16:39:32 | efried | dansmith So who's responsible for creating the shared storage RP and its inventory? | |
| 16:39:39 | dansmith | efried: some storage agent | |
| 16:39:42 | dansmith | like neutron does | |