Earlier  
Posted Nick Remark
#openstack-nova - 2017-10-18
16:16:17 mriedem only sticking point right now is this means you can only specify one set of resource classes/traits which are spread across providers, right?
16:16:25 mriedem because all numbered groups are applied to the same RP
16:16:30 efried mriedem Correct.
16:16:48 efried mriedem Why would you need more than... oh.
16:16:54 mriedem i'm trying to think if that screws us later
16:17:09 efried mriedem Actually, no, I think that's aight.
16:17:23 mriedem i'm mostly worried about the shared storage scenario,
16:17:27 mriedem which i know we're punting, but
16:17:34 dansmith wait, what?
16:17:42 dansmith the point of this is you specifying traits/classes that go together
16:17:45 efried mriedem Well, the spec deliberately leaves out shared RP semantics.
16:18:10 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/nova-specs master: Network bandwidth resource provider https://review.openstack.org/502306
16:18:20 efried mriedem But we do have a couple of pretty big holes even without this, when it comes to traits + aggregates.
16:18:24 efried I'm writing one up now :)
16:19:05 dansmith ah, I see
16:19:06 mriedem so with shared storage, i think we wanted the scheduler to ask for vcpu/memory_mb and disk_gb, and we could get 2 providers back,
16:19:18 mriedem one compute node RP for the vcpu/memory_mb, and one shared storage provider for the disk_gb
16:19:22 mriedem scheduler would claim on both
16:19:37 dansmith with this,
16:19:38 mriedem maybe that still works ok with this new thing
16:19:47 mriedem you request those in an unnumbered resources group
16:19:52 mriedem so the resources can be spread across providers
16:19:53 dansmith you could ask for resources2:DISK=1,required2=SHARED
16:20:06 dansmith to make sure you get a shared disk, right?
16:20:37 mriedem in previous thinking about shared storage i didn't think the flavor had to be specific as to the type of storage
16:20:57 dansmith it doesn't have to be
16:21:21 mriedem so like i said, i think we don't lose the ability to do what we can do today with an unnumbered request,
16:21:30 mriedem you just can't have more than one unnumbered request
16:21:38 mriedem but i can't think of any reason why that might be bad
16:21:52 efried Correct. Except "what we can do today with an unnumbered request" is poorly defined and broken.
16:22:08 dansmith um
16:22:25 dansmith is there some special behavior assigned to resources= that isn't there for resourcesN=?
16:22:31 dansmith I must have missed that if so
16:22:35 mriedem also, now that i'm thinking about this, what is preventing placementing from returning 2 compute node providers, one that satisfies the vcpu requirement and one that satisfies the memory_mb requirement, scheduler claiming on both and the build failing?
16:22:39 efried Sorry, by "today" I actually mean "not today at all". I mean when both traits and aggregates are considered.
16:23:03 mriedem dansmith: my comment here https://review.openstack.org/#/c/510244/6/specs/queens/approved/granular-resource-requests.rst@230
16:23:51 dansmith mriedem: ah, I see, I had read that first bullet differently I guess
16:24:03 dansmith I read this as "if you only have one grouping"
16:24:54 efried "same tree" is the key there, mriedem. That's how you don't get VCPU and MEMORY_MB from separate computes.
16:25:40 mriedem ok so compute node today is a tree of 1 node
16:25:45 mriedem the alpha and omega
16:25:52 efried correct
16:26:01 mriedem whew :)
16:26:09 efried I would think we want to enforce that one allocation can only ever get resources from one tree plus zero or more shared-via-aggregate
16:26:27 mriedem yeah i think so
16:26:31 mriedem otherwise it gets crazy
16:26:50 efried For the unnumbered group, the resources from one class are always from the same RP; but resources from different classes can be spread throughout the tree + aggregates
16:27:08 mriedem i guess we never answered the question from the other day about whether or not a compute node provider can report both local disk_gb and disk_gb via a shared-with-aggregate storage pool
16:27:12 efried For the numbered group, all resources are always from the same RP (one node within a tree)... but I don't know how aggregates come into play there.
16:27:48 efried mriedem Just so. That's the bug I'm writing up now. As currently implemented, we ignore the aggregate if the compute node has DISK_GB.
16:28:03 efried and I don't think that's the right answer, generally/long-term.
16:28:30 mriedem ah, yeah, i'd think we'd pull from the aggregate inventory
16:28:45 mriedem since you're assuming that's a pool the operator wants you to use
16:28:48 efried well, we should return candidates for both.
16:28:54 mriedem maybe that becomes a traits tihng
16:29:10 efried mriedem we're busted there too
16:29:28 efried Because my compute RP can have e.g. RAID trait, and my shared RP can have e.g. SSD trait.
16:29:40 efried But placement has no way to know that those should stick together.
16:29:56 efried So if I ask for RAID+SSD, I'll get candidates, but I shouldn't.
16:30:29 efried that was in fact the bug I was trying to express (by writing a test for it) when I discovered the previous.
16:30:54 mriedem dansmith: so no major issues with you for this unnumbered resource limitation thing?
16:31:06 mriedem as noted, it's no worse than what we have today
16:31:13 dansmith mriedem: no, I didn't have it in my head, but Idon't think it changes anything
16:31:13 dansmith right
16:31:21 cfriesen mriedem: efried: why would the compute node provider report the shared storage amounts? shouldn't it just report that it has access to a particular shared storage?
16:31:37 dansmith I don't like the asymmetry, but I think we probably need to keep it to avoid breakage in the meantime
16:31:38 efried cfriesen Nono, the compute node *has* local storage.
16:31:49 dansmith it's really the resourcesN being different that I hadn't grokked
16:32:28 efried cfriesen So it's got a local disk, and it's also attached to a SAN or whatever. The former is reported in the compute node RP; the latter via the shared RP.
16:32:41 dansmith what efried said
16:32:51 dansmith that's not possible today, but should be eventually
16:33:17 cfriesen efired: yeah, that makes sense. mriedem's comment made it sound like the compute node RP was reporting on sizes of available shared storage
16:34:21 mriedem don't compute nodes that are getting storage from NFS today report disk_gb for the entire NFS cluster?
16:34:33 cfriesen mreidem: yeah, and that's been a bug for a long time I think.
16:34:44 mriedem so it makes it look like you have potentially 100 computes with 1TB of storage each, but not really
16:34:56 efried https://bugs.launchpad.net/nova/+bug/1724613 < there's the first one.
16:34:57 openstack Launchpad bug 1724613 in OpenStack Compute (nova) "AllocationCandidates.get_by_filters ignores shared RPs when the RC exists in both places" [Undecided,New]
16:35:26 cfriesen incidentally, who updates the stats for the shared storage RP? is there an auditor somewhere?
16:35:32 mriedem yeah so i wasn't sure how we are goign to fix the NFS disk reporting thing with shared providers, because the compute service is reporting that disk_gb
16:35:41 mriedem cfriesen: was supposed to be external to nova
16:35:47 mriedem like how neutron reports IP allocation pools
16:36:05 efried mriedem Multiple compute services report the same inventory, but to the same RP, because the shared RP has a UUID that they can all agree on.
16:36:25 efried I think that's what jaypipes update_inventory_if_needed thingy is for (or whatever it's called)
16:36:27 mriedem efried: i don't think that's quite accurate
16:36:48 mriedem each compute node get_inventory is going to report what it thinks it's local disk is,
16:36:52 mriedem but it can't tell if it's shared or not
16:37:08 efried We talking Q or later?
16:37:14 mriedem this is why we have the 'is_shared_storage' ssh stuff during migration
16:37:33 mriedem well, shared storage is not Q, so later,
16:37:39 mriedem but this is why it's not Q, among other reasons
16:38:11 efried Yeah, so each compute node happily reports all the storage, but the conductor (or whatever is doing the rollup) can see that those inventories are coming from the same RP, so it can report the total just once rather than adding it up.
16:38:31 dansmith efried: mriedem right, computes will need to not report shared storage
16:38:33 efried But wait, the compute node shouldn't be reporting inventory in the compute node RP for storage that's in an aggregate.
16:38:34 dansmith not just report the same
16:38:44 efried yeah, that %
16:38:45 dansmith computes will report their storage if they have some, else none
16:38:45 openstackgerrit Merged openstack/nova-specs master: Granular Resource Request Syntax https://review.openstack.org/510244
16:38:48 mriedem efried: the omputes aren't aware of hte aggregate
16:39:32 efried dansmith So who's responsible for creating the shared storage RP and its inventory?
16:39:39 dansmith efried: some storage agent

Earlier   Later