| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-18 | |||
| 16:20:06 | dansmith | to make sure you get a shared disk, right? | |
| 16:20:37 | mriedem | in previous thinking about shared storage i didn't think the flavor had to be specific as to the type of storage | |
| 16:20:57 | dansmith | it doesn't have to be | |
| 16:21:21 | mriedem | so like i said, i think we don't lose the ability to do what we can do today with an unnumbered request, | |
| 16:21:30 | mriedem | you just can't have more than one unnumbered request | |
| 16:21:38 | mriedem | but i can't think of any reason why that might be bad | |
| 16:21:52 | efried | Correct. Except "what we can do today with an unnumbered request" is poorly defined and broken. | |
| 16:22:08 | dansmith | um | |
| 16:22:25 | dansmith | is there some special behavior assigned to resources= that isn't there for resourcesN=? | |
| 16:22:31 | dansmith | I must have missed that if so | |
| 16:22:35 | mriedem | also, now that i'm thinking about this, what is preventing placementing from returning 2 compute node providers, one that satisfies the vcpu requirement and one that satisfies the memory_mb requirement, scheduler claiming on both and the build failing? | |
| 16:22:39 | efried | Sorry, by "today" I actually mean "not today at all". I mean when both traits and aggregates are considered. | |
| 16:23:03 | mriedem | dansmith: my comment here https://review.openstack.org/#/c/510244/6/specs/queens/approved/granular-resource-requests.rst@230 | |
| 16:23:51 | dansmith | mriedem: ah, I see, I had read that first bullet differently I guess | |
| 16:24:03 | dansmith | I read this as "if you only have one grouping" | |
| 16:24:54 | efried | "same tree" is the key there, mriedem. That's how you don't get VCPU and MEMORY_MB from separate computes. | |
| 16:25:40 | mriedem | ok so compute node today is a tree of 1 node | |
| 16:25:45 | mriedem | the alpha and omega | |
| 16:25:52 | efried | correct | |
| 16:26:01 | mriedem | whew :) | |
| 16:26:09 | efried | I would think we want to enforce that one allocation can only ever get resources from one tree plus zero or more shared-via-aggregate | |
| 16:26:27 | mriedem | yeah i think so | |
| 16:26:31 | mriedem | otherwise it gets crazy | |
| 16:26:50 | efried | For the unnumbered group, the resources from one class are always from the same RP; but resources from different classes can be spread throughout the tree + aggregates | |
| 16:27:08 | mriedem | i guess we never answered the question from the other day about whether or not a compute node provider can report both local disk_gb and disk_gb via a shared-with-aggregate storage pool | |
| 16:27:12 | efried | For the numbered group, all resources are always from the same RP (one node within a tree)... but I don't know how aggregates come into play there. | |
| 16:27:48 | efried | mriedem Just so. That's the bug I'm writing up now. As currently implemented, we ignore the aggregate if the compute node has DISK_GB. | |
| 16:28:03 | efried | and I don't think that's the right answer, generally/long-term. | |
| 16:28:30 | mriedem | ah, yeah, i'd think we'd pull from the aggregate inventory | |
| 16:28:45 | mriedem | since you're assuming that's a pool the operator wants you to use | |
| 16:28:48 | efried | well, we should return candidates for both. | |
| 16:28:54 | mriedem | maybe that becomes a traits tihng | |
| 16:29:10 | efried | mriedem we're busted there too | |
| 16:29:28 | efried | Because my compute RP can have e.g. RAID trait, and my shared RP can have e.g. SSD trait. | |
| 16:29:40 | efried | But placement has no way to know that those should stick together. | |
| 16:29:56 | efried | So if I ask for RAID+SSD, I'll get candidates, but I shouldn't. | |
| 16:30:29 | efried | that was in fact the bug I was trying to express (by writing a test for it) when I discovered the previous. | |
| 16:30:54 | mriedem | dansmith: so no major issues with you for this unnumbered resource limitation thing? | |
| 16:31:06 | mriedem | as noted, it's no worse than what we have today | |
| 16:31:13 | dansmith | right | |
| 16:31:13 | dansmith | mriedem: no, I didn't have it in my head, but Idon't think it changes anything | |
| 16:31:21 | cfriesen | mriedem: efried: why would the compute node provider report the shared storage amounts? shouldn't it just report that it has access to a particular shared storage? | |
| 16:31:37 | dansmith | I don't like the asymmetry, but I think we probably need to keep it to avoid breakage in the meantime | |
| 16:31:38 | efried | cfriesen Nono, the compute node *has* local storage. | |
| 16:31:49 | dansmith | it's really the resourcesN being different that I hadn't grokked | |
| 16:32:28 | efried | cfriesen So it's got a local disk, and it's also attached to a SAN or whatever. The former is reported in the compute node RP; the latter via the shared RP. | |
| 16:32:41 | dansmith | what efried said | |
| 16:32:51 | dansmith | that's not possible today, but should be eventually | |
| 16:33:17 | cfriesen | efired: yeah, that makes sense. mriedem's comment made it sound like the compute node RP was reporting on sizes of available shared storage | |
| 16:34:21 | mriedem | don't compute nodes that are getting storage from NFS today report disk_gb for the entire NFS cluster? | |
| 16:34:33 | cfriesen | mreidem: yeah, and that's been a bug for a long time I think. | |
| 16:34:44 | mriedem | so it makes it look like you have potentially 100 computes with 1TB of storage each, but not really | |
| 16:34:56 | efried | https://bugs.launchpad.net/nova/+bug/1724613 < there's the first one. | |
| 16:34:57 | openstack | Launchpad bug 1724613 in OpenStack Compute (nova) "AllocationCandidates.get_by_filters ignores shared RPs when the RC exists in both places" [Undecided,New] | |
| 16:35:26 | cfriesen | incidentally, who updates the stats for the shared storage RP? is there an auditor somewhere? | |
| 16:35:32 | mriedem | yeah so i wasn't sure how we are goign to fix the NFS disk reporting thing with shared providers, because the compute service is reporting that disk_gb | |
| 16:35:41 | mriedem | cfriesen: was supposed to be external to nova | |
| 16:35:47 | mriedem | like how neutron reports IP allocation pools | |
| 16:36:05 | efried | mriedem Multiple compute services report the same inventory, but to the same RP, because the shared RP has a UUID that they can all agree on. | |
| 16:36:25 | efried | I think that's what jaypipes update_inventory_if_needed thingy is for (or whatever it's called) | |
| 16:36:27 | mriedem | efried: i don't think that's quite accurate | |
| 16:36:48 | mriedem | each compute node get_inventory is going to report what it thinks it's local disk is, | |
| 16:36:52 | mriedem | but it can't tell if it's shared or not | |
| 16:37:08 | efried | We talking Q or later? | |
| 16:37:14 | mriedem | this is why we have the 'is_shared_storage' ssh stuff during migration | |
| 16:37:33 | mriedem | well, shared storage is not Q, so later, | |
| 16:37:39 | mriedem | but this is why it's not Q, among other reasons | |
| 16:38:11 | efried | Yeah, so each compute node happily reports all the storage, but the conductor (or whatever is doing the rollup) can see that those inventories are coming from the same RP, so it can report the total just once rather than adding it up. | |
| 16:38:31 | dansmith | efried: mriedem right, computes will need to not report shared storage | |
| 16:38:33 | efried | But wait, the compute node shouldn't be reporting inventory in the compute node RP for storage that's in an aggregate. | |
| 16:38:34 | dansmith | not just report the same | |
| 16:38:44 | efried | yeah, that % | |
| 16:38:45 | openstackgerrit | Merged openstack/nova-specs master: Granular Resource Request Syntax https://review.openstack.org/510244 | |
| 16:38:45 | dansmith | computes will report their storage if they have some, else none | |
| 16:38:48 | mriedem | efried: the omputes aren't aware of hte aggregate | |
| 16:39:32 | efried | dansmith So who's responsible for creating the shared storage RP and its inventory? | |
| 16:39:39 | dansmith | efried: some storage agent | |
| 16:39:42 | dansmith | like neutron does | |
| 16:39:46 | mriedem | thinking back on pike issues, this was also a thing where the scheduler would claim disk_gb on the shared storage RP, but the compute would overwrite the instance disk_gb allocation against it's local compute node | |
| 16:39:51 | mriedem | because it wasn't aware of the aggregate relationship for disk | |
| 16:39:51 | cdent | mriedem: the rt has an aggregate map (as yet unused) | |
| 16:40:00 | cdent | that was supposed to allow it to be able to report inventory correctly | |
| 16:40:06 | cdent | once shared exists | |
| 16:40:13 | mriedem | cdent: ah, fun | |
| 16:40:29 | cdent | “fun” | |
| 16:40:46 | mriedem | i remember working a patch for one afternoon late in pike rc time trying to sort out how to not get the rt to overwrite the shared disk allocation and it went down the hole fast | |
| 16:41:10 | mriedem | involved basically reverse engineering the logic in placement and the scheduler, from the rt | |
| 16:41:25 | cdent | wheeee! | |
| 16:41:34 | mriedem | but, another reason why we don't want the RT trying to figure out allocations | |
| 16:41:43 | cdent | the aggregate_map still leaves open the question of whether a compute node can have both local and shared, which, unsure | |
| 16:42:27 | mriedem | yeah idk, couldn't you mount an NFS share on a compute and configure the instance path to use that share, but leave root and everything else that's local disk for the OS and running nova-compute? | |
| 16:42:42 | mriedem | like, i want all my instance and image crap to go in the NFS share | |
| 16:42:52 | mriedem | leave local disk for everything else | |
| 16:43:03 | dansmith | there's lots of things that need to change on compute to make that duality possible | |
| 16:43:07 | dansmith | the easiest to do today would be local disk + ceph I think, | |
| 16:43:17 | dansmith | since it's not fighting over /var/lib/instances | |
| 16:44:04 | mriedem | ok well this is why no shared storage support in queens :) | |
| 16:44:08 | dansmith | cha | |
| 16:44:12 | mriedem | granular request syntax spec approved | |
| 16:44:19 | cdent | huzzah | |