Earlier  
Posted Nick Remark
#openstack-nova - 2017-10-03
14:20:12 jaypipes RP do not apply to the parent (ancestor) RPs"
14:20:12 jaypipes bauzas: so, your confusion stems from the sentence "However, traits defined on a child
14:20:22 bauzas exactly
14:20:29 jaypipes bauzas: what he's saying is correct, but weird :)
14:20:30 bauzas it doesn't bubble up
14:20:48 bauzas so I just want to make sure we walk in the tree
14:20:57 efried jaypipes bauzas I take that to mean: if the parent RP has *inventory*, you wouldn't get that *inventory*
14:21:13 efried It would be weird (but not forbidden) for a parent to have inventory in the same resource class as its child.
14:21:21 jaypipes bauzas: the trait itself doesn't apply to the parent specifically, but when looking for providers to return in allocation candidates, any traits for any provider in a *tree* will be considered as being "traits of the tree"
14:21:22 efried That's the case where the traits don't propagate upwards.
14:22:49 bauzas well, I'd say I would easily imagine the same resources:FOO:9&required=MISC_FOO not getting the node if only the child is having 8 FOOs per child
14:22:56 mriedem stephenfin: this has been around for a month without any patches https://bugs.launchpad.net/nova/+bug/1714017
14:22:57 openstack Launchpad bug 1714017 in OpenStack Compute (nova) "User guide was not migrated to the nova repo" [Medium,In progress] - Assigned to Pavlukhin Max (mpavlukhin)
14:23:06 jaypipes bauzas: that is correct.
14:23:06 mriedem stephenfin: so if you're keen maybe you want to take that over,
14:23:15 jaypipes bauzas: quantitatively it's not possible.
14:23:16 mriedem stephenfin: keeping in mind we need to backport whatever the changes are
14:23:29 mriedem so maybe step 1 is just import the user guide docs and backport, then step 2 is re-arrange and clean them up
14:23:36 mriedem since some of the config drive user guide docs are really for admins
14:23:44 bauzas jaypipes: but I had the same concern without traits, say I have 2 children with each of them having 8 FOOs
14:23:46 jaypipes bauzas: but remember that the max_unit/min_unit part of hte inventory records will solve that query problem.
14:23:57 stephenfin mriedem: I was going to, but Takashi NATSUME (I don't know his IRC nick) might be on it
14:24:12 bauzas jaypipes: if I'm asking for 9 FOOs, I wouldn't get them even if I could have 8 in a child and 1 in the other child
14:24:14 stephenfin see https://bugs.launchpad.net/nova/+bug/1720873
14:24:15 openstack Launchpad bug 1714017 in OpenStack Compute (nova) "duplicate for #1720873 User guide was not migrated to the nova repo" [Medium,In progress] - Assigned to Pavlukhin Max (mpavlukhin)
14:24:38 bauzas jaypipes: mmm, I don't see how that helps but okay :)
14:24:48 jaypipes bauzas: if you request 9 foo and there are 2 child providers having 8 foo inventories, each of those inventory records would have max_unit = 8 and that would mean a failing WHERE condition on max_unit <= $requested_amount
14:25:11 bauzas jaypipes: I certainly understand *why* it's failing
14:25:32 bauzas jaypipes: but I can easily imagine operators asking for spreading their resources
14:25:34 mriedem stephenfin: ok
14:25:34 efried yeah, that's a bug.
14:25:50 stephenfin If they haven't done anything by next week, I'll pick it up
14:26:14 bauzas snap, I need to disappear for my daily daddy work
14:26:24 jaypipes bauzas: sorry, I'm not following you still. :(
14:26:26 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Move 'ips' field from Subnet object to VIF object https://review.openstack.org/508498
14:27:21 jaypipes efried: do you think bauzas is talking about the resources1=, required1= stuff we discussed at the PTG?
14:28:34 stephenfin jaypipes: Fancy taking a look at this some time today? https://review.openstack.org/#/c/502056/
14:28:39 stephenfin (it's docs)
14:29:31 efried jaypipes A side effect thereof, yes.
14:30:14 efried jaypipes bauzas The two scenarios starting on L83 here: https://etherpad.openstack.org/p/nova-multi-alloc-request-syntax-brainstorm
14:30:33 efried ...as far as spreading resources across RPs.
14:32:45 jaypipes efried: ok. I don't disagree that any of those scenarios are important. just that it's kind of a distraction at this point and an acknowledged thing we'll need to look at at a later time
14:33:52 efried that allocation candidate.<<"
14:33:52 efried For the other discussion - how do traits "propagate" - perhaps it should be phrased: "With nested resource providers, traits defined on a parent RP are assumed to belong to all its child (descendant) RPs >>for purposes of inventory allocation<<. Traits defined on a child RP do not apply to the parent (ancestor) RPs. >>Although inventory is only claimed from a RP having the requested traits, the entire RP tree is returned in the provider summary for
14:34:00 efried or something like that.
14:35:48 cdent 🤐
14:36:36 efried cdent Glad you're staying out of it.
14:37:30 efried bauzas jaypipes We may want to defer this discussion; replace that chunk with "how traits are handled in nested RPs is out of scope and will be discussed in a subsequent spec". But that spec needs to land in Queens.
14:37:47 sdague melwitt: https://review.openstack.org/#/c/486947/ is a quotas spec that would be great to have your commentary on
14:38:06 efried jaypipes The "numbered resources querystring" spec would be a reasonable place for that discussion.
14:39:00 mriedem dansmith: you should probably go through this https://review.openstack.org/#/c/433603/
14:39:32 dansmith mriedem: ugh, do I have to?
14:39:53 mriedem yes
14:39:59 mriedem and finish your peas
14:40:30 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/nova-specs master: Network bandwitdh resource provider https://review.openstack.org/502306
14:42:59 efried jaypipes How does the scheduler determine that a particular RP is a compute host for the purposes of landing an instance?
14:43:26 efried (Yes, this is ultimately relevant to the discussion)
14:45:11 jaypipes efried: right now, the scheduler has a mapping of hostname to compute node UUIDs. the hostname is the same as the nova-compute service RPC topic queue
14:45:33 dansmith mriedem: gawd moooom
14:45:55 jaypipes efried: https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L640
14:46:14 efried jaypipes Okay, cool. Cause when we have more than just compute host RPs, that's going to be pretty crucial.
14:46:26 jaypipes efried: we already do...
14:46:49 jaypipes efried: Ironic nodes and shared storage providers are not compute nodes.
14:47:22 efried jaypipes Nested, then. Where that ties in is, you'll get your inventory from the child RP, but the whole tree is going to be (explicitly or implicitly via root_provider_id) part of your provider summaries. That's going to have to be how the scheduler figures out which compute node it should go to.
14:48:04 mriedem cdent: L105 here https://review.openstack.org/#/c/506552/4/specs/queens/approved/allow-update-instance-keypair.rst
14:48:10 mriedem PUT /servers/{server_id/
14:48:14 efried jaypipes Cause at some point, there's gonna be a model where the compute node RP actually doesn't have *any* resources (e.g. VCPU and MEMORY_MB belong to NUMA node RPs under the compute host; DISK_GB belongs to a shared RP somewhere; etc.)
14:48:25 mriedem if the server is not found by id, it's a 404, but if something in the request body isn't found, then it's a 400, right?
14:48:37 cdent mriedem: correct, that’s the general rule
14:48:46 jaypipes efried: the allocation_requests part of the GET /allocation_candidates HTTP response has the exact allocations against the exact resource providers that the instance will consume from (compute node providers and otherwise)
14:48:53 cdent if the uri fails to hit, 404, otherwise 400
14:50:01 efried jaypipes Right, but see above (..:48:16). The compute host RP may not actually be in the list of exact RPs the instance will consume from).
14:50:26 mriedem sdague: i see you found the competing instance keypair update specs
14:50:32 mriedem and i know you saw the ML thread
14:50:52 mriedem what i don't know is what most deployments do about cloud-init behavior
14:51:03 mriedem if they refresh on reboot or only on new build
14:51:04 jaypipes efried: the compute node provider will always be the root_provider_id, though.
14:51:39 efried jaypipes Right. That's my point. Generically, the scheduler will have to find the compute host by backtracking to the root_provider_id.
14:51:44 jaypipes efried: and the provider_summaries part of the GET /allocation_candidates HTTP response informs the caller that a particular provider is a child of another.
14:52:05 jaypipes efried: yes, that is true.
14:52:11 jaypipes efried: and expected.
14:52:26 efried jaypipes Right. That's another question I've had, though: will provider summaries include the whole tree, or only the "exact resource providers that the instance will consume from"?
14:52:33 jaypipes efried: the placement doesn't know or care that a particular provider represents a compute node.
14:52:49 efried right - the scheduler has to figure that out, I get that.
14:53:01 efried hence my original leading question.
14:53:02 jaypipes efried: placement doesn't care about anything other than inventory and allocation records...
14:53:16 jaypipes efried: so it's the scheduler's responsibility to understand those things.
14:53:24 efried Yup.
14:53:52 jaypipes efried: currently, the scheduler does that by keeping that map of service RPC to compute node UUID in memory. It will need to be adapted to understand nested/root providers
14:54:12 efried jaypipes Will provider summaries include the whole tree, or only the exact RPs being consumed from?
14:54:38 jaypipes efried: the exact RPs that could be consumed from and their parents.
14:54:50 efried parents/ancestors
14:54:50 jaypipes efried: which is needed for callers to piece the tree together.
14:54:53 jaypipes yes
14:54:54 efried up to the root
14:54:57 efried okay.
14:54:57 jaypipes yes
14:55:16 efried I dig that. Saves an extra call from the scheduler.
14:55:28 jaypipes efried: I wasn't planning on making provider_summaries into a tree structure, though. was relying on the caller to do that as needed.
14:56:04 efried I don't see the need for the whole tree - except that's easier to build, cause you'll already have that code for that API (forget which one) that returns a whole tree.

Earlier   Later