| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-03-09 | |||
| 15:51:58 | leakypipes | mriedem: no allocations. | |
| 15:52:02 | leakypipes | mriedem: no aggregates. | |
| 15:52:09 | leakypipes | mriedem: only inventory and traits. | |
| 15:52:10 | mriedem | so just what the RT would report today | |
| 15:52:15 | leakypipes | correct. | |
| 15:52:26 | mriedem | and then the RT will overwrite those, b/c that's what it does today | |
| 15:52:58 | leakypipes | mriedem: I have only submitted a spec describing the behaviour of a single override -- the resource provider's UUID -- in the resource tracker. | |
| 15:53:13 | mriedem | leakypipes: the spec says it can also be inventory and traits | |
| 15:53:31 | leakypipes | mriedem: I haven't discussed overriding inventory or traits in the resource tracker yet. I've only commented on a possible format for describing those inventory/traits overrides. | |
| 15:53:38 | mriedem | https://review.openstack.org/#/c/551315/1/specs/rocky/approved/override-compute-node-uuid.rst@84 | |
| 15:54:01 | mriedem | then ^ should probably come out | |
| 15:54:07 | leakypipes | mriedem: why? | |
| 15:54:29 | mriedem | we need to be clear about what happens if an operator is going to put inventory and trait information in this yaml file, | |
| 15:54:32 | superdan | leakypipes: if we just let them explicitly set the name on the provider, then can't we avoid providing a generic override mechanism so that other services can use the name to sync up? | |
| 15:54:36 | mriedem | because today the RT is going to override that | |
| 15:54:54 | mriedem | if you plan to support out of band inventory/traits definitions, then the spec needs to discuss fixing the override so that nova merges those in | |
| 15:55:05 | leakypipes | mriedem: I'm happy to remove that from the use cases list. | |
| 15:55:28 | leakypipes | mriedem: I planned on another spec to discuss the implementation of inventory and trait overrides in the resource tracker. | |
| 15:55:37 | leakypipes | mriedem: that spec is just for the compute node UUID override. | |
| 15:56:39 | leakypipes | superdan: UUID or name, I don't really care. But there has to be a way of signaling to the resource tracker not to use the CONF.host value. | |
| 15:57:11 | superdan | leakypipes: or we just define a new conf option for the name of the RP, default=None means use CONF.host | |
| 15:57:19 | mriedem | ^ was just thinking that | |
| 15:57:25 | superdan | and then roll from there, without having to provide generic override mechanisms, hard coding uuids, etc | |
| 15:58:03 | leakypipes | superdan: and how does that signal to other agents running in other containers what the value should be that they look up? | |
| 15:58:36 | superdan | the same way as if this override yaml file is present and has a uuid set I guess? | |
| 15:59:26 | superdan | for the container file namespace issue, you can put that one option in its own conf file in a common location | |
| 15:59:39 | superdan | *filesystem namespace I mean | |
| 15:59:43 | leakypipes | superdan: how so? the agent may a) be managing multiple compute nodes and need information on multiple providers and b) needs some way of indicating what those compute node identifiers are. | |
| 16:00:10 | leakypipes | superdan: "one option in its own conf file" vs a standardized descriptor file format for resource providers? | |
| 16:00:25 | superdan | I'm confused. | |
| 16:00:31 | leakypipes | superdan: I would have thought you'd be supportive of not adding yet more configuration options. | |
| 16:00:51 | mriedem | leakypipes: this spec doesn't talk about non-nova agents managing multiple compute nodes | |
| 16:01:08 | openstackgerrit | Balazs Gibizer proposed openstack/nova-specs master: Network bandwidth resource provider https://review.openstack.org/502306 | |
| 16:01:23 | mriedem | i'm reading this all as 1:1 with a single compute node | |
| 16:02:21 | leakypipes | mriedem: I can add text about 1:M because that's what neutron agents do... many of them run on controller nodes and manage resources for hundreds of compute nodes. | |
| 16:03:46 | mriedem | sorry for my lack of neutron agent know-how, but that's not the case for OVS and LB is it? | |
| 16:03:55 | mriedem | otherwise how would os-vif plugins work for those? | |
| 16:04:03 | mriedem | since we don't use rpc or rest apis in os-vif | |
| 16:04:44 | mriedem | i'm not trying to be snarky, i just don't know the details on how neutron agents can all be deployed, and for which types of neutron backends | |
| 16:06:38 | leakypipes | mriedem: imagine a neutron agent that is managing network bandwidth resources for a rack of compute nodes. it needs to know the UUIDs (or compute node names) of those compute node resource providers so that it can add child providers to each compute node that represent the PFs that have a limited supply of ingress/egress bandwidth for physical networks. That agent could run on a single compute node or it could run on the TOR, or it could run | |
| 16:06:38 | leakypipes | in a container on a controller... | |
| 16:08:20 | mriedem | is that a thing today with the traditional OVS agent? or is this a neutron agent of the future? | |
| 16:08:32 | giblet | Kevin_Zheng, alex_xu_: I've updated the bandwidth spec based on the comments so far https://review.openstack.org/#/c/502306 | |
| 16:08:33 | mriedem | like, a new agent service for managing placement type resource information | |
| 16:10:55 | leakypipes | mriedem: I'm not entirely sure where you're going with these questions... I can remove the multi-provider stuff from the spec, but frankly, I don't see why we'd want to limit ourselves to agents that only work with a single resource provider. | |
| 16:11:49 | mriedem | leakypipes: i'm just trying to understand. i don't know if i'm just retarded and don't know how neutron agents work *today*, or if you're talking about something that could be built off of this spec for a new neutron agent in the future. but if it's a problem then i'll just stop asking questions. | |
| 16:11:58 | leakypipes | superdan: to answer your question above about "if we just let them explicitly set the name on the provider, then can't we avoid providing a generic override mechanism so that other services can use the name to sync up?", we will need to provide trait override information in the (near) future -- think about the whole forbidden traits and cpu_model stuff. Do you propose adding more CONF options for trait overrides as well? | |
| 16:12:34 | leakypipes | mriedem: I'm feeling ganged up on. | |
| 16:12:37 | mriedem | operators, or external services, can set traits | |
| 16:12:50 | mriedem | nova just needs to not overwrite them | |
| 16:13:07 | mriedem | i don't think we need config options for that | |
| 16:13:12 | leakypipes | mriedem: and how do we signal to nova not to override *some* traits but not others? | |
| 16:13:13 | mriedem | it's why we have the rest api | |
| 16:13:21 | mriedem | CUSTOM_? | |
| 16:13:48 | mriedem | i thought there was some discussion about that at the ptg about overriding standard vs custom traits | |
| 16:13:50 | mriedem | would have to look | |
| 16:13:58 | leakypipes | mriedem: that's problematic because the specific traits we need to prevent overriding are things like CPU features when CONF.cpu_model is set. | |
| 16:15:04 | mriedem | L494 https://etherpad.openstack.org/p/nova-ptg-rocky | |
| 16:15:44 | superdan | leakypipes: I'm not ganging up on you, I promise | |
| 16:16:13 | superdan | leakypipes: I think we can avoid overwriting traits easily without any config by having things like the virt driver assert which ones it is authoritative over and leaving others alone | |
| 16:16:36 | mriedem | wrt merging traits i'd be fine with just not touching any traits that aren't in the set that we (nova-compute) says should be on the provider | |
| 16:16:39 | mriedem | the compute node provider | |
| 16:16:46 | mriedem | standard or custom | |
| 16:17:15 | mriedem | but if n-cpu says trait X should be on the provider, and someone removed it, then we're going to end up putting it back on | |
| 16:17:27 | leakypipes | superdan: how are you proposing that the virt driver indicate which traits it is authoritative over? | |
| 16:17:38 | superdan | leakypipes: like I said in the room, | |
| 16:18:05 | superdan | some sort of contract between compute and virt, where virt returns all the traits and a "assert", "de-assert", "dont-care" sort of thing | |
| 16:18:19 | openstackgerrit | Balazs Gibizer proposed openstack/nova-specs master: Network bandwidth resource provider https://review.openstack.org/502306 | |
| 16:18:25 | superdan | compute looks at the traits currently on a provider, and doesn't touch any that are not mentioned or listed as dont-care | |
| 16:18:31 | superdan | same for any compute-manager-owned traits | |
| 16:18:34 | leakypipes | superdan: and for non-virt things? | |
| 16:18:38 | superdan | same | |
| 16:19:00 | mriedem | i'm not yet sure why we need that level of granularity | |
| 16:19:11 | leakypipes | superdan: this "contract" of which you speak... that is my proposed provider config file. | |
| 16:19:25 | superdan | leakypipes: no, it's not | |
| 16:19:29 | leakypipes | yes, it is :) | |
| 16:20:10 | mriedem | the difference is the file is defined out of band, while the virt driver is in control of what it says should be on the provider | |
| 16:20:36 | cdent | presumably the contract has to live _somewhere_, or is just just "known" somehow? | |
| 16:20:43 | mriedem | the file is a proxy iow, so that nova can report stuff that the operator setup, w/o the operator having to set the traits on the provider directly in the rest api | |
| 16:22:00 | mriedem | cdent: i'm not yet sure why we need a contract, why we don't just merge in the common sense of what merge means | |
| 16:22:27 | mriedem | RP has traits x,y,z, nova-compute says it should also have 1,2,3, so we merge those and the RP has 1,2,3,x,y,z | |
| 16:22:53 | mriedem | nova-compute gets its traits from the virt driver | |
| 16:23:07 | leakypipes | mriedem: the file is saying "for this resource provider, ignore these auto-discovered traits or always use these traits even if you didn't auto-discover them. | |
| 16:23:38 | mriedem | leakypipes: so the file would have the forbidden traits logic built in? | |
| 16:23:59 | leakypipes | mriedem: yes, that's what the always and ignore sections are. | |
| 16:24:00 | mriedem | like pci whitelist, but traits blacklist | |
| 16:24:47 | leakypipes | mriedem: though the forbidden traits thing is more about *requesting* a specific traits *not* be present on a provider. | |
| 16:24:56 | mriedem | idk, i can't say i like the proxy aspect of this when we have a rest api for placement | |
| 16:25:19 | leakypipes | mriedem: which is different from the provider config YAML file which is saying what always should be set or always should be ignored for a provider. | |
| 16:25:23 | mriedem | it would be one thing if the traits were only set in the compute_nodes table and we didn't have a REST API to manipulate that data externally | |
| 16:26:13 | mriedem | the ignore case is the CPU features thing when host_model is set? | |
| 16:26:29 | mriedem | the 'always set these traits' thing can be done externally from nova | |
| 16:26:41 | leakypipes | mriedem: we absolutely *do* have a REST API for manipulating things. the provider config YAML stuff is a proposal for how to solve two primary problems: 1) notifying external agents about the identifier for a provider it needs to know about and 2) resolving the "admin set this thing and agent overwrote it" problem. | |
| 16:27:14 | mriedem | leakypipes: i'm talking about different rest apis | |
| 16:27:19 | openstackgerrit | Merged openstack/nova master: Fix indentation in doc/source/cli/* https://review.openstack.org/549166 | |
| 16:27:36 | mriedem | my point is that since we do have a placement rest api to set traits on a provider, we do'nt need nova to proxy that information via a yaml file | |
| 16:28:05 | mriedem | i get (1) | |
| 16:28:13 | leakypipes | mriedem: and what happens with the compute node auto-discover traits stuff? | |
| 16:28:21 | mriedem | for (2) if we simply merge the traits set externally with the traits that nova-compute reports, then we should be fine | |
| 16:28:53 | leakypipes | mriedem: somebody is going to overwrite somebody else. the provider config file is a solution to instruct things not to overwrite (or to always overwrite) certain attributes. | |