| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-02-04 | |||
| 14:47:30 | sean-k-mooney | ya i agree | |
| 14:47:34 | efried | Right, that's what I thought we were going to do. If FooHost doesn't ever care about NUMA instances, it can just leave that flag off forever, and never have a flavor with hw:numa* in it, and go on happily forever. | |
| 14:47:48 | ralonsoh | stephenfin, yes, but I'm reviewing the code and if the extension is not enabled, this information won't be there, in the FIP | |
| 14:47:53 | sean-k-mooney | i dont really like basing the behavior off what the current vms are using | |
| 14:48:03 | efried | UNTIL we can design & implement proper granular anti-affinity syntax in placement, which IMO is too ambitious for Ussuri. | |
| 14:48:03 | efried | gibi, bauzas: To close on group_policy: sean-k-mooney and I talked about it a bit yesterday and decided that, in order to support both NUMA and bandwidth, we need to retain the group_policy=none default and do the anti-affinity in the NTF | |
| 14:48:09 | efried | sean-k-mooney: +1 | |
| 14:48:19 | sean-k-mooney | efried: keep in mind that almost all host are numa hosts | |
| 14:48:32 | dansmith | sean-k-mooney: in reality, ALL of them are | |
| 14:48:38 | efried | meaning almost all hosts are *capable* of NUMA. | |
| 14:48:45 | efried | Not that all hosts are hosting NUMA-aware instances. | |
| 14:48:49 | sean-k-mooney | dansmith: yes | |
| 14:48:59 | dansmith | not capable, they are numa and if not configured, are losing performance to that fact | |
| 14:49:00 | sean-k-mooney | efried: correct | |
| 14:49:12 | efried | right, which some VMs don't care about. | |
| 14:49:13 | stephenfin | ralonsoh: WDYM? As I understood it, we won't have the 'port_details' field but the 'device_id' field will still be there in the ports API response, right? | |
| 14:49:41 | dansmith | efried: no, they all care about it, just some operators don't care to take on the burden of configuring it because we make it so difficult | |
| 14:49:41 | sean-k-mooney | efried: no i think dansmit ment at the host level. not the vm level | |
| 14:49:48 | ralonsoh | stephenfin, no no | |
| 14:50:00 | dansmith | nobody is opting out of numa for any reason other than the cost benefit isn't there for configuring it | |
| 14:50:02 | ralonsoh | stephenfin, if you don't have 'port_details', you won't have this info | |
| 14:50:14 | stephenfin | ohhhh, they're tied together | |
| 14:50:21 | stephenfin | ? | |
| 14:50:23 | ralonsoh | stephenfin, if you really need the host but you don't have 'port_details' | |
| 14:50:33 | efried | dansmith: oh? Not because opting out lets them squeeze more VMs onto their host? | |
| 14:50:35 | ralonsoh | then you have the "port_id" always | |
| 14:50:55 | sean-k-mooney | efried: well it depends | |
| 14:50:57 | ralonsoh | stephenfin, and then, you can call the Neutron API to retrieve the port info (and the host) | |
| 14:51:04 | sean-k-mooney | in general no | |
| 14:51:08 | efried | yes, it depends, that's what I'm saying. | |
| 14:51:20 | sean-k-mooney | but if you use pinning or hugepage that disable oversubsription or cpus or memroy | |
| 14:51:23 | dansmith | efried: only because complexity of what we provide makes that increasingly difficult as density increases | |
| 14:51:36 | sean-k-mooney | if you jsut use numa over subsription is allowed | |
| 14:51:43 | efried | right, and that's not a fitting problem we're going to solve any time soon. | |
| 14:52:20 | bauzas | I feel we need to define a clear upgrade path | |
| 14:52:33 | stephenfin | ralonsoh: I'm confused. You seem to be saying the same thing as me :) Let me try again | |
| 14:52:35 | bauzas | 1/ for NUMA-aware instances | |
| 14:52:42 | bauzas | 2/ for non-NUMA-aware instances | |
| 14:52:46 | efried | 2/ don't | |
| 14:52:46 | efried | 1/ flip the switch | |
| 14:53:01 | efried | If we've already decided we're going to segregate, it really is that simple. | |
| 14:53:08 | sean-k-mooney | so the boolean config option allows us to punt on the final decission on this for a cycle or two. | |
| 14:53:09 | dansmith | is there some reason that the switch can't be flipped for everyone? | |
| 14:53:23 | sean-k-mooney | e.g. default to on | |
| 14:53:27 | sean-k-mooney | and you opt out? | |
| 14:53:44 | dansmith | having compute nodes behave like one thing or another isn't something I'd like to see us doing long-term | |
| 14:53:49 | bauzas | we could also have an implicit switch | |
| 14:54:13 | bauzas | like, if you ask for mempages, then basically you ask for NUMA things | |
| 14:54:28 | bauzas | (even if, and I hate to say, is unrelated) | |
| 14:54:32 | bauzas | or cpu_pin | |
| 14:54:36 | efried | dansmith: the reason we didn't want to do that is because, if your priority is "land my VM", that becomes difficult/impossible as your cloud reaches saturation. | |
| 14:54:36 | bauzas | or whatever | |
| 14:54:50 | sean-k-mooney | bauzas: those are properties fo teh guest | |
| 14:54:54 | sean-k-mooney | bauzas: not of the host | |
| 14:54:54 | dansmith | efried: that is specifically called out in our project scope as not a problem nova solves | |
| 14:55:06 | stephenfin | ralonsoh: I can do 'GET /floatingips' (or whatever the API is). Each item in the response may contain a 'port_details' field but only if this extension is enabled. Correct? | |
| 14:55:07 | dansmith | efried: i.e. fitting the last VM into available memory | |
| 14:55:10 | bauzas | sean-k-mooney: for cpu pinset, it's for host | |
| 14:55:23 | bauzas | but I agree with large pages | |
| 14:55:46 | ralonsoh | stephenfin, exactly, the key 'port_details' will be there only if the extension is enabled | |
| 14:55:46 | bauzas | dammit, I'm torn | |
| 14:55:49 | sean-k-mooney | ya so we could enable it by default if you have defiend cpu_dedicated_set | |
| 14:56:05 | efried | bauzas: that's backwards though. If you ask the scheduler for large pages, you're asking to land on a host that knows how to do that. You can't use that question to decide to make a host large page-aware. | |
| 14:56:17 | sean-k-mooney | if cpu_dedicated_set is defiend then you are reporting PCPU and the only instance that can consume PCPUs have a numa toplogy | |
| 14:56:51 | sean-k-mooney | so that is once case where yes we can implcitly enable the numa reporting | |
| 14:57:36 | efried | dansmith: So how did we get to a point where we decided it was important to segregate the cloud? Hate to drag you into a second simultaneous discussion stephenfin, but weren't you part of that? | |
| 14:57:37 | stephenfin | ralonsoh: Okay. So if I do 'GET /ports/{port_id}', will that response always contain a 'device_id' field? If not, is this because it's provided by the same extension? | |
| 14:57:40 | sean-k-mooney | having hugepages allocated on the host is not really a good enough reason in my book and we should not assume that all numa hosts use cpu pinning | |
| 14:57:54 | bauzas | sean-k-mooney: right, that's my point | |
| 14:58:09 | bauzas | you could only care about large pages or just standard NUMA sharding | |
| 14:58:20 | dansmith | efried: segregate what? numa and non-numa instances? | |
| 14:58:31 | efried | yes | |
| 14:58:31 | sean-k-mooney | bauzas: well you could be useing the hugepages on the host for a dpdk vswitch for example | |
| 14:58:36 | bauzas | like, "I want 2 vCPUs on the same NUMA node" doesn't absolutely require CPU pinning | |
| 14:58:36 | sean-k-mooney | it might not be for the vms | |
| 14:58:38 | stephenfin | efried: Because we can't make everything have NUMA | |
| 14:58:50 | efried | stephenfin: right, so dansmith wants to know why not | |
| 14:58:50 | ralonsoh | stephenfin, exactly https://github.com/openstack/neutron/blob/master/neutron/db/db_base_plugin_common.py#L231 | |
| 14:58:56 | sean-k-mooney | bauzas: right that just need hw:numa_nodes=1 | |
| 14:59:02 | ralonsoh | stephenfin, the device_id will be there | |
| 14:59:20 | sean-k-mooney | bauzas: i use hugepages on my home system but not pinning | |
| 14:59:36 | stephenfin | efried: 2 sockets w/ a 32 core CPU in each socket (no HyperThreading). Go boot a 33 core instance | |
| 14:59:40 | bauzas | either way, I think we need to move forward | |
| 14:59:45 | sean-k-mooney | bauzas: beacue i want cpu oversubscition but not memory over subscrtion | |
| 15:00:08 | bauzas | I'll write the new revision with a bool flag and mention the alternative of an automatic all-NUMA world in the spec | |
| 15:00:09 | stephenfin | You can't because the instance will no longer split across the NUMA nodes, and we don't let an instance oversubscribe against itself | |
| 15:00:17 | bauzas | people will chime in and we'll see | |
| 15:00:24 | efried | dansmith: --^ | |
| 15:00:39 | dansmith | stephenfin: efried: because what nova currently provides is "no numa means 1 numa" yeah? | |
| 15:00:44 | sean-k-mooney | efried: yes we are aware that is the giant vm case. | |
| 15:00:50 | stephenfin | ralonsoh: It will *always* be there? | |
| 15:00:57 | stephenfin | dansmith: No, no NUMA means no NUMA | |
| 15:00:58 | ralonsoh | stephenfin, yes | |
| 15:00:59 | stephenfin | currently | |
| 15:01:18 | bauzas | dansmith: nova currently provides "I can spread my VM across many NUMA nodes if I don't care' | |
| 15:01:20 | dansmith | stephenfin: how is no numa and 1 numa node different to theguest/ | |
| 15:01:23 | ralonsoh | stephenfin, in the port this key, "device_id", is mandatory | |
| 15:01:33 | sean-k-mooney | dansmith: in the guest it does not | |
| 15:01:35 | dansmith | bauzas: right, we lie and say it's one node when it's not you mean? | |
| 15:01:38 | dansmith | sean-k-mooney: right exactly | |