| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-04-12 | |||
| 13:10:39 | stephenfin | bauzas: typically those guest nodes would be scheduled to different NUMA nodes, but if that's not possible they'll be squeezed onto the same one | |
| 13:11:00 | stephenfin | bauzas: At least that's the case to the best of my recollection | |
| 13:11:00 | bauzas | efried: that's correct | |
| 13:11:09 | efried | ahjeez. "anti-affinity preferred" is a use case we haven't even thought about yet. | |
| 13:11:14 | bauzas | efried: trying to get a consensus for Rocky-1 or Rocky-2 | |
| 13:11:25 | bauzas | efried: while I'm implementing other things | |
| 13:11:35 | efried | bauzas: roger that | |
| 13:11:42 | bauzas | efried: but once I'm done with the other things, going back to NUMA for implementing | |
| 13:12:04 | bauzas | honestly, if we don't need new Placement version, I think it can be done very quickly | |
| 13:12:07 | bauzas | efried: ^ | |
| 13:12:21 | efried | yes | |
| 13:12:25 | bauzas | efried: it's just about translating specs into Placement queries | |
| 13:12:32 | efried | yes | |
| 13:12:48 | bauzas | so, at least the NUMAFilter would still check the hosts, but it would only check the accepted ones | |
| 13:13:10 | efried | And I actually like the idea of translating the existing numa-related flavor specs into placement queries, rather than asking folks to rewrite their flavors with placement-y syntax. | |
| 13:13:16 | bauzas | and given the default flag value for NUMA will be None, nothing should change actually | |
| 13:13:40 | bauzas | unless someone wants to test that magic bullet to restrict hosts passed to the scheduler | |
| 13:13:43 | efried | Because that way we can make placement queries that would be... unreasonable for humans to come up with. | |
| 13:14:04 | bauzas | it's more about an upgrade question for me | |
| 13:14:19 | bauzas | efried: modifying flavors could be a problem for upgrading | |
| 13:14:29 | bauzas | if we support the existing flavors, that's better | |
| 13:14:36 | efried | ++ | |
| 13:15:04 | bauzas | stephenfin: so, to clarify, you mean that we don't explicitely shard between NUMA nodes ? | |
| 13:15:43 | stephenfin | bauzas: We do but it's best effort | |
| 13:16:09 | bauzas | stephenfin: tbc, if you're asking for hw:numa_nodes=2 but you only have one NUMA node (or even a UMA topology), then we accept the host, right?N | |
| 13:16:19 | stephenfin | bauzas: yup | |
| 13:16:25 | bauzas | perfect, efried ^ | |
| 13:17:00 | efried | Cool beans. One down, three to go. | |
| 13:17:07 | bauzas | stephenfin: so, it's more about NUMA "affinity" of multiple resources, rather than NUMA 'anti-affinity' of different CPUs | |
| 13:17:19 | bauzas | still right? | |
| 13:17:35 | bauzas | if so, numbered request groups is the perfect expression | |
| 13:17:49 | stephenfin | It's not really anything to do with NUMA affinity, tbh. It's purely to do with the guest topology | |
| 13:17:54 | bauzas | perfect | |
| 13:18:16 | bauzas | it doesn't guaranttee you'll land all your CPUs on specific NUMA nodes | |
| 13:18:23 | stephenfin | The fact that we don't split a guest NUMA node across a host NUMA node is a performance improvement but not essential by any means | |
| 13:18:26 | bauzas | you could have all your CPUs on the same | |
| 13:18:33 | stephenfin | correct | |
| 13:18:41 | bauzas | excellent, good news | |
| 13:18:49 | bauzas | and tbh, I understand the reasoning | |
| 13:19:01 | stephenfin | so booting an instance with hw:numa_nodes=4 on a dual socket system is a valid thing to do | |
| 13:19:02 | bauzas | you could discover the host topology as an end-user if we were restricting | |
| 13:19:12 | stephenfin | assuming total vCPUs < total host CPUs | |
| 13:19:20 | stephenfin | 'zactly | |
| 13:19:24 | bauzas | just by trying to boot a couple of different topologies | |
| 13:19:27 | stephenfin | There's also zero reason to allow it | |
| 13:19:53 | bauzas | efried: so you agree the numbered request groups feature exactly matches the above ? ^ | |
| 13:20:57 | sahid | bauzas, stephenfin, if user is asking for hw:numa_node=2 and host does not have at least 2 NUMA node we don't accept the host | |
| 13:21:08 | sahid | or it's a bug and we should fix it | |
| 13:21:27 | stephenfin | sahid: Why? | |
| 13:21:36 | bauzas | sahid: that's very different from what stephenfin says | |
| 13:21:48 | bauzas | http://www.quickmeme.com/meme/361gwd | |
| 13:22:00 | sahid | because it's all performance related | |
| 13:22:22 | sahid | it's all about distance between pci devices, memory channel an cpus | |
| 13:24:07 | sahid | stephenfin: we should have in hardware.py constrainsts something which check that | |
| 13:24:14 | stephenfin | bauzas: Ah, crap. sahid's correct there | |
| 13:24:42 | stephenfin | bauzas, sahid: https://github.com/openstack/nova/blob/master/nova/virt/hardware.py#L1548-L1553 | |
| 13:25:35 | bauzas | sahid: so, like I said, say I'm an evil folk named Eve, I could guess the host topology by booting a couple of test instances having various NUMA guest topologies, right? | |
| 13:25:58 | stephenfin | bauzas: and https://github.com/openstack/nova/blob/master/nova/objects/instance_numa_topology.py#L217-L219 | |
| 13:26:23 | bauzas | I understand the performance reason for affinitizing the guest, but not for anti-affinitizing it | |
| 13:26:27 | stephenfin | I thought that was comparing the length of cell.cpuset, a la https://github.com/openstack/nova/blob/master/nova/objects/instance_numa_topology.py#L82-L83 | |
| 13:26:39 | stephenfin | sahid: Yeah, why do we do that? It seems unnecessary | |
| 13:27:34 | stephenfin | Like I said above, I get the reason for not splitting guest NUMA nodes across host NUMA nodes | |
| 13:27:37 | stephenfin | and for placing at least one guest NUMA node's CPUs on the host NUMA node associated with a PCI device | |
| 13:27:52 | bauzas | stephenfin: checking whether the host has the same number of NUMA nodes that the guest is one thing | |
| 13:28:10 | bauzas | stephenfin: the other thing being that we would restrict that to different nodes | |
| 13:28:18 | sahid | what is the use-cases? I mean why you want to have differente NUMA nodes if the CPU is using a memory channel which is on an different NUMA node? as I said it's all performance related | |
| 13:28:59 | stephenfin | sahid: Right, but I have two NUMA nodes and I boot the instance with four NUMA nodes. Why can't two go on one host node and two on the other? | |
| 13:29:03 | bauzas | sahid: performance is affinity, and I don't disagree with you | |
| 13:29:24 | stephenfin | Sounds like you'd get the same performance as with an instance with two NUMA nodes that's split across the two host nodes | |
| 13:29:25 | sahid | stephenfin: why you want do that? | |
| 13:29:34 | stephenfin | sahid: So I can boot my instance | |
| 13:29:53 | bauzas | sahid: in theory, the user doesn't know the host topology | |
| 13:30:00 | bauzas | sahid: he's just booting a flavor | |
| 13:30:17 | bauzas | but what he knows is that he'll get a guest having its own topology | |
| 13:30:20 | stephenfin | Yeah, what bauzas said | |
| 13:30:20 | openstackgerrit | Matthew Booth proposed openstack/nova master: Rename recreate to evacuate in driver signatures https://review.openstack.org/560900 | |
| 13:30:33 | sahid | yes and if he wants 2 numa where don"t want to fake that | |
| 13:30:39 | openstackgerrit | Eric Fried proposed openstack/os-traits master: normalize_name helper https://review.openstack.org/560107 | |
| 13:30:43 | stephenfin | He wants two _guest_ NUMA nodes | |
| 13:30:50 | sahid | yes and if he wants 2 numa nodes for it's guest we don"t want to fake that | |
| 13:31:00 | stephenfin | The entire thing is fake | |
| 13:31:14 | bauzas | sahid: so some evil guy could guess there are no left hosts having 2 numa nodes | |
| 13:31:20 | stephenfin | There aren't actually two guest NUMA nodes. That's just QEMU/KVM mocking it | |
| 13:32:07 | sahid | stephenfin: it's not... the guest memory and vcpus threads are running on a uniq node | |
| 13:32:35 | stephenfin | sahid: I get that. I'm not suggesting splitting a guest's NUMA node across a host NUMA node | |
| 13:32:56 | stephenfin | What I'm saying is there's not reason we shouldn't fit N guest NUMA nodes on the same host NUMA node | |
| 13:32:57 | efried | but the other way around is okay | |
| 13:33:09 | stephenfin | efried: Yeah, precisely | |
| 13:33:15 | sahid | stephenfin: it's not what we do | |
| 13:33:29 | bauzas | so the big thing is | |
| 13:33:33 | bauzas | now we know that | |
| 13:33:38 | bauzas | should Placement do that ? | |
| 13:33:41 | bauzas | I don't think so | |
| 13:33:45 | stephenfin | sahid: Indeed, and I think we can change that | |
| 13:33:47 | bauzas | leave that logic to the filter | |
| 13:33:53 | stephenfin | bauzas: No, neither do I | |
| 13:34:01 | bauzas | what Placement would return is just a list of potential targets | |
| 13:34:12 | stephenfin | So long as we can still ensure at least one guest NUMA node is affined with a given host NUMA node | |
| 13:34:17 | bauzas | some can run that workload but on a single node, some can do that better | |
| 13:34:19 | stephenfin | For things like PCI devices and the likes | |