| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-06 | |||
| 22:13:27 | edleafe | efried: according to the people who seemed to know more about it than I do, a machine with NUMA can be configured in a variety of ways. | |
| 22:14:05 | edleafe | efried: so if all the resources were provided by a plain compute node, wouldn't that also get returned? | |
| 22:14:21 | sean-k-mooney | edleafe: yes it they can. i know more about this topic then is good for my health | |
| 22:15:04 | sean-k-mooney | edleafe: if that compute node has only one memoy controler then by definiton all resocues on that node belong to the same numa node | |
| 22:15:07 | efried | edleafe: Okay, that's a fair point. Which we could resolve, as you say, by tagging the NUMA RPs with a NUMA trait. | |
| 22:15:20 | edleafe | ah good. So efried, just direct your questions to sean-k-mooney | |
| 22:15:32 | sean-k-mooney | so all compute nodes always have at least 1 numa node | |
| 22:16:29 | efried | Okay, so the only way to get non | |
| 22:16:47 | efried | non-NUMA-ness is to go to a system that has more than one NUMA node and get resources spread out across them. | |
| 22:17:05 | efried | In which case we *don't* need the trait. | |
| 22:17:11 | sean-k-mooney | efried: and if you incldue a distance to restict it | |
| 22:18:07 | sean-k-mooney | if you have multiple numa node and put no constraitn on the distance between resocues in the tree then you get resouce from different numa nodes | |
| 22:18:17 | sean-k-mooney | your performace will be less but it will still work | |
| 22:18:37 | sean-k-mooney | and you never asked for any afintiy so it fine | |
| 22:21:11 | efried | Okay, so this is a workable thing for NUMA affinity. I don't think the same mechanism can be made to work for affinity in general, or for anti-affinity. | |
| 22:21:55 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Move the nova-next job in-tree and update it https://review.openstack.org/541474 | |
| 22:22:03 | sean-k-mooney | efried: if you dont have to deal with aggreates then i think it will work not sure how distance is compute for aggregates however. | |
| 22:22:40 | efried | sean-k-mooney: I'm talking about using one aggregate per NUMA node (so nothing to do with distance). | |
| 22:23:03 | sean-k-mooney | when modeling afinity via distance affinty becoures deistnace < X and anti afinity is distance > X | |
| 22:23:48 | sean-k-mooney | efried: oh then ya if numa nodes are aggrages then that is different | |
| 22:23:50 | efried | Yeah, I get that, but *modeling* and *expressing* distance seems really complicated for people to understand. | |
| 22:24:10 | cfriesen | efried: one aggregate per numa node? would you then associate a pool of RAM, CPUs, PCI devices, etc, with that aggregate? | |
| 22:24:48 | cfriesen | at probably 2-4 numa nodes per compute node that seems like a lot of aggregates | |
| 22:24:55 | sean-k-mooney | you would make the cpus/ram/pci devices resouce providers members of the aggregate | |
| 22:25:03 | efried | cfriesen: You associate the resource *providers* with the aggregate. So whatever providers provide the resources associated with your NUMA node. | |
| 22:25:06 | efried | yeah, what sean said. | |
| 22:25:18 | efried | No, it would be one aggregate per NUMA node, period. | |
| 22:25:39 | sean-k-mooney | efried: not one aggregate per numa node per host? | |
| 22:25:51 | efried | sean-k-mooney: Yes. | |
| 22:26:02 | efried | One aggregate per NUMA node, in the global sense :) | |
| 22:26:24 | sean-k-mooney | oh ame that aggreate will be am large | |
| 22:26:35 | efried | Is that "a lot" of aggregates? Meh. Aggregates are cheap. | |
| 22:27:16 | efried | I feel like we're not talking in the same terms. Here: If you have 10 hosts, and each host has 2 NUMA nodes, you have 20 aggregates total. | |
| 22:27:18 | sean-k-mooney | i think i would prefer lots of small aggregate then a few giant cluster wide ones | |
| 22:27:58 | sean-k-mooney | efried: ah ok that what i was thinking we should have if we used aggregates so im cool with that | |
| 22:29:04 | sean-k-mooney | we dont have traits on aggregates but at least for numa i dont think we need them on teh aggreages so it could work | |
| 22:29:21 | cfriesen | so you'd have these per-numa-node aggregates for system use, then per-host aggregates to match against flavor/image properties and act as availability zones? | |
| 22:29:44 | efried | sean-k-mooney: We can always stamp the involved providers with traits. But not sure we need 'em. At least to mark NUMA nodes. | |
| 22:30:02 | efried | cfriesen: I know nothing about host aggregates or availability zones. | |
| 22:30:16 | sean-k-mooney | cfriesen: well an aggreaget is just a bag of resouce providers nothing more or less. you can use them to group anything you want | |
| 22:30:18 | efried | cfriesen: Except that placement aggregates != host aggregates | |
| 22:30:39 | efried | cfriesen: The other main use case for aggregates is to associate sharing resource providers, like a shared storage pool. | |
| 22:31:25 | efried | ^ placement aggregates, that is. | |
| 22:32:07 | cfriesen | efried: okay...I'm not up on the placement terminology. :) | |
| 22:32:35 | efried | cfriesen: I'm not up on... pretty much everything else. Between the two of us, we're gonna need a pidgin. | |
| 22:33:22 | sean-k-mooney | cfriesen: a sharing resouce provider is jsut a normal resouce provider with a trait and then its inventores are consumable by anything in the same aggreates as it | |
| 22:34:36 | cfriesen | so how do we handle the idea of an instance consuming cpu/ram from one placement aggregate but a PCI device from another placement aggregate? (for the non-strict PCI affinity case) | |
| 22:34:44 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Test websocketproxy with TLS in the nova-next job https://review.openstack.org/513160 | |
| 22:34:49 | mriedem | melwitt: cleaned that up ^ | |
| 22:35:28 | sean-k-mooney | cfriesen: instance consume resocues form invetories provided by resouce providers | |
| 22:35:40 | melwitt | mriedem: thanks. sorry I lost track of that one | |
| 22:35:47 | sean-k-mooney | cfriesen: and resouce provires can be member of 0-n aggregates | |
| 22:36:30 | sean-k-mooney | so you are not consuming form the aggregates | |
| 22:37:22 | efried | cfriesen: In the design I'm strawpersoning, if you care about CPU/mem affinity and you don't care about VF affinity, you would say affine_resources=VCPU,MEMORY_MB but not SRIOV_NET_VF | |
| 22:37:46 | efried | cfriesen: s/VF/PCI device/ if you wish. | |
| 22:39:08 | sean-k-mooney | efried: that may break down a bit with anti afinity | |
| 22:39:33 | efried | sean-k-mooney: I'm saying we're not handling anti-affinity at all with this model. | |
| 22:39:34 | sean-k-mooney | e.g. i want 2 VF whoes parent PF are anti affined | |
| 22:39:48 | sean-k-mooney | efried: ah ok | |
| 22:39:52 | efried | sean-k-mooney: Yeah, I get that. But that's not a NUMA use case, right? | |
| 22:40:01 | efried | It's more about HA | |
| 22:40:05 | efried | redundancy | |
| 22:40:15 | efried | powerfail grouping | |
| 22:40:18 | efried | that kind of thing. | |
| 22:40:21 | sean-k-mooney | ya its for ha bonding so my bond dont come from the same card | |
| 22:40:51 | efried | So yeah, this doesn't cover that. And if there's a generic solution that encompasses both, great. But I'm gonna have to pick Jay's brain for that. | |
| 22:40:51 | sean-k-mooney | e.g. pulling out one cable wont bring donw the bond | |
| 22:42:24 | sean-k-mooney | efried: if you can show your propasl support affinity in a clean way and does not block future addtion of anti afinity by design then its better then nothing | |
| 22:43:02 | sean-k-mooney | e.g. we dont have to support everything in one go in rocky | |
| 22:43:33 | sean-k-mooney | it would be nice however | |
| 22:44:25 | efried | That's kinda what I was thinking. Propose it this way, show that it wouldn't prevent a more generic solution in the future, and then Jay can propose (or not) that more generic solution. | |
| 22:46:14 | cfriesen | efried: what about multi-numa-node guests? they would need to essentially do numa anti-affinity (but host affinity) between the two pools of memory | |
| 22:47:02 | efried | cfriesen: What does that mean? Multi-numa-node guests? You mean guests that don't care about NUMA affinity? Or do they explicitly care about *anti* affinity? | |
| 22:47:05 | sean-k-mooney | cfriesen: multi numa guest do not need to be on multiple host numa nodes | |
| 22:47:14 | cfriesen | efried: no, I mean guests with multi numa nodes | |
| 22:47:22 | sean-k-mooney | cfriesen: i consider that limitation in libvirt to be a bug not a feature | |
| 22:47:30 | sean-k-mooney | libvirt virt driver that is | |
| 22:47:45 | cfriesen | sean-k-mooney: one reason why you might "require" multi-numa-node guests is to get access to more memory bandwidth | |
| 22:48:48 | sean-k-mooney | cfriesen: yes but there is a flavour extra spec defined for that that only the hyperv dirver support | |
| 22:49:00 | sean-k-mooney | cfriesen: i think the libvirt one should too | |
| 22:49:17 | cfriesen | sean-k-mooney: sorry, don't follow. | |
| 22:50:09 | sean-k-mooney | cfriesen: hw:numa_mempolicy | |
| 22:50:13 | cfriesen | sean-k-mooney: I can currently specify a multi-numa-node guest with libvirt, and I get access to that many host numa nodes, with accordingly more host memory bandwidth. (presumably combined with dedicated cpus) | |
| 22:50:59 | cfriesen | numa_mempolicy isn't mentioned anywhere in nova | |
| 22:51:41 | sean-k-mooney | cfriesen: i used to be in the hyperv driver | |
| 22:52:06 | sean-k-mooney | anyway that is off topic | |
| 22:52:25 | sean-k-mooney | efried: a multi numa gues is a guest with multiple virtual numa nodes | |
| 22:52:55 | cfriesen | efried: and currently they map to different physical numa nodes of the same host. | |
| 22:53:03 | sean-k-mooney | efried: you can choose what virtual cores are associated to each virtual numa node and you can device the memoy between each virtual numa node | |
| 22:54:09 | efried | Well, if you wanted to get that specific, you could probably swang something with granular request groups. | |
| 22:54:29 | efried | But we still wouldn't be able to guarantee you anti-affinity. | |
| 22:54:49 | cfriesen | if you can't guarantee numa-antiaffinity it'll change the current behaviour | |
| 22:55:02 | cfriesen | (ie if both virtual numa nodes could end up on the same host numa node) | |
| 22:55:06 | sean-k-mooney | placement does not need to give that guarentee | |
| 22:55:14 | sean-k-mooney | the virt driver can via the resouce tracker | |
| 22:55:40 | efried | sean-k-mooney: Yeahhhh, but we eventually *should* be able to implement that filter in placement. | |
| 22:55:44 | sean-k-mooney | or the nova numa topology filter | |
| 22:56:09 | sean-k-mooney | efried: talk to jay about that. in principal i agree | |
| 22:57:23 | sean-k-mooney | i would like placement to fully handel numa. but if it can handel 80% of numa in rocky with the rest in the filetr i would be happy with that too | |
| 22:57:46 | efried | Okay. I'll try to put up that spec later this week. | |