Earlier  
Posted Nick Remark
#openstack-nova - 2020-01-29
17:37:47 stephenfin defaults to off
17:38:00 efried We force deployers to split their data center.
17:38:00 efried N hosts are NUMA-modeled. All and only VMs with a NUMA topo land on those hosts.
17:38:00 efried M hosts are flat. All and only *non* NUMA topo VMs land on those hosts.
17:38:36 sean-k-mooney so in pricipal you shoudl be doing that partioning alreday
17:38:37 sean-k-mooney for reason i wont go into today
17:38:41 efried exactly.
17:38:50 stephenfin she borked.
17:38:57 stephenfin (if you don't do that partitioning)
17:39:01 sean-k-mooney elequently said
17:39:04 efried there may be a way we can support a small amount of overlap in the future, but for the moment I think it wouldn't be a terrible idea to add (under the covers) a required or forbidden trait that makes sure you land on the right kind of host.
17:39:27 efried like the NUMA_ROOT trait on the provider from which you're requesting $resource.
17:39:43 efried so let's get back to pages.
17:39:44 stephenfin that might be needed to handle the upgrade impact
17:39:59 efried How many "abstract" page sizes are there?
17:40:07 efried small/large/huge?
17:40:14 stephenfin we can't expect users to go set the "this is NUMA node" flag ahead of time
17:40:18 stephenfin efried: small, large
17:40:29 stephenfin hey, I documented that for my validator
17:40:30 efried stephenfin: totally not, that happens under the covers, now and forever, if it happens at all.
17:40:45 efried okay, and how many *discrete* page sizes are there within those buckets?
17:40:45 sean-k-mooney small/large/any or a specifc page size expressed as an integer with an optional suffix
17:41:04 stephenfin https://review.opendev.org/#/c/704643/2/nova/api/validation/extra_specs/hw.py@113
17:41:09 efried \o/
17:41:12 stephenfin noting sean-k-mooney's comment
17:41:15 sean-k-mooney the integuer is unbounded but in pratice about 12ish
17:41:32 stephenfin efried: architecture dependent
17:41:40 stephenfin Intel supports 2M and 1G
17:41:47 stephenfin *x86
17:41:49 efried lovely. Is there overlap on what's considered small and large?
17:41:57 stephenfin POWER supports 8G, iirc
17:41:57 sean-k-mooney efried: jay put up a patch to os resouces for the common ones a while ago
17:42:01 stephenfin Who knows what ARM supports
17:42:04 efried that is, are there discrete values that are considered small on some systems and large on others?
17:42:10 stephenfin ALL the pagesizes
17:42:22 sean-k-mooney efried: in practice not really but technically there can be
17:42:30 stephenfin sean-k-mooney: 4k is a small page _everywhere_, right?
17:42:32 sean-k-mooney small is almost always 4k
17:42:38 sean-k-mooney no
17:42:40 efried those are different statements
17:42:54 sean-k-mooney small is the native pagesize of the host
17:43:01 sean-k-mooney and the smallest in the available set
17:43:25 sean-k-mooney it is almost alwasy 4k and raely 16k or 64k
17:43:47 stephenfin sean-k-mooney: The internet tells me 4k is hardcoded as the default page size in Linux
17:43:57 stephenfin https://unix.stackexchange.com/a/128218
17:44:18 sean-k-mooney yes on x86, aarch64 and power9
17:44:41 stephenfin okay, we don't need to care about anything else, realistically
17:44:56 sean-k-mooney that why i said its almost alwasy 4k
17:45:01 stephenfin we can stick a giant TODO in somewhere in case someone wants to run openstack on obscure architecture
17:45:03 sean-k-mooney we can proably treat it as such
17:45:07 stephenfin agreed
17:45:15 stephenfin s/TODO/NOTE/
17:45:36 stephenfin efried: what are you referring to?
17:45:44 sean-k-mooney large is defiend as any pagesize that is not the same as small on the plathform
17:45:49 efried okay, so I'm afraid the compromise we need to make is this:
17:45:49 efried Represent all the memory as MEMORY_MB with a step_size that's the least common denominator of all the page sizes.
17:45:49 efried Supply traits saying things like I_HAVE_$SIZE_PAGES_HERE, where $SIZE can include both discrete and abstract sizes.
17:45:49 efried Use ^ where possible to make the placement pass a bit better
17:45:49 efried but accept that it's not going to be perfect, and that we have to do the rest in the NTF.
17:46:12 efried ...which may entail bouncing a host late.
17:46:53 stephenfin I don't understand how step_size would work
17:46:54 sean-k-mooney efried: if we do that we need to keep the hugepage code in the resouce tracker and numa toplogy filter for ever
17:47:24 efried I'm afraid that may be unavoidable, unless we make big changes in placement, or people agree to stop needing that shit.
17:47:29 sean-k-mooney efried: if we had a placement exteion weree we could pass a step size in the querry that would help
17:47:39 stephenfin I have a host that have 16 1GB hugepages and the rest are normal small pages
17:47:45 efried stephenfin: example, if we have 4K, 1M, and 2G pages, the step_size has to be 4K
17:47:59 stephenfin you will always use 4k so
17:48:06 stephenfin every host is going to report small pages
17:48:13 efried okay, so be it.
17:48:15 sean-k-mooney efried: not all the moroy can be allocated in 4k
17:48:15 sean-k-mooney on that host
17:48:23 sean-k-mooney hugepages are preallcoted
17:48:31 sean-k-mooney with a given page size
17:48:32 stephenfin I don't get how this solves anything
17:48:32 efried yeah, I understand that. We have to do that part via the NTF. I don't see a way around it.
17:48:40 efried "solves"
17:48:47 efried it gives us a way to move forward
17:49:03 efried and support all the variants we need
17:49:14 sean-k-mooney well the thing is if we had different resocue classes hugepage was one of the things we had determin could be fully done by placment
17:49:19 stephenfin this is essentially not modelling different page sizes in placement
17:49:21 sean-k-mooney if we modeld the different pages sizes in placment
17:49:25 efried stephenfin: correct.
17:49:46 stephenfin NUMA is all about memory
17:49:47 sean-k-mooney so thats a lot of work for very littel benifit
17:49:50 sean-k-mooney yep
17:49:53 efried we cannot do that and also support requests for MEMORY_MB=$X that don't care about pages.
17:50:03 stephenfin what's the point in modelling NUMA in placement if we don't do the memory modelling?
17:50:08 efried affinity
17:50:17 sean-k-mooney efried: we can if MEMORY_MB is always for non numa memory
17:50:19 stephenfin we have the NUMA topology filter for that
17:50:28 stephenfin good enough
17:50:41 efried stephenfin: this allows us to move a big swatch of the affinity filtering to the placement query.
17:50:48 efried just not all of it.
17:50:50 sean-k-mooney efried: no it does not
17:50:54 efried and not as much as we would have liked
17:50:56 efried wha?
17:50:58 stephenfin very little of it
17:51:00 efried wha??
17:51:01 sean-k-mooney we will have to do it all again in the ntf
17:51:08 stephenfin yup

Earlier   Later