| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-01-29 | |||
| 17:49:03 | efried | and support all the variants we need | |
| 17:49:14 | sean-k-mooney | well the thing is if we had different resocue classes hugepage was one of the things we had determin could be fully done by placment | |
| 17:49:19 | stephenfin | this is essentially not modelling different page sizes in placement | |
| 17:49:21 | sean-k-mooney | if we modeld the different pages sizes in placment | |
| 17:49:25 | efried | stephenfin: correct. | |
| 17:49:46 | stephenfin | NUMA is all about memory | |
| 17:49:47 | sean-k-mooney | so thats a lot of work for very littel benifit | |
| 17:49:50 | sean-k-mooney | yep | |
| 17:49:53 | efried | we cannot do that and also support requests for MEMORY_MB=$X that don't care about pages. | |
| 17:50:03 | stephenfin | what's the point in modelling NUMA in placement if we don't do the memory modelling? | |
| 17:50:08 | efried | affinity | |
| 17:50:17 | sean-k-mooney | efried: we can if MEMORY_MB is always for non numa memory | |
| 17:50:19 | stephenfin | we have the NUMA topology filter for that | |
| 17:50:28 | stephenfin | good enough | |
| 17:50:41 | efried | stephenfin: this allows us to move a big swatch of the affinity filtering to the placement query. | |
| 17:50:48 | efried | just not all of it. | |
| 17:50:50 | sean-k-mooney | efried: no it does not | |
| 17:50:54 | efried | and not as much as we would have liked | |
| 17:50:56 | efried | wha? | |
| 17:50:58 | stephenfin | very little of it | |
| 17:51:00 | efried | wha?? | |
| 17:51:01 | sean-k-mooney | we will have to do it all again in the ntf | |
| 17:51:08 | stephenfin | yup | |
| 17:51:14 | efried | um | |
| 17:51:14 | efried | no | |
| 17:51:16 | efried | look | |
| 17:51:20 | sean-k-mooney | we may have less host to check | |
| 17:51:26 | stephenfin | what's the objection to modelling different page sizes as different resource types? | |
| 17:51:31 | stephenfin | they're different things | |
| 17:51:36 | sean-k-mooney | but we still need to check all aspecs again | |
| 17:51:38 | stephenfin | something that ask for one can't consume the other | |
| 17:51:54 | efried | Because it makes it impossible to request a certain amount of memory without asking for specific pages. | |
| 17:51:58 | sean-k-mooney | there is another way to do it | |
| 17:52:06 | sean-k-mooney | we can do it in 1 resouce class | |
| 17:52:10 | sean-k-mooney | MEMORY_MB | |
| 17:52:11 | stephenfin | but it doesn't | |
| 17:52:18 | efried | how would you do it stephenfin? | |
| 17:52:19 | sean-k-mooney | but we need 1 RP per page size | |
| 17:52:31 | stephenfin | we report MEMORY_MB for normal 4k pages | |
| 17:52:33 | sean-k-mooney | and each RP would be under a numa node | |
| 17:52:49 | stephenfin | and then PAGES_4M, PAGES_1G, etc. for the various hugepage sizes | |
| 17:53:00 | sean-k-mooney | that wont work | |
| 17:53:10 | stephenfin | if instances aren't requesting page sizes, they get MEMORY_MB | |
| 17:53:11 | sean-k-mooney | because it breaks large | |
| 17:53:22 | stephenfin | it does, and we'll need to figure out a migration plan for that | |
| 17:53:27 | sean-k-mooney | mem_page_size=large is the most common case | |
| 17:53:32 | efried | Okay, so if you don't ask for large pages, you only get small pages, and if the only way to satisfy your memory requirement would be by using the large pages, you just bounce. | |
| 17:53:41 | stephenfin | you can't use the large pages | |
| 17:53:46 | stephenfin | unless you *explicitly* ask for them | |
| 17:53:50 | efried | is that what happens today? | |
| 17:53:52 | stephenfin | yup | |
| 17:54:01 | efried | okay cool. | |
| 17:54:09 | sean-k-mooney | yes | |
| 17:54:16 | stephenfin | given a host with 8GB of RAM, of which 4GB is allocated to 1GB huge pages | |
| 17:54:21 | sean-k-mooney | which is why you have to partion the host today | |
| 17:54:37 | sean-k-mooney | we have 2 memory tracker in nova today | |
| 17:54:45 | efried | is it possible today to ask for both small and large pages in the same request, or is mem_page_size=large all or nothing? | |
| 17:54:46 | stephenfin | only 4GB is actually usable | |
| 17:54:57 | sean-k-mooney | one is numa and page size aware and the other just looks at total memory | |
| 17:55:07 | stephenfin | nope, all or nothing | |
| 17:55:10 | efried | cool | |
| 17:55:17 | efried | then yes, split into two resource classes. | |
| 17:55:18 | sean-k-mooney | efried: there is mempage_size=any | |
| 17:55:25 | efried | ugh | |
| 17:55:30 | sean-k-mooney | but that just means the image gets to choose | |
| 17:55:36 | stephenfin | N resource classes | |
| 17:55:39 | efried | and if the image doesn't specify... bounce? | |
| 17:55:39 | sean-k-mooney | and it will be small if not stated | |
| 17:55:50 | efried | stephenfin: Still can't do N resource classes. Have to have two. | |
| 17:55:56 | stephenfin | Can't do two | |
| 17:55:56 | efried | one for small, one for large. | |
| 17:56:10 | efried | Each one has a step_size that's the least common denominator of the pages in that range. | |
| 17:56:11 | stephenfin | Host has some 1G hugepages and some 2M pages | |
| 17:56:22 | sean-k-mooney | efried: any was ment to give you the smallest pagesize available but we skimed on it and just do 4k pages i think | |
| 17:56:25 | stephenfin | I can't request some of the former and some of the latter | |
| 17:56:30 | stephenfin | It has to be one or the other | |
| 17:56:42 | sean-k-mooney | stephenfin: that is a nova limitation | |
| 17:56:44 | efried | okay, but if you say LARGE, we get to pick | |
| 17:56:44 | stephenfin | So placement could say "oh, I have hugepages", but it turns out they're different sizes | |
| 17:56:55 | sean-k-mooney | not a libvirt one but we should not mix them | |
| 17:56:57 | stephenfin | and it can't actually fulfil that request | |
| 17:57:00 | efried | yeah, the resource class would not indicate number of pages, it's still a number of MB. | |
| 17:57:08 | efried | right, that's the part we would have to defer to the NTF | |
| 17:57:27 | efried | Otherwise you *still* can't say memory=4GB,page_size=large because we would have no way to translate that request. | |
| 17:57:34 | sean-k-mooney | yes | |
| 17:57:50 | efried | ...without doing multiple placement queries, which is a hard no. | |
| 17:57:55 | stephenfin | then find a way to kill page_size=large or convert it at the API level | |
| 17:57:58 | sean-k-mooney | so there is a way to do this with only one resocue class + traits | |
| 17:58:01 | stephenfin | api config option | |
| 17:58:11 | stephenfin | large_mem_page_size = PAGE_2M | |
| 17:58:15 | sean-k-mooney | but to do that it need a 3 level resouce provider tree | |
| 17:58:26 | stephenfin | that's required from N+1 | |
| 17:58:46 | stephenfin | as an interim step in N, we make two requests, one for PAGE_2M, one for PAGE_1G | |
| 17:58:51 | sean-k-mooney | stephenfin: mem_page_size=large is the most comonly used value | |
| 17:58:52 | stephenfin | like we're doing for VCPU and PCPU rn | |
| 17:59:02 | stephenfin | sean-k-mooney: yup, and we can continue supporting it | |
| 17:59:07 | efried | is that not config-driven API behavior? | |
| 17:59:13 | efried | the thing you just told me was a no-no | |
| 17:59:30 | stephenfin | we just have to ask operators to define what large aliases to | |
| 17:59:40 | sean-k-mooney | no | |