Earlier  
Posted Nick Remark
#openstack-nova - 2020-01-29
17:55:10 efried cool
17:55:17 efried then yes, split into two resource classes.
17:55:18 sean-k-mooney efried: there is mempage_size=any
17:55:25 efried ugh
17:55:30 sean-k-mooney but that just means the image gets to choose
17:55:36 stephenfin N resource classes
17:55:39 efried and if the image doesn't specify... bounce?
17:55:39 sean-k-mooney and it will be small if not stated
17:55:50 efried stephenfin: Still can't do N resource classes. Have to have two.
17:55:56 stephenfin Can't do two
17:55:56 efried one for small, one for large.
17:56:10 efried Each one has a step_size that's the least common denominator of the pages in that range.
17:56:11 stephenfin Host has some 1G hugepages and some 2M pages
17:56:22 sean-k-mooney efried: any was ment to give you the smallest pagesize available but we skimed on it and just do 4k pages i think
17:56:25 stephenfin I can't request some of the former and some of the latter
17:56:30 stephenfin It has to be one or the other
17:56:42 sean-k-mooney stephenfin: that is a nova limitation
17:56:44 efried okay, but if you say LARGE, we get to pick
17:56:44 stephenfin So placement could say "oh, I have hugepages", but it turns out they're different sizes
17:56:55 sean-k-mooney not a libvirt one but we should not mix them
17:56:57 stephenfin and it can't actually fulfil that request
17:57:00 efried yeah, the resource class would not indicate number of pages, it's still a number of MB.
17:57:08 efried right, that's the part we would have to defer to the NTF
17:57:27 efried Otherwise you *still* can't say memory=4GB,page_size=large because we would have no way to translate that request.
17:57:34 sean-k-mooney yes
17:57:50 efried ...without doing multiple placement queries, which is a hard no.
17:57:55 stephenfin then find a way to kill page_size=large or convert it at the API level
17:57:58 sean-k-mooney so there is a way to do this with only one resocue class + traits
17:58:01 stephenfin api config option
17:58:11 stephenfin large_mem_page_size = PAGE_2M
17:58:15 sean-k-mooney but to do that it need a 3 level resouce provider tree
17:58:26 stephenfin that's required from N+1
17:58:46 stephenfin as an interim step in N, we make two requests, one for PAGE_2M, one for PAGE_1G
17:58:51 sean-k-mooney stephenfin: mem_page_size=large is the most comonly used value
17:58:52 stephenfin like we're doing for VCPU and PCPU rn
17:59:02 stephenfin sean-k-mooney: yup, and we can continue supporting it
17:59:07 efried is that not config-driven API behavior?
17:59:13 efried the thing you just told me was a no-no
17:59:30 stephenfin we just have to ask operators to define what large aliases to
17:59:40 sean-k-mooney no
17:59:47 efried but that only works if the deployment only has one large page size?
17:59:51 stephenfin on the assumption that people aren't mixing and matching different page sizes
17:59:55 efried and if that's true, then the whole issue is moot
17:59:55 sean-k-mooney because its not uncommon for the same cloud to have both
18:00:09 efried if we can assume that, we're golden.
18:00:10 stephenfin why would they do that?
18:00:12 sean-k-mooney that is not a safe assumtion
18:00:24 sean-k-mooney because differnt applciation perfrom better
18:00:35 sean-k-mooney with different page sizes
18:00:53 stephenfin by application you mean something in an instance?
18:01:02 stephenfin if so, are they actually using 'large'?
18:01:07 stephenfin vs. the size that performs better
18:01:10 stephenfin *best
18:01:10 sean-k-mooney also 1G was added a lot later then 2mb
18:01:11 sean-k-mooney it requires specifc cpu support
18:01:51 stephenfin okay, so this will force operators to choose one or the other for their deployment
18:01:54 sean-k-mooney they use large for applciation that can support either and 1G when needed
18:02:14 sean-k-mooney yes which i dont think we need to do
18:02:24 sean-k-mooney im going to skech somthing up for ye to review
18:02:31 sean-k-mooney give me 2 mins
18:02:32 stephenfin yeah, good idea
18:02:39 stephenfin I know what efried's approach is and I don't like it
18:02:45 stephenfin Don't grok yours yet though
18:03:11 stephenfin Also, I'm supposed to be going out for dinner so I should probably scarper for now
18:03:12 sean-k-mooney what is the ascii site we use for specs
18:03:21 stephenfin ascii?
18:03:25 stephenfin ascii site?
18:03:33 stephenfin for drawing?
18:03:33 sean-k-mooney http://asciiflow.com/
18:03:50 sean-k-mooney ya i litrally ment draw since its clearer
18:04:00 stephenfin oh, I was going to say just use https://www.draw.io/
18:04:15 stephenfin or https://docs.google.com/drawings
18:04:25 stephenfin anyway, let me know what you draft
18:09:09 sean-k-mooney stephenfin: im thinking 3 layers like this
18:09:10 sean-k-mooney https://etherpad.openstack.org/p/mem_page_size_and_placement
18:14:39 sean-k-mooney efried: bauzas ^
18:14:41 efried sean-k-mooney: step_size=4 doesn't make sense unless we use MEMORY_KB
18:15:04 sean-k-mooney well actuly i gues it would be 1024
18:15:20 sean-k-mooney since we limit flavor to 1mb granuarity
18:15:39 efried tbc, it's not really a problem to introduce new RCs for this, since we're doing the translation under the covers and only allowing NUMA-modeled VMs on NUMA-modeled hosts.
18:15:46 sean-k-mooney its only there becasue you can say mem_page_size=4 or 4k today
18:15:53 efried okay, that's fine.
18:16:33 efried The three-tiered approach works IF you always get exactly one page size
18:16:35 sean-k-mooney this will allow all the sentinels to wrok and we can remove the hugepage page tracking from the numa toplogy filter/resouce tracker
18:16:51 sean-k-mooney efried: yes today we only allow 1 page size
18:17:13 sean-k-mooney so if we dont enable more flexiblity then today we can make that assumtion a requirement
18:18:39 efried so if I have three different hugepage sizes on the same numa cell, say 1G, 2G, 4G, then memory=8G,page_size=large will only ever get me one of
18:18:40 efried [2 x 4G]
18:18:40 efried [4 x 2G]
18:18:40 efried [8 x 1G]
18:18:40 efried but never e.g. [1 x 4G] + [4 x 1G]
18:18:47 efried right?
18:19:03 sean-k-mooney yes
18:19:10 efried cool, then this works, I like it.
18:19:13 sean-k-mooney libvirt support mixing and we intentionaly do not
18:19:49 sean-k-mooney i propsoed this in the past and the main push back is an extra layer adds well an extra layer
18:19:49 efried shall I write it up in the spec comments?
18:20:06 efried yeah, the layer doesn't bother me. It's totally abstracted from the user.
18:20:09 sean-k-mooney sure that would be awsome
18:20:16 efried cool, on it.

Earlier   Later