Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-04
10:22:34 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Cleanup of existing index pages https://review.openstack.org/498819
10:22:34 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add configuration index page https://review.openstack.org/498818
10:22:35 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add user index page https://review.openstack.org/498817
10:25:36 stephenfin sahid: RE: https://review.openstack.org/#/c/472633/, I still don't agree, I'm afraid. I'm not blocking it but I can't +2
10:25:49 stephenfin Think my comments are still valid
10:29:05 sean-k-mooney stephenfin: sahid eek im not sure https://review.openstack.org/#/c/472633 is a good idea
10:29:40 stephenfin sean-k-mooney: Do tell?
10:30:17 sean-k-mooney stephenfin: sahid looked memory cannot be swapped out by the kernel even when the host is running out of memory. the limit is there to prevent a memory leak in the guest or a malitious guest form exauxting the host memory and crashing the host
10:31:18 sahid sean-k-mooney: yes we know that, but we don't have any way to compute the necessary amount of memory needed by QEMU
10:31:57 sean-k-mooney we can put a better upper bound then infinity though
10:32:21 sahid sean-k-mooney: like which one?
10:33:01 sean-k-mooney its rather arbitry but memory request + 1G would be better. the hard limit is there to account for qemu overhead. 1G is overkill but again better the infinity
10:33:52 sahid sean-k-mooney: not sure that is make sense since that 1G can run out of memory
10:34:26 sean-k-mooney sahid: its true it can but only if the reserved memory on the host is less then 1 GB
10:35:17 sahid yes but you also are adding a limit which can make the process to be killed for any reason
10:35:56 sean-k-mooney sahid: well if the host is really running out of memory even with locked memory the OOM killer willl possibly do that anyway
10:36:08 sahid are you sure that 1G is enought for QEMU if running a virtual machine of 128GB?
10:36:27 sahid ans what about 256GB?
10:36:50 sahid basically we try to follow with that patch what libvirt is doind right now
10:37:01 sean-k-mooney sure not but it is a qusteion that i would like to ask the qemu comunity to comment on or make it a config option rather then a hardcoded limit
10:37:06 sahid this patch is fixing a bug in older version of libvirt
10:37:38 openstackgerrit Merged openstack/nova master: doc: fix online_data_migrations option in upgrades doc https://review.openstack.org/500124
10:38:50 sahid sean-k-mooney: https://www.redhat.com/archives/libvir-list/2017-March/msg01092.html
10:39:40 sean-k-mooney 9Yetabytes is likely too much. 1G might be too small but im not convicned of that.
10:40:07 sean-k-mooney sahid: yes the "Use with extreme care" bit is why im not sure its a good idea to do this by default with locked memory in openstack
10:40:31 sean-k-mooney ill be back soon have to go to scrum
10:42:05 sahid sean-k-mooney: it's the current behavior with newer version of libvirt, that patch is fixing an issue for older version. since we do not have any way to compute the necessary amount of memory needed by QEMU we can't arbitrary set a limit
10:59:57 kashyap mdbooth: When you get a moment, this is in your wheelhouse. Would appreciate your view - https://review.openstack.org/#/c/498983/
11:00:34 mdbooth Why does that ring a bell?
11:05:49 sean-k-mooney sahid: if the intent of the path is just to match the bevavior of new libvirt i guess that is ok. can you add a release not with a security section thoguh for this in https://review.openstack.org/#/c/472633
11:06:52 sean-k-mooney sahid: incidentally what happens today if you just dont set the hardlimit in the xml at all?
11:09:17 sahid sean-k-mooney: libvirt is going to add it for you
11:10:03 sahid sean-k-mooney: seems reasonable to have a reno note yes, let me update the patch
11:10:12 sean-k-mooney sahid: and for new libvirt its unlimited and old it used to add a gig memKB = virDomainDefGetMemoryTotal(def) + 1024 * 1024;
11:11:41 sean-k-mooney ok if there is no other way to calulate a safe hard limit the i guess this is the best we can do
11:12:06 sahid sean-k-mooney: not sure to have understand, for old libvirt that is not set at all where it's something mandatory when memory is locked on host
11:12:27 sahid for new libvirt it's set when you do not have specifically set it in domain XML
11:13:06 sahid sean-k-mooney: yeah, thanks.. let me update the patch to add a reno note and see if that is going to make moving things
11:13:29 sean-k-mooney oh i taught the hard_limit was optional for some reason... any way the effect of your patch is to make the behavior the same regradless of the libvirt you are using
11:13:52 sahid exactly
11:14:23 sean-k-mooney well from a debuging perspecitve that alone is a good thing
11:44:41 sean-k-mooney stephenfin: i need to check the placement code again but currently when making a placement alocation can i say which specific resouce from a resouce pool i am allocating
11:46:09 sean-k-mooney stephenfin: i.g. for a 8 core cpu can i allocate core 1 and 7 to a vm or can i only allocate 2 cores to a vm? with nested resouce providers that is.
11:46:39 sean-k-mooney stephenfin: i belive we will be able to do the former correct
11:53:39 openstackgerrit Balazs Gibizer proposed openstack/nova stable/pike: reno: mention that customer resource are not supported https://review.openstack.org/500521
11:59:41 ps_jadhav gegelio
12:22:30 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/500011
12:41:02 stephenfin sean-k-mooney: Sorry - was gone for lunch
12:41:52 stephenfin sean-k-mooney: To the best of my recollection, you won't be able to do anything as specific as that. We won't be doing things like handling PCI-NUMA affinity in placement - that'll all remain a compute-node level operation
13:06:57 kashyap stephenfin: Crazy nit - I accidentally noticed -- you asked for lower-casing of Nova here. Shouldn't the 'N' in Nova always be in caps? - https://review.openstack.org/#/c/476188/3/doc/source/user/serial_console.rst
13:07:54 stephenfin kashyap: Nope - it's a weird thing the docs team have. Project names are not capitalized
13:08:17 kashyap stephenfin: Ah, okay. You're trying to be consistent with the pre-existing rule
13:08:46 stephenfin kashyap: 'zactly. I don't know why it's that way, but if everyone else is doing then we best do it too
13:09:13 kashyap Sure, consistency is nice. Just felt tripped by the crazy
13:18:48 sean-k-mooney stephenfin: then that a major limitation of placement.
13:19:48 sean-k-mooney i was hopign ot not only do pci affinity in placement but also cpu thread policies and other fetures... if we cannot move them to placement we will need to keep adding new code to nova to handel this
13:20:47 stephenfin sean-k-mooney: Sec. I'm pretty sure I've a mail from jaypipes about this
13:20:58 jaypipes stephenfin: quoi?
13:21:12 sean-k-mooney hi jaypipes o/
13:21:17 jaypipes heyo :)
13:21:20 sean-k-mooney are you not ment to be on vacation
13:21:34 jaypipes sean-k-mooney: yes, all those things would be traits against resource providers.
13:21:41 jaypipes sean-k-mooney: meh :)
13:21:47 jaypipes sean-k-mooney: it's labor day. I'm laboring.
13:22:30 sean-k-mooney jaypipes: well traits may not be enough to model some of the specific things we will need to schedle on.
13:22:40 jaypipes sean-k-mooney: example?
13:23:16 sean-k-mooney i added some examples to the ptg schedule i think
13:23:23 jaypipes sean-k-mooney: ah, k
13:23:24 sean-k-mooney let me check
13:23:39 stephenfin sean-k-mooney: So this is what jaypipes and I discussed after the last summit. Not sure if it's still all relevant, but jaypipes can correct me if not http://paste.openstack.org/show/620336/
13:24:48 jaypipes that is still correct, yes, stephenfin
13:25:13 sean-k-mooney jaypipes: see lines 106-125 https://etherpad.openstack.org/p/nova-ptg-queens
13:25:50 jaypipes sean-k-mooney: the idea is that we prevent the reschedule problem by claiming NUMA resources (a quantity of cores, threads, sockets) but we don't reserve specific cores/threads/pinset in placement. Instead, we rely on the existing code to do that.
13:26:50 sean-k-mooney so we need to keep a second set of tabels in nova to track which specific cores are allocated
13:27:52 jaypipes sean-k-mooney: no... we already do... it's the numa_topology field in compute_nodes table.
13:28:14 sean-k-mooney yes and for things other then numa?
13:29:36 jaypipes sean-k-mooney: not entirely sure what example you're giving there.
13:30:08 jaypipes sean-k-mooney: are you referring to a specific VF having bandwidth resources doled out by placement?
13:31:21 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544
13:31:30 sean-k-mooney that is one usecase yes. we have partly implemented the neutron half to enfore the mimium bandwith qos policy on the vf but seperate form that intel has a set of technologies that we have supported in our severs since haswell such as cache allocation technology which allow to confine/allocate cache to a specific porcsess
13:31:56 sean-k-mooney the cache allocation has socket affinity but not numa affintiy
13:32:37 sean-k-mooney we could track that with a trait but its not moddled in the numa_topology struture in nova today
13:33:10 jaypipes le sigh... hardware-defined software at its finest. :)
13:34:40 sean-k-mooney our cascade lake server plathform will have memory bandwith allocation support also which is even more fun.
13:35:18 sean-k-mooney and yes you know i dont consider my self that much of a hardware guy just working at intel i get exposed to all this stuff...
13:35:30 jaypipes sean-k-mooney: well, technically, it's indeed possible to model all this stuff with the nested resource providers modeling...
13:36:10 sean-k-mooney jaypipes: yes the bandwith is straigt forward its a nested resouce provider that is a child of the pf
13:36:33 jaypipes sean-k-mooney: I think we probably need to add some standard resource class types that represent this stuff. doesn't seem like the existing NUMA_SOCKET, NUMA_CORE, and NUMA_THREAD resource classes are going to be enough.
13:37:38 sean-k-mooney well those three classes shoudl not have numa in the name in the first palce
13:37:53 jaypipes sean-k-mooney: why not?
13:38:16 jaypipes sean-k-mooney: remember, these are the classes of resources the *guest* gets
13:38:22 sean-k-mooney numa is soly about your memory layout and is not related to you cpu core topology implcitly
13:38:42 jaypipes sean-k-mooney: a guest doesn't get LPCU#12 and LPCU#16. It gets 2 NUMA_CORE
13:39:12 jaypipes LCPU...
13:39:29 sean-k-mooney maybe form a guest perspective it will work but you can have 0-n numa nodes per socket and you can have numa nodes not connected to a socket
13:39:49 sean-k-mooney you bassicaly have 1 numa node per memory contoler
13:40:07 jaypipes sean-k-mooney: doesn't sound like any of this is quantitative.
13:40:23 jaypipes sean-k-mooney: it's all just to represent distance to memory, no?
13:41:00 sean-k-mooney am sort of, it distance(latency) and bandwith to memory
13:41:50 jaypipes gah, I wish we could punt this to k8s. oh wait, they don't want it either...

Earlier   Later