Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-04
13:18:48 sean-k-mooney stephenfin: then that a major limitation of placement.
13:19:48 sean-k-mooney i was hopign ot not only do pci affinity in placement but also cpu thread policies and other fetures... if we cannot move them to placement we will need to keep adding new code to nova to handel this
13:20:47 stephenfin sean-k-mooney: Sec. I'm pretty sure I've a mail from jaypipes about this
13:20:58 jaypipes stephenfin: quoi?
13:21:12 sean-k-mooney hi jaypipes o/
13:21:17 jaypipes heyo :)
13:21:20 sean-k-mooney are you not ment to be on vacation
13:21:34 jaypipes sean-k-mooney: yes, all those things would be traits against resource providers.
13:21:41 jaypipes sean-k-mooney: meh :)
13:21:47 jaypipes sean-k-mooney: it's labor day. I'm laboring.
13:22:30 sean-k-mooney jaypipes: well traits may not be enough to model some of the specific things we will need to schedle on.
13:22:40 jaypipes sean-k-mooney: example?
13:23:16 sean-k-mooney i added some examples to the ptg schedule i think
13:23:23 jaypipes sean-k-mooney: ah, k
13:23:24 sean-k-mooney let me check
13:23:39 stephenfin sean-k-mooney: So this is what jaypipes and I discussed after the last summit. Not sure if it's still all relevant, but jaypipes can correct me if not http://paste.openstack.org/show/620336/
13:24:48 jaypipes that is still correct, yes, stephenfin
13:25:13 sean-k-mooney jaypipes: see lines 106-125 https://etherpad.openstack.org/p/nova-ptg-queens
13:25:50 jaypipes sean-k-mooney: the idea is that we prevent the reschedule problem by claiming NUMA resources (a quantity of cores, threads, sockets) but we don't reserve specific cores/threads/pinset in placement. Instead, we rely on the existing code to do that.
13:26:50 sean-k-mooney so we need to keep a second set of tabels in nova to track which specific cores are allocated
13:27:52 jaypipes sean-k-mooney: no... we already do... it's the numa_topology field in compute_nodes table.
13:28:14 sean-k-mooney yes and for things other then numa?
13:29:36 jaypipes sean-k-mooney: not entirely sure what example you're giving there.
13:30:08 jaypipes sean-k-mooney: are you referring to a specific VF having bandwidth resources doled out by placement?
13:31:21 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544
13:31:30 sean-k-mooney that is one usecase yes. we have partly implemented the neutron half to enfore the mimium bandwith qos policy on the vf but seperate form that intel has a set of technologies that we have supported in our severs since haswell such as cache allocation technology which allow to confine/allocate cache to a specific porcsess
13:31:56 sean-k-mooney the cache allocation has socket affinity but not numa affintiy
13:32:37 sean-k-mooney we could track that with a trait but its not moddled in the numa_topology struture in nova today
13:33:10 jaypipes le sigh... hardware-defined software at its finest. :)
13:34:40 sean-k-mooney our cascade lake server plathform will have memory bandwith allocation support also which is even more fun.
13:35:18 sean-k-mooney and yes you know i dont consider my self that much of a hardware guy just working at intel i get exposed to all this stuff...
13:35:30 jaypipes sean-k-mooney: well, technically, it's indeed possible to model all this stuff with the nested resource providers modeling...
13:36:10 sean-k-mooney jaypipes: yes the bandwith is straigt forward its a nested resouce provider that is a child of the pf
13:36:33 jaypipes sean-k-mooney: I think we probably need to add some standard resource class types that represent this stuff. doesn't seem like the existing NUMA_SOCKET, NUMA_CORE, and NUMA_THREAD resource classes are going to be enough.
13:37:38 sean-k-mooney well those three classes shoudl not have numa in the name in the first palce
13:37:53 jaypipes sean-k-mooney: why not?
13:38:16 jaypipes sean-k-mooney: remember, these are the classes of resources the *guest* gets
13:38:22 sean-k-mooney numa is soly about your memory layout and is not related to you cpu core topology implcitly
13:38:42 jaypipes sean-k-mooney: a guest doesn't get LPCU#12 and LPCU#16. It gets 2 NUMA_CORE
13:39:12 jaypipes LCPU...
13:39:29 sean-k-mooney maybe form a guest perspective it will work but you can have 0-n numa nodes per socket and you can have numa nodes not connected to a socket
13:39:49 sean-k-mooney you bassicaly have 1 numa node per memory contoler
13:40:07 jaypipes sean-k-mooney: doesn't sound like any of this is quantitative.
13:40:23 jaypipes sean-k-mooney: it's all just to represent distance to memory, no?
13:41:00 sean-k-mooney am sort of, it distance(latency) and bandwith to memory
13:41:50 jaypipes gah, I wish we could punt this to k8s. oh wait, they don't want it either...
13:43:18 sean-k-mooney the vf sellection policy section i added in line 106-125 of https://etherpad.openstack.org/p/nova-ptg-queens are actull more straight forword the modeling what we were just talking about
13:43:28 sean-k-mooney the featuer bases scheduling is just traits.
13:43:36 jaypipes yeah, it's the NUMA stuff that kills everything.
13:43:38 sean-k-mooney bandwith is jsut childe of pf
13:43:41 jaypipes yep
13:43:50 sean-k-mooney numa could be a trait i guess
13:44:03 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add plugin names as constants. https://review.openstack.org/500111
13:44:42 sean-k-mooney pf affinity is for bonding. if you bound 2 sriov nics in the guest you want vf to have anti affinity at the pf level
13:45:04 jaypipes sean-k-mooney: NUMA is definitely not a trait...
13:45:12 jaypipes sean-k-mooney: you don't "have NUMA" or not.
13:45:36 sean-k-mooney jaypipes: well actully if your plathform only has one memory controler then you dont have numa
13:46:13 jaypipes sean-k-mooney: you know what I meant :)
13:46:42 sean-k-mooney :) yes i do
13:47:05 sean-k-mooney that should read unfortunetly i do :)
13:47:11 jaypipes hehe
13:47:43 jaypipes sean-k-mooney: before we get to Denver, I'd appreciate if you and ralonsoh can prioritize these placement features.
13:48:11 jaypipes sean-k-mooney: we're def not going to get to them all, and I'd like to work on the top priority items that actually have a chance of being completed in Queens
13:48:36 sean-k-mooney yep understood.
13:48:55 jaypipes sean-k-mooney: and remember that we already have two higher priorities in placement-land: getting all the move operations straight (migrations, etc) and fully supporting shared storage
13:48:55 sean-k-mooney ill try and see what once have the most bang for there buck
13:49:12 jaypipes sean-k-mooney: and integrating traits into the virt driver and scheduling layers./
13:49:21 jaypipes sean-k-mooney: cool, thanks man
13:50:05 sean-k-mooney i mainly added the vf selection topic becase fo the pci passthrough overhaul section.
13:50:33 jaypipes sean-k-mooney: you mean the generic device manager thing? :)
13:50:35 openstackgerrit Dmitry Tantsur proposed openstack/nova master: Correct examples in "Manage Compute services" documentation https://review.openstack.org/500551
13:50:57 dtantsur just saw a person hit by this ^^^
13:51:58 jaypipes dtantsur: +2
13:52:19 dtantsur thnx
13:55:00 stephenfin alex_xu: Just wait til you get the two of them in a room together :P
13:58:20 stephenfin gibi: Easy "kill dead code" patch here, if you fancied taking a look https://review.openstack.org/#/c/483030/
14:07:29 gibi stephenfin: looking...
14:09:37 sean-k-mooney jaypipes: yes the generic device manager thing + there are a few other spects related to sriov in general such as attach/detach support.
14:09:49 jaypipes yup
14:22:45 sbezverk sean-k-mooney : ping
14:24:20 sean-k-mooney sbezverk: hi
14:25:21 sbezverk sean-k-mooney hi, quick question for you if you have a second. Are you aware of any in depth libvirt debugging/troubleshooting guide?
14:25:29 openstackgerrit Merged openstack/os-traits master: Add a new parameter ``suffix`` to function ``get_traits`` https://review.openstack.org/483451
14:26:08 sean-k-mooney sbezverk: not as such no. are you trying to find a guide to debug problems using libvirt or developing libvirt
14:27:54 sbezverk sean-k-mooney I see "x-files" type of problem with one VM and need to dig way deeper in libvirt to understand what is going on
14:29:01 sean-k-mooney well libvirt iteslef does almost nothing excepth take an xml and convert it into a qemu ars list
14:29:38 sean-k-mooney perhaps thats a bit unfair to it but usually the issue is not in libvirt but in the thing its calling of the thing that is driving it
14:30:25 sean-k-mooney sbezverk: what is the problem you are having?
14:30:47 sbezverk sean-k-mooney you asked yourself ;)
14:31:33 sbezverk sean-k-mooney one vm gets paused without any particular reason, I checked everything I could no issue with storage, plenty in vm and on the host
14:32:04 sbezverk run extensive diagnostic on the host memory/cpu/disk subsystem comes clean
14:32:45 sbezverk I wanted to see what process/even signals to libvirt to pause that vm
14:34:55 sean-k-mooney ok is there anyting in the instance log in /var/log/libvirt/qemu/instance-<id>.log
14:34:56 sbezverk even answering whether it is triggered outside on VM or some internal thing in the VM is problematic as no logs indicate anything abnormal
14:35:34 sean-k-mooney if libvirt recived a api call to pause the vm it should be in the libvirtd log in /var/log/libvirtd.log
14:36:45 sbezverk sean-k-mooney ok thanks I will double check, I enabled debug level of log for libvirt
14:52:03 sbezverk sean-k-mooney what do you think about this http://paste.openstack.org/show/620346/
14:56:06 sean-k-mooney sbezverk: you would not happen to be trying to use a new intel purly plathform or new rayzen server would you
14:56:51 sean-k-mooney sbezverk: i have seen this before but only when you are trying to use a gust cpu model that uses feature not availabel on the host
14:57:09 sbezverk sean-k-mooney nope it is cisco ucs box with 2600 something v4

Earlier   Later