| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-04 | |||
| 13:24:48 | jaypipes | that is still correct, yes, stephenfin | |
| 13:25:13 | sean-k-mooney | jaypipes: see lines 106-125 https://etherpad.openstack.org/p/nova-ptg-queens | |
| 13:25:50 | jaypipes | sean-k-mooney: the idea is that we prevent the reschedule problem by claiming NUMA resources (a quantity of cores, threads, sockets) but we don't reserve specific cores/threads/pinset in placement. Instead, we rely on the existing code to do that. | |
| 13:26:50 | sean-k-mooney | so we need to keep a second set of tabels in nova to track which specific cores are allocated | |
| 13:27:52 | jaypipes | sean-k-mooney: no... we already do... it's the numa_topology field in compute_nodes table. | |
| 13:28:14 | sean-k-mooney | yes and for things other then numa? | |
| 13:29:36 | jaypipes | sean-k-mooney: not entirely sure what example you're giving there. | |
| 13:30:08 | jaypipes | sean-k-mooney: are you referring to a specific VF having bandwidth resources doled out by placement? | |
| 13:31:21 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544 | |
| 13:31:30 | sean-k-mooney | that is one usecase yes. we have partly implemented the neutron half to enfore the mimium bandwith qos policy on the vf but seperate form that intel has a set of technologies that we have supported in our severs since haswell such as cache allocation technology which allow to confine/allocate cache to a specific porcsess | |
| 13:31:56 | sean-k-mooney | the cache allocation has socket affinity but not numa affintiy | |
| 13:32:37 | sean-k-mooney | we could track that with a trait but its not moddled in the numa_topology struture in nova today | |
| 13:33:10 | jaypipes | le sigh... hardware-defined software at its finest. :) | |
| 13:34:40 | sean-k-mooney | our cascade lake server plathform will have memory bandwith allocation support also which is even more fun. | |
| 13:35:18 | sean-k-mooney | and yes you know i dont consider my self that much of a hardware guy just working at intel i get exposed to all this stuff... | |
| 13:35:30 | jaypipes | sean-k-mooney: well, technically, it's indeed possible to model all this stuff with the nested resource providers modeling... | |
| 13:36:10 | sean-k-mooney | jaypipes: yes the bandwith is straigt forward its a nested resouce provider that is a child of the pf | |
| 13:36:33 | jaypipes | sean-k-mooney: I think we probably need to add some standard resource class types that represent this stuff. doesn't seem like the existing NUMA_SOCKET, NUMA_CORE, and NUMA_THREAD resource classes are going to be enough. | |
| 13:37:38 | sean-k-mooney | well those three classes shoudl not have numa in the name in the first palce | |
| 13:37:53 | jaypipes | sean-k-mooney: why not? | |
| 13:38:16 | jaypipes | sean-k-mooney: remember, these are the classes of resources the *guest* gets | |
| 13:38:22 | sean-k-mooney | numa is soly about your memory layout and is not related to you cpu core topology implcitly | |
| 13:38:42 | jaypipes | sean-k-mooney: a guest doesn't get LPCU#12 and LPCU#16. It gets 2 NUMA_CORE | |
| 13:39:12 | jaypipes | LCPU... | |
| 13:39:29 | sean-k-mooney | maybe form a guest perspective it will work but you can have 0-n numa nodes per socket and you can have numa nodes not connected to a socket | |
| 13:39:49 | sean-k-mooney | you bassicaly have 1 numa node per memory contoler | |
| 13:40:07 | jaypipes | sean-k-mooney: doesn't sound like any of this is quantitative. | |
| 13:40:23 | jaypipes | sean-k-mooney: it's all just to represent distance to memory, no? | |
| 13:41:00 | sean-k-mooney | am sort of, it distance(latency) and bandwith to memory | |
| 13:41:50 | jaypipes | gah, I wish we could punt this to k8s. oh wait, they don't want it either... | |
| 13:43:18 | sean-k-mooney | the vf sellection policy section i added in line 106-125 of https://etherpad.openstack.org/p/nova-ptg-queens are actull more straight forword the modeling what we were just talking about | |
| 13:43:28 | sean-k-mooney | the featuer bases scheduling is just traits. | |
| 13:43:36 | jaypipes | yeah, it's the NUMA stuff that kills everything. | |
| 13:43:38 | sean-k-mooney | bandwith is jsut childe of pf | |
| 13:43:41 | jaypipes | yep | |
| 13:43:50 | sean-k-mooney | numa could be a trait i guess | |
| 13:44:03 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add plugin names as constants. https://review.openstack.org/500111 | |
| 13:44:42 | sean-k-mooney | pf affinity is for bonding. if you bound 2 sriov nics in the guest you want vf to have anti affinity at the pf level | |
| 13:45:04 | jaypipes | sean-k-mooney: NUMA is definitely not a trait... | |
| 13:45:12 | jaypipes | sean-k-mooney: you don't "have NUMA" or not. | |
| 13:45:36 | sean-k-mooney | jaypipes: well actully if your plathform only has one memory controler then you dont have numa | |
| 13:46:13 | jaypipes | sean-k-mooney: you know what I meant :) | |
| 13:46:42 | sean-k-mooney | :) yes i do | |
| 13:47:05 | sean-k-mooney | that should read unfortunetly i do :) | |
| 13:47:11 | jaypipes | hehe | |
| 13:47:43 | jaypipes | sean-k-mooney: before we get to Denver, I'd appreciate if you and ralonsoh can prioritize these placement features. | |
| 13:48:11 | jaypipes | sean-k-mooney: we're def not going to get to them all, and I'd like to work on the top priority items that actually have a chance of being completed in Queens | |
| 13:48:36 | sean-k-mooney | yep understood. | |
| 13:48:55 | sean-k-mooney | ill try and see what once have the most bang for there buck | |
| 13:48:55 | jaypipes | sean-k-mooney: and remember that we already have two higher priorities in placement-land: getting all the move operations straight (migrations, etc) and fully supporting shared storage | |
| 13:49:12 | jaypipes | sean-k-mooney: and integrating traits into the virt driver and scheduling layers./ | |
| 13:49:21 | jaypipes | sean-k-mooney: cool, thanks man | |
| 13:50:05 | sean-k-mooney | i mainly added the vf selection topic becase fo the pci passthrough overhaul section. | |
| 13:50:33 | jaypipes | sean-k-mooney: you mean the generic device manager thing? :) | |
| 13:50:35 | openstackgerrit | Dmitry Tantsur proposed openstack/nova master: Correct examples in "Manage Compute services" documentation https://review.openstack.org/500551 | |
| 13:50:57 | dtantsur | just saw a person hit by this ^^^ | |
| 13:51:58 | jaypipes | dtantsur: +2 | |
| 13:52:19 | dtantsur | thnx | |
| 13:55:00 | stephenfin | alex_xu: Just wait til you get the two of them in a room together :P | |
| 13:58:20 | stephenfin | gibi: Easy "kill dead code" patch here, if you fancied taking a look https://review.openstack.org/#/c/483030/ | |
| 14:07:29 | gibi | stephenfin: looking... | |
| 14:09:37 | sean-k-mooney | jaypipes: yes the generic device manager thing + there are a few other spects related to sriov in general such as attach/detach support. | |
| 14:09:49 | jaypipes | yup | |
| 14:22:45 | sbezverk | sean-k-mooney : ping | |
| 14:24:20 | sean-k-mooney | sbezverk: hi | |
| 14:25:21 | sbezverk | sean-k-mooney hi, quick question for you if you have a second. Are you aware of any in depth libvirt debugging/troubleshooting guide? | |
| 14:25:29 | openstackgerrit | Merged openstack/os-traits master: Add a new parameter ``suffix`` to function ``get_traits`` https://review.openstack.org/483451 | |
| 14:26:08 | sean-k-mooney | sbezverk: not as such no. are you trying to find a guide to debug problems using libvirt or developing libvirt | |
| 14:27:54 | sbezverk | sean-k-mooney I see "x-files" type of problem with one VM and need to dig way deeper in libvirt to understand what is going on | |
| 14:29:01 | sean-k-mooney | well libvirt iteslef does almost nothing excepth take an xml and convert it into a qemu ars list | |
| 14:29:38 | sean-k-mooney | perhaps thats a bit unfair to it but usually the issue is not in libvirt but in the thing its calling of the thing that is driving it | |
| 14:30:25 | sean-k-mooney | sbezverk: what is the problem you are having? | |
| 14:30:47 | sbezverk | sean-k-mooney you asked yourself ;) | |
| 14:31:33 | sbezverk | sean-k-mooney one vm gets paused without any particular reason, I checked everything I could no issue with storage, plenty in vm and on the host | |
| 14:32:04 | sbezverk | run extensive diagnostic on the host memory/cpu/disk subsystem comes clean | |
| 14:32:45 | sbezverk | I wanted to see what process/even signals to libvirt to pause that vm | |
| 14:34:55 | sean-k-mooney | ok is there anyting in the instance log in /var/log/libvirt/qemu/instance-<id>.log | |
| 14:34:56 | sbezverk | even answering whether it is triggered outside on VM or some internal thing in the VM is problematic as no logs indicate anything abnormal | |
| 14:35:34 | sean-k-mooney | if libvirt recived a api call to pause the vm it should be in the libvirtd log in /var/log/libvirtd.log | |
| 14:36:45 | sbezverk | sean-k-mooney ok thanks I will double check, I enabled debug level of log for libvirt | |
| 14:52:03 | sbezverk | sean-k-mooney what do you think about this http://paste.openstack.org/show/620346/ | |
| 14:56:06 | sean-k-mooney | sbezverk: you would not happen to be trying to use a new intel purly plathform or new rayzen server would you | |
| 14:56:51 | sean-k-mooney | sbezverk: i have seen this before but only when you are trying to use a gust cpu model that uses feature not availabel on the host | |
| 14:57:09 | stephenfin | sean-k-mooney, sbezverk | |
| 14:57:09 | sbezverk | sean-k-mooney nope it is cisco ucs box with 2600 something v4 | |
| 14:57:26 | sean-k-mooney | such as when you are using an old kvm/qemu that does not understand the plathform your are on | |
| 14:57:30 | stephenfin | sean-k-mooney, sbezverk: That does look very similar. I'm pretty sure there was a patch up somewhere to do with this | |
| 14:57:32 | sean-k-mooney | stephenfin: yes? | |
| 14:57:41 | stephenfin | sean-k-mooney: Sorry - git return too early :P | |
| 14:57:45 | stephenfin | *hit | |
| 14:59:24 | sean-k-mooney | sbezverk: one thing you could do to rule out that the cpu model was the issue is enablle host passhtrough by seting cpu_mode = host-passthrough in the libvirt section of /etc/nova/nova.conf | |
| 14:59:25 | sbezverk | stephenfin : do you have a link by any chance? | |
| 14:59:44 | sbezverk | it is pass through | |
| 14:59:50 | sbezverk | sean-k-mooney ^^^ | |
| 15:00:05 | sean-k-mooney | really? huh that is strange indeed | |
| 15:00:59 | sean-k-mooney | sbezverk: well according to that log the host does not suport vmx or smx so you have not hadware virtualiation support | |
| 15:01:28 | sean-k-mooney | can you show me the output of cat /proc/cpuinfo and also check that /dev/kvm is presnet | |
| 15:01:34 | sbezverk | sean-k-mooney : :) there are 5 other vms happily running on the host | |
| 15:01:51 | sean-k-mooney | haha of course there is :) | |
| 15:02:45 | sbezverk | sean-k-mooney that is why I call it the x-files case ;) | |