| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-08-07 | |||
| 15:19:11 | efried | yeah, see, I think NUMA affinity Just Happens (tm) on Power. | |
| 15:19:12 | jaypipes | cdent: perhaps, yes | |
| 15:19:38 | sean-k-mooney | efried: i dont think vsphere exposes it either. its just libvirt that exposes it as far as i can tell | |
| 15:19:55 | efried | frickin libvirt <rolls eyes> | |
| 15:20:43 | sean-k-mooney | hehe well libvirt exposes it because libvirt is not a hypervior its an hypervior abstraction layer and isnce qemu does not do this by default it fell to nova to fill in the gapps | |
| 15:21:04 | efried | :) | |
| 15:22:08 | jaypipes | sean-k-mooney: WRONG! libvirt is an XML file management system. | |
| 15:22:40 | sean-k-mooney | jaypipes: lol | |
| 15:22:58 | sean-k-mooney | jaypipes: it does a little bit more then that but more or less | |
| 15:24:24 | sean-k-mooney | you know it could be argured that a new service could be insrted between nova and libvirt that did all the nfv stuff and the libvirt driver could be made way simpeler | |
| 15:25:24 | cdent | I sure hope those unicorns are free range | |
| 15:35:10 | s10 | lyarwood: with https://review.openstack.org/589513 update_available_resource() lasts 10 seconds for 100 instances instead of 20 seconds without. | |
| 15:39:13 | lyarwood | s10: kk, that's a single qemu-img info call per disk to avoid the other issues fixed by the original changes | |
| 15:39:52 | dansmith | s10: this sounds like a silly question, but what does it matter how long that takes? | |
| 15:39:57 | dansmith | it's not blocking other work right? | |
| 15:40:36 | efried | stephenfin: I think https://review.openstack.org/#/c/588422/ is ready for your +A now | |
| 15:40:46 | stephenfin | ack | |
| 15:40:53 | lyarwood | dansmith: I did ask above and it's causing issues with LM apparently | |
| 15:41:09 | dansmith | lyarwood: oh sorry I missed that.. maybe because it's holding the RT semaphore? | |
| 15:41:42 | lyarwood | dansmith: yes I think so | |
| 15:41:53 | dansmith | okay makes sense I guess | |
| 15:42:19 | dansmith | lyarwood: the new call checks the allocated value, which doesn't change over the lifecycle of the image right? | |
| 15:42:54 | sean-k-mooney | dansmith: correct it should not | |
| 15:43:26 | lyarwood | dansmith: allocated can, virtual shouldn't | |
| 15:43:26 | sean-k-mooney | dansmith: unless the call is checking the actul used space on the host and we are not preallocting ? | |
| 15:43:36 | lyarwood | yeah pretty much | |
| 15:43:53 | dansmith | lyarwood: but which are we looking at now? | |
| 15:44:09 | lyarwood | dansmith: now it's just a single qemu-img info call that grabs both | |
| 15:44:15 | sean-k-mooney | lyarwood: the resouce track should be tracking the virtual amont right not the currently allcoated amount? | |
| 15:44:33 | dansmith | lyarwood: ah, both so one of them is dynamic and we can't really cache it yeah? | |
| 15:44:34 | lyarwood | sean-k-mooney: the allocated amount is used to work out over commit etc | |
| 15:44:58 | sean-k-mooney | lyarwood: that seams wrong the maxium it could use should be used for that | |
| 15:45:30 | dansmith | sean-k-mooney: IIRC, this is not for reporting to placement but for some of the legacy values (right lyarwood ?) | |
| 15:45:32 | openstack | Launchpad bug 1784826 in OpenStack Compute (nova) "Guest remain in origin host after evacuate and unset force-down nova-compute" [Undecided,In progress] - Assigned to huanhongda (hongda) | |
| 15:45:32 | jaypipes | dansmith: what are your thoughts on https://bugs.launchpad.net/nova/+bug/1784826? I can't tell what the expected behaviour there should be... | |
| 15:45:40 | dansmith | because auditing on the compute node doesn't change what placement is counting | |
| 15:46:04 | lyarwood | dansmith: yeah correct, this doesn't make it to placement AFAIK | |
| 15:47:04 | dansmith | lyarwood: so we really only need to be collecting this info for old RT, which is ignored if you don't have DiskFilter enabled, and only useful for people using the deprecated CachingScheduler yeah? | |
| 15:47:28 | sean-k-mooney | lyarwood: what advantage is there to using allocated though? since placement would already be using its allocation ratio to filter the hosts before the compute ever need to check its over_subsciption ratio setting | |
| 15:47:54 | lyarwood | dansmith: in master but this series was backported all the way back to ocata | |
| 15:48:16 | lyarwood | dansmith: wasn't it still used back then? | |
| 15:48:29 | dansmith | lyarwood: yeah I know, but diskfilter hasn't been needed since claims in the scheduler | |
| 15:48:36 | dansmith | *maybe* ocata, but not pike or later IIRC | |
| 15:50:40 | lyarwood | sean-k-mooney: right, I think the only reason this came up before is that it's visable from the CLI | |
| 15:50:46 | dansmith | lyarwood: anyway, just wondering if maybe we should add a workaround config tweak to disable this extra inspection so you can turn it off if you're not using DiskFilter and/or are on a new enough version | |
| 15:51:04 | lyarwood | dansmith: yeah that sounds like the way to go tbh | |
| 15:51:19 | dansmith | lyarwood: ah, well, that makes it even more pain than gain if it was just "make the numbers line up" :) | |
| 15:52:04 | dansmith | lyarwood: probably best to get some sign-offs from the PTL(s) of the affected releases, but that's kinda where I'm thinking | |
| 15:52:14 | sean-k-mooney | lyarwood: so if its visable via the cli i think thats even more reason to use the virtual not allocated disk size to have it corralte with placement | |
| 15:52:53 | sean-k-mooney | lyarwood: that said that would be a behavior change i guess. | |
| 15:53:22 | sean-k-mooney | lyarwood: i assume this is visable via the hyperviors api? | |
| 15:54:34 | openstack | Launchpad bug 1764489 in OpenStack Compute (nova) queens "Preallocated disks are deducted twice from disk_available_least when using preallocated_images = space" [Medium,Fix committed] - Assigned to Lee Yarwood (lyarwood) | |
| 15:54:34 | lyarwood | sean-k-mooney: yeah via a hypervisor-show - https://bugs.launchpad.net/nova/+bug/1764489 | |
| 15:56:41 | sean-k-mooney | right so its messing up disk_available_least not local_gb_used... | |
| 16:01:44 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Use placement 1.28 in scheduler report client https://review.openstack.org/583667 | |
| 16:11:02 | mdbooth | lyarwood: I'm partially through the functional failure. Got past the shelve/unshelve failure, error is now in _test_attach_volume_error | |
| 16:11:51 | mdbooth | lyarwood: However, I'm going to leave shortly, it doesn't pass yet, and there are still print statements all over it | |
| 16:12:06 | melwitt | . | |
| 16:12:19 | mdbooth | lyarwood: So I'm not inclined to push it this evening unless that would be especially helpful to you | |
| 16:15:04 | lyarwood | mdbooth: yeah np leave it for now | |
| 16:50:48 | openstackgerrit | Sergii Golovatiuk proposed openstack/nova master: Fix URI for IPv6 https://review.openstack.org/589548 | |
| 16:57:56 | melwitt | dansmith: this review looks in your wheelhouse https://review.openstack.org/589425 | |
| 16:58:37 | dansmith | yeah, gibi and I have been discussing it | |
| 16:58:51 | dansmith | I'll hit it in a few | |
| 16:59:02 | melwitt | k | |
| 17:10:07 | efried | stephenfin, sean-k-mooney: I found out a little more about NUMA affinity on Power. It is indeed done dynamically; you can't ask for it to be done a certain way. However, the NUMA cell and affinity information *is* exposed to the guest, so if the guest has the savvy and ability to use such information for clever scheduling of threads/processes, it can do so. | |
| 17:11:39 | efried | stephenfin, sean-k-mooney: Also of note, the system may swizzle affinities dynamically for various reasons (including, apparently, at the behest of a generic "improve my affinity" background job you can run) so the guest needs to be able to hang with that. | |
| 17:26:22 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix resize revert to use non-legacy alloc handling https://review.openstack.org/589425 | |
| 17:38:07 | sean-k-mooney | efried: stephenfin so am i right in saying i can request a partcalar virtual numa topology for the guest but powervm will dynmically map that as it sees fit and adjust as needed | |
| 17:39:18 | sean-k-mooney | brb | |
| 18:25:28 | efried | sean-k-mooney: You can *not* request a particular topology for a guest. PowerVM will set it up dynamically and adjust as needed. | |
| 18:29:51 | openstackgerrit | Lee Yarwood proposed openstack/nova master: WIP libvirt: Reduce calls to qemu-img during update_available_resource https://review.openstack.org/589513 | |
| 18:29:52 | openstackgerrit | Lee Yarwood proposed openstack/nova master: WIP libvirt: Add workaround to stop use of qemu-img by the RT https://review.openstack.org/589567 | |
| 18:29:58 | lyarwood | dansmith: ^ was that what you had in mind for the qemu-img workaround? If it is I'll clean it up and ping the ML to see if others agree. | |
| 18:30:49 | dansmith | lyarwood: kinda, but we weren't just returning zero before right? what were we doing? | |
| 18:31:00 | dansmith | I was thinking just make the workaround trigger "do the old thing" | |
| 18:31:49 | lyarwood | dansmith: well before we were calling os.path.getsize | |
| 18:32:13 | dansmith | ah right, so .. I would just make the workaround revert to that I think | |
| 18:45:09 | sean-k-mooney | efried: does that imply that from within the guest i can see that the topoloyg is chaning? | |
| 18:45:27 | efried | sean-k-mooney: Yeah. Is that crazy? | |
| 18:45:32 | sean-k-mooney | efried: yes | |
| 18:45:53 | efried | sean-k-mooney: It's possible that the topo only changes if you run that job I mentioned. | |
| 18:46:01 | sean-k-mooney | windows and linux are really not going to do the right thing | |
| 18:46:42 | sean-k-mooney | im pretry sure that neither will check the numa topology on an ongoing basis | |
| 18:47:26 | sean-k-mooney | so the linux schduler is going to detect the numa topology when it start up and then contiue to use the same mappings for the liftime of the vm | |
| 18:51:06 | sean-k-mooney | if the hyperviors hide the phyical topology from the guest but dynamically rempas the ram across the host numa nodes then atleast the info present to the guest is consitent and it can try to make sane decisions but if the virtual tolopogy the guest sees is dynamic things like numactl are going to not work correctly in the guest | |
| 18:53:27 | efried | sean-k-mooney: Yeah, I don't know to this level of detail. I imagine it'd be in response to other LPARs going away or whatever, and thus freeing up the system to shuffle my stuff to make it more efficient. Or maybe I hot plug a device and now it's better to move my procs/mem to the cell that's affined to that device. But I don't know how that actually appears to the guest. | |
| 18:55:25 | sean-k-mooney | efried: well the guest view is what i was asking about when i said can you request a virtual numa toplogy. as long as the geust view does not change the hypervior can do what it like underneath without worriying it will break the guest | |
| 18:56:19 | sean-k-mooney | the guest can still make bad desisions if the hyperviour dicides to split a virutal numa node across host phyical numa nodes but it wont break anything | |
| 18:57:09 | sean-k-mooney | if the gust view is changing things like hugepages within the guest would break | |
| 18:57:20 | efried | sean-k-mooney: When you say "request a virtual numa topology" do you mean "ask what the topology is" or do you mean "request that the virtual topology look like XYZ" ? | |
| 18:57:40 | efried | s/look like/get set up as/ | |
| 18:58:50 | sean-k-mooney | i mean "power vm please create a vm that to the guest looks like it has 2 numa node, do whatever you like on the host" | |
| 18:59:25 | sean-k-mooney | that is effectivly what hw:numa_nodes=2 means | |
| 18:59:26 | efried | sean-k-mooney: Right, okay, so I'm saying you definitely can't do that in PowerVM. I'm not sure whether you can do it in PowerKVM. | |
| 18:59:42 | sean-k-mooney | ok cool | |
| 19:00:25 | sean-k-mooney | you can do that with hyperv i think but i dont think i guarntees they are mapped to 2 phyical host numa nodes | |
| 19:01:14 | sean-k-mooney | or rather if they are mapped/affinitised to host numa nodes is governed by a hypervior config option as far as i know | |
| 19:14:52 | sean-k-mooney | efried: its burried in the docs but ya it looks like you can do numa affinity with powerkvm https://www.ibm.com/support/knowledgecenter/SSZJY4_3.1.0/liabp/liabphotplugcpunuma.htm | |
| 19:15:11 | sean-k-mooney | you can also do cpu pinning | |