Earlier  
Posted Nick Remark
#openstack-nova - 2019-03-19
17:13:27 aspiers sean-k-mooney: we got Brijesh from AMD to respond to your hugepages suggestion
17:13:37 sean-k-mooney mdbooth: stephenfin so yes this only happens on python3.6 which is why it passes on fedora and not on ubuntu or centos which default to python 3.6
17:13:50 mdbooth Well, at least they're all in the same place, now
17:13:53 sean-k-mooney aspiers: oh ok i havent read it yet
17:14:06 sean-k-mooney aspiers: was it favorable or not?
17:14:20 aspiers sean-k-mooney: I don't think so. Bottom of https://review.openstack.org/#/c/641994/2
17:15:02 aspiers SEV by design prevents the kernel copying encrypted pages around
17:15:34 aspiers he says using hugepages would pin it within the hugepages pool, but not pin each page to a single address
17:15:44 aspiers I'm on a call with AMD now
17:16:48 sean-k-mooney aspiers: ya dan berrange replied and said hugepages may not be sufficent too as the roms and uefi pflash would not be pinned
17:16:55 aspiers right
17:18:04 aspiers sean-k-mooney: how does it work for device passthrough?
17:18:47 aspiers apparently that's a similar scenario in terms of pinning
17:20:34 sean-k-mooney the iommu maps the device mmio region to a fixed set of addresses that are maped to the memory regions listed in the pci VF config space i think. but honestly i tought that a hugepage was similarly fixed and could not migrate when we use explcit not transparent hugepages
17:21:30 mdbooth It seems there's no need for a double-import guard unless you borked the python path, because python doesn't appear to import a second time anyway
17:21:39 mdbooth stephenfin: Is ^^^ your understanding?
17:21:53 sean-k-mooney mdbooth: we dont need the import guard correct
17:22:20 stephenfin mdbooth: That sounds reasonable, aye
17:22:31 stephenfin Based on my understanding
17:22:55 sean-k-mooney mdbooth: the issue is we never monkeypatched the docs before and under python 3.6 only it breaks things if we do. apparently
17:24:20 mdbooth Incidentally, that likely only breaks because we didn't patch earlier...
17:24:34 mdbooth Perhaps we should update sphinx to always monkey patch ;)
17:25:38 sean-k-mooney mdbooth: you joke but we could proably write a tiny sphinx extenion to monky patch earlir. that said didnt you plane to patch in nova/__init__.py
17:26:34 sean-k-mooney oh your not doing that in that patch
17:27:29 mdbooth sean-k-mooney: Yeah, that breaks everything for different reasons.
17:27:58 aspiers sean-k-mooney: yeah, I get the impression Brijesh (AMD) thinks otherwise but TBH I'm out of my depth here :)
17:27:59 sean-k-mooney ah ok :) well it is eventlets
17:28:09 aspiers definitely not my area of expertise
17:28:32 sean-k-mooney aspiers: well we still lock the memory expiclitly when we use hugepages
17:28:45 aspiers ah
17:28:53 sean-k-mooney i.e. we set the locked element https://libvirt.org/formatdomain.html#elementsMemoryBacking
17:29:08 openstackgerrit Matthew Booth proposed openstack/nova master: Eventlet monkey patching should be as early as possible https://review.openstack.org/626952
17:29:10 mdbooth Thar she blows.
17:29:18 sean-k-mooney or maybe we only do that for realtime i can check
17:29:28 mdbooth Lets see if that passes muster.
17:30:03 aspiers sean-k-mooney: ah OK, is that also done for device pass-through?
17:30:47 sean-k-mooney no its not done for device pass-though
17:32:37 sean-k-mooney aspiers: this is the only time we current lock the memory explcitly whic is for realtime guests
17:32:39 sean-k-mooney http://git.openstack.org/cgit/openstack/nova/tree/nova/virt/libvirt/driver.py#n4810
17:34:20 sean-k-mooney the way dpdk work i am prtty sure the hugepage does not migrate to different phyical pages on the host as it is used for dma transfer for the vhost-user nics
17:34:54 sean-k-mooney aspiers: in anycase if you have hardware you can test it on it would be worth testing with hugepages
17:36:04 sean-k-mooney aspiers: but if dan's and brijesh's assertion regarding the rom+pflash issue are correct it may not be enough
17:36:04 aspiers sean-k-mooney: AMD already tried hugepages in one of their earliest iterations
17:36:15 aspiers that's what I just heard
17:36:49 aspiers they are also saying that mlock() is just a hint not a guarantee
17:37:31 sean-k-mooney ok well in that case i dont know what the best path forword is on the hardlimit issue
17:37:55 aspiers one suggestion was to expose the right hardlimit via QEMU / libvirt
17:38:01 aspiers so nova could just query it
17:39:29 sean-k-mooney aspiers: if qemu/libvirt can provide it that sounds resonable but it would make the relevent libvirt/qemu version that provides that the minium qemu/libivrt for that feature
17:39:47 aspiers correct, there's already a minimum version requirement anyway
17:40:06 aspiers I think everyone is aware of the issue now though, so hopefully Dan and Brijesh can figure something out :)
17:45:47 sean-k-mooney aspiers: so reading dan's comment again if we just remove the hard_limmit entirly then that may be suffienct to adress his orignial comment here https://review.openstack.org/#/c/641994/2/specs/train/approved/amd-sev-libvirt-support.rst@167
17:46:32 sean-k-mooney ah but that wont work as that is how you are pinning the memory
17:46:47 aspiers right
17:49:15 aspiers bryan_stephenson: catch up on the conversation via http://eavesdrop.openstack.org/irclogs/%23openstack-nova/%23openstack-nova.2019-03-19.log.html#t2019-03-19T17:13:27
17:50:22 sean-k-mooney aspiers: honestly looking at the documentation of hard_limit i think relying on it for the behavior you desire is undocumented and not part of the contract at least form the libvirt perspective
17:50:56 sean-k-mooney aspiers: hard_limit does not guarntee teh meoeoy will not be swapped or that it will be preallocated
17:51:37 sean-k-mooney aspiers: thos poperties are conntoeld via allocation and locked in the memory backing
17:53:20 bbobrov lets just ask on the libvirt mailing list
17:53:31 bbobrov (hi)
17:53:37 aspiers I agree with bbobrov
17:53:46 aspiers this probably requires libvirt changes
17:54:35 aspiers bbobrov: Do you want to kick off the thread there?
17:54:39 sean-k-mooney bbobrov: aspiers sure let us know what they say
17:55:25 bbobrov aspiers: yep, i will do it, but starting tomorrow
17:55:28 aspiers we can just link to the spec review and say "please find us a solution" :-)
17:55:33 aspiers bbobrov: great thanks
17:55:33 sean-k-mooney bbobrov: aspiers but as i said the behavior that is being asserted for hard_limit. i.e. that it prevents the page form migrating is not documented in teh libvirt docs https://libvirt.org/formatdomain.html#elementsMemoryTuning
17:55:42 aspiers sean-k-mooney: ack
17:56:16 aspiers https://libvirt.org/formatdomain.html#elementsMemoryBacking seems to be the bit for pinning but if that just does an mlock() then apparently it's not enough
17:56:23 aspiers from what I heard earlier
17:56:36 aspiers although that seems strange to me, since then why would it be good enough for realtime?
17:56:57 aspiers sean-k-mooney: real-time also needs a hard guarantee, right?
17:57:06 bbobrov why have we decided to use memtune for pinning in the first place?
17:57:22 aspiers because that's what the libvirt/QEMU guys recommended initially
17:57:27 bbobrov if there is no direct indication that it actually pins memory
17:57:35 aspiers they said it did
17:57:40 bbobrov is there a mailing list thread about it?
17:57:47 aspiers yes, on the internal list
17:58:02 sean-k-mooney aspiers: not entirly. for minium latency variance yes but it wont break the feature if there was host page migration in the background but it may break SLAs
17:58:08 bbobrov eh, internal lists
17:58:20 aspiers bbobrov: our sev list
17:58:39 bbobrov yeah, i got it, still ba
17:58:41 bbobrov *bad
17:59:10 sean-k-mooney bbobrov: well this is what the amdese repo does https://github.com/AMDESE/AMDSEV/blob/master/xmls/sample.xml#L218
17:59:17 sean-k-mooney it jsue set hard_limit
17:59:57 sean-k-mooney there is no other memory tuning or pinning but i would not have assumed that pinned any memory
18:00:43 aspiers "The value of the domain/memtune/hard_limit element will be used to setrlimit(RLIMIT_MEMLOCK, hard_limit) on the qemu process and in memory.limit_in_bytes setting of the processes memory controller (/sys/fs/cgroup/memory/machine.slice/machine-<vmid><vmname>.scope/memory.limit_in_bytes)"
18:00:56 aspiers that's from our libvirt guy
18:01:28 aspiers also https://libvirt.org/git/?p=libvirt.git;a=blob;f=src/qemu/qemu_domain.c;h=ba3fff607a93533b9b47956cc2cfa70237e7c041;hb=HEAD#l10134
18:02:35 gibi mriedem: I've replied in https://review.openstack.org/#/c/640390/5/doc/source/admin/config-qos-min-bw.rst@98
18:07:26 bryan_stephenson Accoring to https://libvirt.org/formatdomain.html#elementsMemoryBacking setting "locked" reserves/locks the physical memory. One use is for DMA, which would only be useful if the memory did not move. Do we think that memory might move locations even if it is locked?
18:07:49 aspiers bryan_stephenson: that's what Brijesh said on the call. he said mlock(2) was just a hint
18:08:02 aspiers the man page makes it sound like more than that, but I don't know
18:08:15 bryan_stephenson Then how does DMA not get goofed up?
18:08:33 sean-k-mooney bryan_stephenson we use hugepage memory for dma trasfer with dpdk so i think either hugepage or setting locked expclitly would be suffient
18:08:45 aspiers Honestly it's beyond my paygrade ;-)
18:09:20 bryan_stephenson So the "locked" just calls mlock() which may not be sufficient. Is that a correct understanding?
18:09:34 aspiers That's what I think I heard from Brijesh
18:12:06 sean-k-mooney aspiers: so looking at that funtion it runtrun Unlimited if the memory is locked to this would not caluatle a valide limit for realtime guests

Earlier   Later