Earlier  
Posted Nick Remark
#openstack-nova - 2022-11-29
17:01:22 sean-k-mooney by limiting it to one process
17:01:31 clarkb or just bound your buffers
17:01:34 clarkb and read in chunks
17:01:54 sean-k-mooney maybe i havent really looked at the channel implemenation closely
17:02:07 clarkb I suspect this is a case of python makes it easy to read abritrary sized buffers into memroy without much fuss
17:02:24 clarkb it might also be inefficient compilation of the ruleset
17:02:28 clarkb (regexes aren't free either)
17:02:54 sean-k-mooney the impelmation is here https://github.com/openstack/oslo.privsep/blob/83870bd2655f3250bb5d5aed7c9865ba0b5e4770/oslo_privsep/comm.py
17:03:41 sean-k-mooney self.writesock.sendall(buf)
17:03:53 sean-k-mooney so its takeign the serialsed payload and sending it
17:04:13 sean-k-mooney using the msgpack format for serialisation
17:05:40 sean-k-mooney its using 4k buffers https://github.com/openstack/oslo.privsep/blob/83870bd2655f3250bb5d5aed7c9865ba0b5e4770/oslo_privsep/comm.py#L81
17:08:35 sean-k-mooney clarkb: honestly i have looked at privsep a couple of time but dont have enough context of the code to have a feel for how much memory it shoudl be using and if its bounded or not
17:09:02 dansmith also not sure what we might be calling via privsep that would return large buffers
17:09:04 sean-k-mooney clarkb: but i suspect that if we are using it ofr any file operatiosn then it might need to process guest images
17:09:14 dansmith it's mostly for doing things and maybe pulling the qemu-img info blob
17:09:23 bauzas sean-k-mooney: https://review.opendev.org/c/openstack/tempest/+/866049
17:09:27 sean-k-mooney dansmith: the console log or writing images to disk would be the two the come to mind
17:10:01 dansmith I don't think we write images to disk via privsep. console log.. maybe? I thought we can get that via libvirt
17:10:10 clarkb ya I'm not sure. Its just in aggregate privsep uses more memory than most other openstack services
17:10:30 sean-k-mooney dansmith: no we read the file that is written to disk im pretty sure
17:10:34 clarkb I think cinder? and neutron use more then privsep is next. Its been a while since I looked at hte memory profiling though
17:11:01 bauzas I don't know if we could somehow pdb the running privsep process thru a backdoor, because if we could, like we do with nova services, then we could monitor the growing memory
17:11:24 sean-k-mooney dansmith: i would hope for the images that we write it to somewhere we own then move it and change the permission if needed
17:11:33 bauzas I personnally use tracemalloc to persist the memory state and compare between snapshots
17:11:51 sean-k-mooney like in most case i woudl expect nova to put it in the image cache then use qemu-image to create a qcow with the iamge as the backing file
17:11:52 bauzas but this requires access to the process
17:12:16 sean-k-mooney so privsep shoudl only be needed for invokeign qemu-img and not the actul image downlaod
17:12:29 sean-k-mooney but not sure about the same codepath for raw images
17:13:11 sean-k-mooney bauzas: you coudl do that locally but i dont think privsep uses eventlet
17:13:35 dansmith sean-k-mooney: we can't possibly be streaming multiple gigs of image data over the privsep socket
17:13:35 bauzas sean-k-mooney: afaik it doesn't
17:13:57 sean-k-mooney dansmith: ya that woudl not make sense to me eitehr
17:14:16 sean-k-mooney i know the console log is stream over it as we had a security issue with that in the past
17:14:24 sean-k-mooney but that is small
17:15:24 dansmith well, that might be large, but not tens of gigs
17:15:53 dansmith and still not sure why we're doing that really, but AFAIK nova has to eat the whole thing and return it over RPC anyway, right?
17:15:57 sean-k-mooney in production yes in ci the vms dont run that long so i would be surpsied if we ever go close to a mb
17:16:35 dansmith oh sure
17:17:07 sean-k-mooney https://github.com/openstack/nova/blob/master/nova/config.py#L78-L80
17:17:31 sean-k-mooney so in produciton we drop privesep abck to info
17:17:50 clarkb Looking at https://6546e150b851cf607d58-90bcb133e16a34d804e19428ec130682.ssl.cf5.rackcdn.com/865479/1/gate/tempest-full-py3/9076120/controller/logs/screen-memory_tracker.txt there are three different nova-cpu privsep processes and together use about 150MB of rss. Similar story for neutron ovn metadata. Buts its ~250MB rss
17:17:51 sean-k-mooney im not sure if we override that in ci
17:18:31 sean-k-mooney clarkb: two fo the prive sepc process i can accoutn for
17:19:04 sean-k-mooney nova has one privsep context for everything and os-vif has one for the linux bridge and ovs plugin but i only expect 1 of those two to be created in any given ci job
17:19:05 clarkb sean-k-mooney: why do we need multiple process per service though?
17:19:16 sean-k-mooney we need one per security context
17:19:43 sean-k-mooney and we shoudl have mulitpel security context per service
17:19:46 clarkb sean-k-mooney: but if your security contro lmehtod is file permissions (any maybe selinux rules) then you can only limit by pid user or pid selinux context
17:20:09 clarkb you aren't any more secure if you have three different processes that nova cpu can talk to
17:20:13 sean-k-mooney privesp drops permissions to only does in the context when it spawns
17:20:45 sean-k-mooney clarkb: if you have muplie context you can limit the permission of each funciton
17:21:21 clarkb sean-k-mooney: but if the same entity can tell each of those three to do the work then functioanlly its the same context
17:21:27 sean-k-mooney that is the permission model that privsep provides and the big delta form rootwarp
17:21:32 clarkb you're limited by what you can limit access to on the filesystem
17:21:46 clarkb and if I have permission to talk to one I have permisssion to talk to all since a process shares those attributes
17:22:16 clarkb and if I somehow break that security layer to promote myself to nova-cpu user or selinux context I can talk to all of them
17:22:18 sean-k-mooney you have to decorate the funciton with the context they can use
17:22:40 clarkb right but does that give the system any additional security since it sounds like we are using fielsystem controls I think not
17:22:50 clarkb the level of scrutiny is process attribute (user, selinux context etc)
17:22:59 sean-k-mooney we have had this dusicssion multiple time
17:23:15 sean-k-mooney prive sep is about limiting the permission of indivual function calls
17:23:28 sean-k-mooney it is not intended to adress all security concerns
17:24:03 sean-k-mooney in many respec rootwarp was more secure
17:24:16 sean-k-mooney but it was unmaintained and slow
17:24:22 clarkb ok, the result i that ~5% of system memory on a tempest job run is consumed by privsep for nova cpu and neutron ovn metadata
17:25:18 sean-k-mooney the nova_ovn_metadata on is kind fo interesting it will need cap net admin for setting the iptables rules to nat the traffic
17:25:32 sean-k-mooney but it shoudl not need much memory for that
17:25:44 bauzas I remember some internal customer bug about the metadata service being greedy
17:25:46 clarkb nova cpu uses about 200MB of memory compared to 150 for privsep for nova cpu
17:25:48 clarkb as a comparison
17:25:57 clarkb it essentially doubles the memory footprint of nova cpu
17:26:17 bauzas we tried to look into it, but given the issue was not reproducable after a reboot, we were unable to conclude
17:26:26 sean-k-mooney yep it has to load many of the same nova depencies
17:26:44 sean-k-mooney it is runnign part of the nova codebase after all
17:27:11 sean-k-mooney bauzas: the customer issue was becasue they had 0 swap
17:27:25 clarkb sure, but that is why I remember it being outsized. Seems like the current memory tracking continues to show this
17:27:34 sean-k-mooney so the python interperty was using more memory then when swap was avaiable
17:28:19 clarkb mysqld then journald are the top consumers. No surprise there given how mysql uses memory and the quantity of logs we are dumping into the journal
17:28:39 sean-k-mooney i wonder if there is any way to have privsep free memeory
17:28:59 sean-k-mooney i dont know if it suffers form the same fragmentation thing that cinder? glance? had
17:29:04 bauzas sean-k-mooney: no, unrelated case
17:29:22 clarkb looks like cinder-backup is also running in he base job. I thought there was no testing of it in he base jobs and we were removing it
17:29:26 sean-k-mooney oh sorry metadata not nova-compute
17:30:11 sean-k-mooney bauzas: ya metadata memory usage is varible
17:30:37 sean-k-mooney we need to build the fully metadta respocen in memory for every request
17:30:43 sean-k-mooney which is why we store it in memcache
17:31:29 bauzas anyway, I need to quit, my wife feels ill and I need to visit the drugstore
17:31:41 sean-k-mooney im not sure exactuly how we handel user-data and if that is incldued in the cached value or if we stream that seperatly
17:31:55 clarkb in that example job we have about 100MB free memory when the log file is collected
17:32:04 sean-k-mooney the user-data file can be up to 64mb
17:32:23 clarkb and we are using about 600MB of swap fi I read that correctly
17:32:40 clarkb (out of about 1GB of swap total)
17:41:18 clarkb other than mysqld there isn't any single process getting close to dobule digit memory consumption. This is deaht by a thousand cuts (whcih makes sense given micro services etc)
17:42:20 dansmith I'm pretty sure "death by a thousand cuts" is the actual marketing tagline of the microservice approach :)
17:44:40 clarkb oh wait memavailable is distinct to memfree. Its closer to 750MB of memory available. That isn't too bad. I wonder why we're so deep into swap then
17:49:43 sean-k-mooney free is actully unacllocated
17:50:00 sean-k-mooney avaiable include cachces and maybe buffers
19:25:57 dansmith sean-k-mooney: bauzas if you're still around, have you seen anything from rajat about changing nova to use a service account to look at image locations?
19:52:26 opendevreview Merged openstack/nova-specs master: spec: allowing target state for evacuate https://review.opendev.org/c/openstack/nova-specs/+/857838

Earlier   Later