| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-29 | |||
| 17:13:11 | sean-k-mooney | bauzas: you coudl do that locally but i dont think privsep uses eventlet | |
| 17:13:35 | dansmith | sean-k-mooney: we can't possibly be streaming multiple gigs of image data over the privsep socket | |
| 17:13:35 | bauzas | sean-k-mooney: afaik it doesn't | |
| 17:13:57 | sean-k-mooney | dansmith: ya that woudl not make sense to me eitehr | |
| 17:14:16 | sean-k-mooney | i know the console log is stream over it as we had a security issue with that in the past | |
| 17:14:24 | sean-k-mooney | but that is small | |
| 17:15:24 | dansmith | well, that might be large, but not tens of gigs | |
| 17:15:53 | dansmith | and still not sure why we're doing that really, but AFAIK nova has to eat the whole thing and return it over RPC anyway, right? | |
| 17:15:57 | sean-k-mooney | in production yes in ci the vms dont run that long so i would be surpsied if we ever go close to a mb | |
| 17:16:35 | dansmith | oh sure | |
| 17:17:07 | sean-k-mooney | https://github.com/openstack/nova/blob/master/nova/config.py#L78-L80 | |
| 17:17:31 | sean-k-mooney | so in produciton we drop privesep abck to info | |
| 17:17:50 | clarkb | Looking at https://6546e150b851cf607d58-90bcb133e16a34d804e19428ec130682.ssl.cf5.rackcdn.com/865479/1/gate/tempest-full-py3/9076120/controller/logs/screen-memory_tracker.txt there are three different nova-cpu privsep processes and together use about 150MB of rss. Similar story for neutron ovn metadata. Buts its ~250MB rss | |
| 17:17:51 | sean-k-mooney | im not sure if we override that in ci | |
| 17:18:31 | sean-k-mooney | clarkb: two fo the prive sepc process i can accoutn for | |
| 17:19:04 | sean-k-mooney | nova has one privsep context for everything and os-vif has one for the linux bridge and ovs plugin but i only expect 1 of those two to be created in any given ci job | |
| 17:19:05 | clarkb | sean-k-mooney: why do we need multiple process per service though? | |
| 17:19:16 | sean-k-mooney | we need one per security context | |
| 17:19:43 | sean-k-mooney | and we shoudl have mulitpel security context per service | |
| 17:19:46 | clarkb | sean-k-mooney: but if your security contro lmehtod is file permissions (any maybe selinux rules) then you can only limit by pid user or pid selinux context | |
| 17:20:09 | clarkb | you aren't any more secure if you have three different processes that nova cpu can talk to | |
| 17:20:13 | sean-k-mooney | privesp drops permissions to only does in the context when it spawns | |
| 17:20:45 | sean-k-mooney | clarkb: if you have muplie context you can limit the permission of each funciton | |
| 17:21:21 | clarkb | sean-k-mooney: but if the same entity can tell each of those three to do the work then functioanlly its the same context | |
| 17:21:27 | sean-k-mooney | that is the permission model that privsep provides and the big delta form rootwarp | |
| 17:21:32 | clarkb | you're limited by what you can limit access to on the filesystem | |
| 17:21:46 | clarkb | and if I have permission to talk to one I have permisssion to talk to all since a process shares those attributes | |
| 17:22:16 | clarkb | and if I somehow break that security layer to promote myself to nova-cpu user or selinux context I can talk to all of them | |
| 17:22:18 | sean-k-mooney | you have to decorate the funciton with the context they can use | |
| 17:22:40 | clarkb | right but does that give the system any additional security since it sounds like we are using fielsystem controls I think not | |
| 17:22:50 | clarkb | the level of scrutiny is process attribute (user, selinux context etc) | |
| 17:22:59 | sean-k-mooney | we have had this dusicssion multiple time | |
| 17:23:15 | sean-k-mooney | prive sep is about limiting the permission of indivual function calls | |
| 17:23:28 | sean-k-mooney | it is not intended to adress all security concerns | |
| 17:24:03 | sean-k-mooney | in many respec rootwarp was more secure | |
| 17:24:16 | sean-k-mooney | but it was unmaintained and slow | |
| 17:24:22 | clarkb | ok, the result i that ~5% of system memory on a tempest job run is consumed by privsep for nova cpu and neutron ovn metadata | |
| 17:25:18 | sean-k-mooney | the nova_ovn_metadata on is kind fo interesting it will need cap net admin for setting the iptables rules to nat the traffic | |
| 17:25:32 | sean-k-mooney | but it shoudl not need much memory for that | |
| 17:25:44 | bauzas | I remember some internal customer bug about the metadata service being greedy | |
| 17:25:46 | clarkb | nova cpu uses about 200MB of memory compared to 150 for privsep for nova cpu | |
| 17:25:48 | clarkb | as a comparison | |
| 17:25:57 | clarkb | it essentially doubles the memory footprint of nova cpu | |
| 17:26:17 | bauzas | we tried to look into it, but given the issue was not reproducable after a reboot, we were unable to conclude | |
| 17:26:26 | sean-k-mooney | yep it has to load many of the same nova depencies | |
| 17:26:44 | sean-k-mooney | it is runnign part of the nova codebase after all | |
| 17:27:11 | sean-k-mooney | bauzas: the customer issue was becasue they had 0 swap | |
| 17:27:25 | clarkb | sure, but that is why I remember it being outsized. Seems like the current memory tracking continues to show this | |
| 17:27:34 | sean-k-mooney | so the python interperty was using more memory then when swap was avaiable | |
| 17:28:19 | clarkb | mysqld then journald are the top consumers. No surprise there given how mysql uses memory and the quantity of logs we are dumping into the journal | |
| 17:28:39 | sean-k-mooney | i wonder if there is any way to have privsep free memeory | |
| 17:28:59 | sean-k-mooney | i dont know if it suffers form the same fragmentation thing that cinder? glance? had | |
| 17:29:04 | bauzas | sean-k-mooney: no, unrelated case | |
| 17:29:22 | clarkb | looks like cinder-backup is also running in he base job. I thought there was no testing of it in he base jobs and we were removing it | |
| 17:29:26 | sean-k-mooney | oh sorry metadata not nova-compute | |
| 17:30:11 | sean-k-mooney | bauzas: ya metadata memory usage is varible | |
| 17:30:37 | sean-k-mooney | we need to build the fully metadta respocen in memory for every request | |
| 17:30:43 | sean-k-mooney | which is why we store it in memcache | |
| 17:31:29 | bauzas | anyway, I need to quit, my wife feels ill and I need to visit the drugstore | |
| 17:31:41 | sean-k-mooney | im not sure exactuly how we handel user-data and if that is incldued in the cached value or if we stream that seperatly | |
| 17:31:55 | clarkb | in that example job we have about 100MB free memory when the log file is collected | |
| 17:32:04 | sean-k-mooney | the user-data file can be up to 64mb | |
| 17:32:23 | clarkb | and we are using about 600MB of swap fi I read that correctly | |
| 17:32:40 | clarkb | (out of about 1GB of swap total) | |
| 17:41:18 | clarkb | other than mysqld there isn't any single process getting close to dobule digit memory consumption. This is deaht by a thousand cuts (whcih makes sense given micro services etc) | |
| 17:42:20 | dansmith | I'm pretty sure "death by a thousand cuts" is the actual marketing tagline of the microservice approach :) | |
| 17:44:40 | clarkb | oh wait memavailable is distinct to memfree. Its closer to 750MB of memory available. That isn't too bad. I wonder why we're so deep into swap then | |
| 17:49:43 | sean-k-mooney | free is actully unacllocated | |
| 17:50:00 | sean-k-mooney | avaiable include cachces and maybe buffers | |
| 19:25:57 | dansmith | sean-k-mooney: bauzas if you're still around, have you seen anything from rajat about changing nova to use a service account to look at image locations? | |
| 19:52:26 | opendevreview | Merged openstack/nova-specs master: spec: allowing target state for evacuate https://review.opendev.org/c/openstack/nova-specs/+/857838 | |
| 20:44:57 | sean-k-mooney | dansmith: i have not seen that but i dont think nova would need to do anything would it | |
| 20:45:11 | sean-k-mooney | the account we have in the glance section just needs to have the service role | |
| 20:45:25 | dansmith | sean-k-mooney: we have no such account | |
| 20:45:49 | sean-k-mooney | oh ok because we currently dont use admin for this | |
| 20:46:42 | dansmith | right, everything we do with glance is with the user's token | |
| 20:46:47 | dansmith | this is a pretty fundamental change from that | |
| 20:47:27 | sean-k-mooney | ya ok im not sure if the service_user stuff woudl help there although thats really just for working aound token exeriation | |
| 20:48:20 | dansmith | sean-k-mooney: it would require it | |
| 20:48:25 | sean-k-mooney | this would really be that large a change in nova either a specless blueprint or short spec i guess | |
| 20:48:27 | dansmith | sean-k-mooney: maybe just read the spec :) | |
| 20:48:42 | sean-k-mooney | oh we have a spec already then sure | |
| 20:49:00 | dansmith | sean-k-mooney: no, the spec on the glance side I copied you on | |
| 20:49:01 | sean-k-mooney | the current service user config we have is not for that however | |
| 20:49:03 | dansmith | no spec on the nova side, | |
| 20:49:18 | sean-k-mooney | ah righit i saw the email havent looked at it yet | |
| 20:49:26 | sean-k-mooney | https://review.opendev.org/c/openstack/glance-specs/+/863209 | |
| 20:49:31 | sean-k-mooney | the new locations api spec | |
| 20:49:43 | dansmith | but since everything in nova assumes that the user's token is used for glance interaction, I just want to be super careful that we don't accidentally use the service account for anything related to the image other than the ceph location thing | |
| 20:50:50 | sean-k-mooney | ya so the same way we have for geting an admin client for neutron we likely need to add a get_service_client funtion or similar and use it for that call explcitly | |
| 20:51:41 | dansmith | right | |
| 20:52:07 | dansmith | we discussed this earlier related to the dual internal/external glance thing was brought up | |
| 20:52:22 | dansmith | and I thought you had a good reason for why there's a gotcha there, but I don't remember what it was | |
| 20:53:19 | sean-k-mooney | the internal endpoint is also used to provide unmeetered acces to the api for tenant workloads | |
| 20:53:21 | sean-k-mooney | in public clouds | |
| 20:53:40 | sean-k-mooney | so really it woudl be nice if keystone addded a service endpoint | |
| 20:53:48 | sean-k-mooney | for service to service comunications | |
| 20:54:02 | sean-k-mooney | although if this requried the service user | |
| 20:54:02 | dansmith | yeah, that's unrelated to this | |
| 20:54:04 | dansmith | I mean, this is to avoid needing that | |