| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-29 | |||
| 16:40:33 | bauzas | sean-k-mooney: that's simple to do and that's like 4 weeks I promised it | |
| 16:41:13 | bauzas | sean-k-mooney: you know what ? I'll end this meeting by now so everyone can do what they want, including me writing a zuul patch :) | |
| 16:41:25 | sean-k-mooney | :) | |
| 16:41:35 | bauzas | having said it, | |
| 16:41:39 | bauzas | thanks folks | |
| 16:41:43 | opendevmeet | Log: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.log.html | |
| 16:41:43 | opendevmeet | Minutes (text): https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.txt | |
| 16:41:43 | opendevmeet | Minutes: https://meetings.opendev.org/meetings/nova/2022/nova.2022-11-29-16.00.html | |
| 16:41:43 | opendevmeet | Meeting ended Tue Nov 29 16:41:43 2022 UTC. Information about MeetBot at http://wiki.debian.org/MeetBot . (v 0.1.4) | |
| 16:41:43 | bauzas | #endmeeting | |
| 16:42:03 | gibi | o/ | |
| 16:42:22 | chateaulav | o/ | |
| 16:42:29 | elodilles | o/ | |
| 16:44:07 | bauzas | like, https://zuul.openstack.org/job/tempest-centos9-stream-fips | |
| 16:48:28 | clarkb | sean-k-mooney: one thing I've found is fairly consistent for our systemic job timeouts is that we've got a number of steps to each job: base setup (configuring mirrors, configuring ssh keys, configuring git repos), test env setup (tox/devstack/whatever), actual testing, and log collection. Each of these tends to have unnecessary slow bits that add up over the course of a job. | |
| 16:48:30 | clarkb | When you then run into something like a slower node or swapping or slowness fetching an external resource it is very easy to tip over the timeout | |
| 16:49:00 | clarkb | sean-k-mooney: imo it would be helpful for us to try and whittle away at that accumulated slowness tech debt to make us less susceptible when we run into an overall slower situation. | |
| 16:49:27 | clarkb | I've worked on that a bit in the base jobs and log collection side of things as that has broad impact. But the downsides here are that it has broad impact so I have to be extremely careful to maintain backward compatibility | |
| 16:49:44 | clarkb | but the same approaches can be taken to improve things like devstack (did you know it installs tempest 3 times!) | |
| 16:50:26 | clarkb | I think improving memory consumption would also help avoid slowness caused by swapping. privsep is a fairly outsized offender here | |
| 16:53:16 | bauzas | clarkb: I'm curious about privsep being memory greedy | |
| 16:53:30 | bauzas | and I wonder why | |
| 16:53:43 | clarkb | I suspect because it grows buffers to handle all the input and output sent through it | |
| 16:54:05 | clarkb | one way to maybe improve things is to stop running a different privsep for each service whcih creates a bunch of large buffers. We might be able to get away with one large buffer | |
| 16:54:05 | bauzas | our internal customers haven't reported such problem, but I guess because of lack of evidence rather than not having a problem | |
| 16:54:29 | clarkb | or buffer things with intentionally smaller buffers | |
| 16:54:40 | bauzas | agreed | |
| 16:54:48 | bauzas | a stream is costly | |
| 16:57:01 | sean-k-mooney | clarkb: yep although we have enough fo a buffer in the nova project that we get a time out failure only 1 or twice a week | |
| 16:57:18 | clarkb | sean-k-mooney: yes, but you've also set your timeout to two hours | |
| 16:57:36 | clarkb | (one hour was the goal once upon a time) | |
| 16:57:50 | sean-k-mooney | ack | |
| 16:58:03 | sean-k-mooney | shareing privesep is a security issue | |
| 16:58:09 | sean-k-mooney | so i dont think we can ever do that | |
| 16:58:22 | sean-k-mooney | nova will have more privespe deamons eventurlaly | |
| 16:59:16 | clarkb | does privsep prevent random processes from connecting to it? If not this is equivalent. If so it could also apply restrictions on what a specific process can do (granted this is maybe a larger attack surface than refusing to talk at all) | |
| 16:59:42 | sean-k-mooney | you are ment to use file permissiosn to limit access at the file system level | |
| 16:59:53 | sean-k-mooney | but no | |
| 17:00:00 | sean-k-mooney | not as part of privsep itself | |
| 17:00:25 | sean-k-mooney | clarkb: you need one privsep deamon per privsep context currently | |
| 17:00:42 | sean-k-mooney | for it to proplry provide the correct permission enforcement/escalation | |
| 17:00:57 | clarkb | sean-k-mooney: in that case my suggestion would be to investigate optimizing privsep memory usage | |
| 17:01:14 | sean-k-mooney | we might be able to reduce it in ci | |
| 17:01:22 | sean-k-mooney | by limiting it to one process | |
| 17:01:31 | clarkb | or just bound your buffers | |
| 17:01:34 | clarkb | and read in chunks | |
| 17:01:54 | sean-k-mooney | maybe i havent really looked at the channel implemenation closely | |
| 17:02:07 | clarkb | I suspect this is a case of python makes it easy to read abritrary sized buffers into memroy without much fuss | |
| 17:02:24 | clarkb | it might also be inefficient compilation of the ruleset | |
| 17:02:28 | clarkb | (regexes aren't free either) | |
| 17:02:54 | sean-k-mooney | the impelmation is here https://github.com/openstack/oslo.privsep/blob/83870bd2655f3250bb5d5aed7c9865ba0b5e4770/oslo_privsep/comm.py | |
| 17:03:41 | sean-k-mooney | self.writesock.sendall(buf) | |
| 17:03:53 | sean-k-mooney | so its takeign the serialsed payload and sending it | |
| 17:04:13 | sean-k-mooney | using the msgpack format for serialisation | |
| 17:05:40 | sean-k-mooney | its using 4k buffers https://github.com/openstack/oslo.privsep/blob/83870bd2655f3250bb5d5aed7c9865ba0b5e4770/oslo_privsep/comm.py#L81 | |
| 17:08:35 | sean-k-mooney | clarkb: honestly i have looked at privsep a couple of time but dont have enough context of the code to have a feel for how much memory it shoudl be using and if its bounded or not | |
| 17:09:02 | dansmith | also not sure what we might be calling via privsep that would return large buffers | |
| 17:09:04 | sean-k-mooney | clarkb: but i suspect that if we are using it ofr any file operatiosn then it might need to process guest images | |
| 17:09:14 | dansmith | it's mostly for doing things and maybe pulling the qemu-img info blob | |
| 17:09:23 | bauzas | sean-k-mooney: https://review.opendev.org/c/openstack/tempest/+/866049 | |
| 17:09:27 | sean-k-mooney | dansmith: the console log or writing images to disk would be the two the come to mind | |
| 17:10:01 | dansmith | I don't think we write images to disk via privsep. console log.. maybe? I thought we can get that via libvirt | |
| 17:10:10 | clarkb | ya I'm not sure. Its just in aggregate privsep uses more memory than most other openstack services | |
| 17:10:30 | sean-k-mooney | dansmith: no we read the file that is written to disk im pretty sure | |
| 17:10:34 | clarkb | I think cinder? and neutron use more then privsep is next. Its been a while since I looked at hte memory profiling though | |
| 17:11:01 | bauzas | I don't know if we could somehow pdb the running privsep process thru a backdoor, because if we could, like we do with nova services, then we could monitor the growing memory | |
| 17:11:24 | sean-k-mooney | dansmith: i would hope for the images that we write it to somewhere we own then move it and change the permission if needed | |
| 17:11:33 | bauzas | I personnally use tracemalloc to persist the memory state and compare between snapshots | |
| 17:11:51 | sean-k-mooney | like in most case i woudl expect nova to put it in the image cache then use qemu-image to create a qcow with the iamge as the backing file | |
| 17:11:52 | bauzas | but this requires access to the process | |
| 17:12:16 | sean-k-mooney | so privsep shoudl only be needed for invokeign qemu-img and not the actul image downlaod | |
| 17:12:29 | sean-k-mooney | but not sure about the same codepath for raw images | |
| 17:13:11 | sean-k-mooney | bauzas: you coudl do that locally but i dont think privsep uses eventlet | |
| 17:13:35 | dansmith | sean-k-mooney: we can't possibly be streaming multiple gigs of image data over the privsep socket | |
| 17:13:35 | bauzas | sean-k-mooney: afaik it doesn't | |
| 17:13:57 | sean-k-mooney | dansmith: ya that woudl not make sense to me eitehr | |
| 17:14:16 | sean-k-mooney | i know the console log is stream over it as we had a security issue with that in the past | |
| 17:14:24 | sean-k-mooney | but that is small | |
| 17:15:24 | dansmith | well, that might be large, but not tens of gigs | |
| 17:15:53 | dansmith | and still not sure why we're doing that really, but AFAIK nova has to eat the whole thing and return it over RPC anyway, right? | |
| 17:15:57 | sean-k-mooney | in production yes in ci the vms dont run that long so i would be surpsied if we ever go close to a mb | |
| 17:16:35 | dansmith | oh sure | |
| 17:17:07 | sean-k-mooney | https://github.com/openstack/nova/blob/master/nova/config.py#L78-L80 | |
| 17:17:31 | sean-k-mooney | so in produciton we drop privesep abck to info | |
| 17:17:50 | clarkb | Looking at https://6546e150b851cf607d58-90bcb133e16a34d804e19428ec130682.ssl.cf5.rackcdn.com/865479/1/gate/tempest-full-py3/9076120/controller/logs/screen-memory_tracker.txt there are three different nova-cpu privsep processes and together use about 150MB of rss. Similar story for neutron ovn metadata. Buts its ~250MB rss | |
| 17:17:51 | sean-k-mooney | im not sure if we override that in ci | |
| 17:18:31 | sean-k-mooney | clarkb: two fo the prive sepc process i can accoutn for | |
| 17:19:04 | sean-k-mooney | nova has one privsep context for everything and os-vif has one for the linux bridge and ovs plugin but i only expect 1 of those two to be created in any given ci job | |
| 17:19:05 | clarkb | sean-k-mooney: why do we need multiple process per service though? | |
| 17:19:16 | sean-k-mooney | we need one per security context | |
| 17:19:43 | sean-k-mooney | and we shoudl have mulitpel security context per service | |
| 17:19:46 | clarkb | sean-k-mooney: but if your security contro lmehtod is file permissions (any maybe selinux rules) then you can only limit by pid user or pid selinux context | |
| 17:20:09 | clarkb | you aren't any more secure if you have three different processes that nova cpu can talk to | |
| 17:20:13 | sean-k-mooney | privesp drops permissions to only does in the context when it spawns | |
| 17:20:45 | sean-k-mooney | clarkb: if you have muplie context you can limit the permission of each funciton | |
| 17:21:21 | clarkb | sean-k-mooney: but if the same entity can tell each of those three to do the work then functioanlly its the same context | |
| 17:21:27 | sean-k-mooney | that is the permission model that privsep provides and the big delta form rootwarp | |
| 17:21:32 | clarkb | you're limited by what you can limit access to on the filesystem | |
| 17:21:46 | clarkb | and if I have permission to talk to one I have permisssion to talk to all since a process shares those attributes | |
| 17:22:16 | clarkb | and if I somehow break that security layer to promote myself to nova-cpu user or selinux context I can talk to all of them | |