Earlier  
Posted Nick Remark
#openstack-nova - 2021-03-03
16:29:53 belmoreira we have an independent rabbit for everything
16:29:58 bauzas some kind of instance that would be pre-created
16:30:46 belmoreira bauzas each cell has it's own rabbit, including one for the support conductor/scheduler. Then we have a rabbit for the notifications
16:31:15 belmoreira support/super
16:31:29 bauzas oh sorry, I missed the fact you were mentioning the notifications rabbit
16:31:42 bauzas and not the API MQ or the cell MQs
16:32:02 bauzas I guess we hold on emitting notifications
16:33:13 belmoreira I was expecting that behaviour (maybe with an error msg), but in reality instances can't be created
16:33:40 belmoreira was looking into the config options that I think this is not handled at all
16:33:51 gibi belmoreira: I think we need to warp the notification sending with some exception handler and log a WARNING if the notification sending failed but not block the actual work
16:35:06 belmoreira gibi +1
16:35:29 gibi belmoreira: I think we simply did not have those error handled properly in the current code
16:36:06 gibi one can argue that not delivering a notification could mean some external system become desynced
16:36:42 belmoreira gibi I agree, but in that case should be configurable
16:37:27 gibi belmoreira: hm, that could work. Something like notification_failure_is_fatal config option
16:38:14 gibi so if somebody use the notification interface for charging customers based on usage then that deployer would likely make this config True
16:38:57 belmoreira makes sense to me
16:41:47 gibi me to
16:41:48 gibi o
16:44:12 belmoreira I'm not familiar with the notifications code. Is this something that someone can have a look?
16:44:44 gibi I can take a look but my backlog is pretty long so it will take time to reach that bug
16:46:44 belmoreira thanks gibi, meanwhile I can have a look but definitely I will need some guidance
16:47:10 gibi belmoreira: sure, let me know if you have questions
16:48:16 belmoreira thanks a lot
16:49:11 bauzas gibi: belmoreira: sorry wrapped into a meeting, but the approach looks good to me
16:57:39 openstackgerrit Balazs Gibizer proposed openstack/nova master: Remove non-libguestfs file injection for libvirt https://review.opendev.org/c/openstack/nova/+/324720
16:58:51 openstackgerrit Balazs Gibizer proposed openstack/nova master: Remove VFSLocalFS https://review.opendev.org/c/openstack/nova/+/778506
16:59:47 kashyap gibi: Interesting that you revived it (I agree). Anything in particular that made you revive?
17:00:35 gibi kashyap: the security bug behind it become public a week ago
17:00:49 gibi kashyap: and also we removed Xen support so one less complication
17:00:50 kashyap Argh, rotting security bugs :-(
17:00:53 gibi yepp
17:01:23 kashyap gibi: Right; fair enough. That problem is real ...
17:03:53 bauzas gibi: urgent review needs on them, I guess ?
17:04:06 gibi bauzas: no, it is a really old security bug
17:04:18 gibi so no need to rush
17:04:51 gibi it just become stuck in private state until the secu team did a spring cleaning recently
17:04:55 bauzas gibi: okay, focusing on blueprints reviews, but I can take a look at them later
17:04:56 gibi and made the bug public
17:05:02 gibi bauzas: sure, thanks
17:10:09 sean-k-mooney thats the one we talked about 2 weeks ago in teh meeting right
17:10:30 sean-k-mooney i assume you revied the old patches
17:11:25 gibi sean-k-mooney: yes, I revived, rebased, and realized the we deleted Xen since so I put a cleanup top of it
17:11:32 gibi sean-k-mooney: but the basic idea is the same
17:11:49 gibi sean-k-mooney: fail to boot if file injection is requested but libguestfs is not available on the compute
17:11:51 sean-k-mooney cool ill try and review this this week
17:11:56 gibi sean-k-mooney: thanks
17:12:48 sean-k-mooney by the way i saw your question on the port numa patches. ill hopefully get time to rebase that tomorrow to address it i just need to fix that env i was using it for something else but ill do that in the morning
17:12:59 sean-k-mooney thanks for taking a look
17:13:03 gibi ack
17:19:49 bauzas sean-k-mooney: just a quick q, why do we need to pass the list of ARQs when shelving an instance ? I guess this is for the cyborg-agent to free up the resources ?
17:19:55 bauzas context : https://review.opendev.org/c/openstack/nova/+/778440/1/nova/compute/manager.py
17:25:25 sean-k-mooney bauzas: we need to free them yes
17:25:31 sean-k-mooney so its for unbinding them
17:25:46 sean-k-mooney technially its only needed for the shelve_offload part
17:26:30 bauzas yup, that's what I guessed
17:36:11 sean-k-mooney i commented in line but i dont think this is a ddos vector really
17:36:46 sean-k-mooney the new api query only happens if the instance has cyborg resoucs. and it will happen only once per shleved instance
17:38:30 sean-k-mooney bauzas: we also prefilter the list by the timeout and only do this for instance that have exceed the time out so we wont check this on every iteration of the perodic
17:39:14 sean-k-mooney pulling the client out of the loop is not a bad idea
17:39:43 bauzas sure, but I wonder whether some malicious user could create 10000 small instances by one and wait for 3600secs
17:40:07 sean-k-mooney i mean they would hit there instance quota right
17:40:16 sean-k-mooney shelved instances still count to that
17:40:27 bauzas the problem is that we call N times the cyborg api
17:40:31 bauzas at the same time
17:40:45 bauzas and you multiply by the periodic value
17:41:07 sean-k-mooney ya but you cant avoid that wihout caching the info in nova which we do not do intentionally
17:41:20 bauzas maybe not an attack vector but some performance impact for sure
17:41:31 sean-k-mooney i dont think it will be
17:42:10 bauzas on a large cloud with 10000 instances being shelved at the same time from the same tenant, cyborg will face 10000 times a connection roundtrip
17:42:15 sean-k-mooney if the nova api is beefy enought to handel the 10000 shleve api calls then the cyborg one should be able to handel 10000 arq lookups
17:42:19 bauzas from different tenants*
17:42:37 bauzas that's a periodic
17:42:41 bauzas not an API straight call
17:42:52 sean-k-mooney sure i know
17:42:54 bauzas during those 3600 secs, you can create and shelve as much instances as you want
17:43:11 sean-k-mooney right but we defualt to 0
17:43:16 bauzas but once the periodic runs, it will pick all the shelved instances during this window
17:43:19 sean-k-mooney e.g. offloading without a delay
17:44:24 sean-k-mooney so for it to be an issue the operator has to opt in to offloading after a period of time and increase it enouch for the shelved instance to build up enough to ddos the cyborg api
17:44:55 sean-k-mooney pragmatically i dont think we will enough user of cyborg+shelve +that non default config for this to realisticlly happen
17:45:29 sean-k-mooney it could but the instance.save() before this would propably ddos the db before the cyborg issue was hit
17:46:23 sean-k-mooney im not saying it not a valid concern i just dont think it makes it substantailly worse then it would be already
17:46:29 bauzas sean-k-mooney: I'm just saying "doc it"
17:46:43 sean-k-mooney well we should doc the instance.save then too right
17:46:54 bauzas because shelving has a very specific implication now if you use cyborg
17:47:23 sean-k-mooney well shelve offload already calls out to cinder
17:47:24 bauzas well, I'm pretty sure ops size the DB write performance on the nova database
17:47:33 sean-k-mooney to remove the volumn attaments right?
17:47:54 sean-k-mooney and it should be calling neutron to unbind the neutron ports
17:48:06 sean-k-mooney so im wonderign wny cyborg is special in this case
17:48:26 bauzas actually, you're making a point
17:48:57 bauzas but that doesn't infirm my theory
17:49:16 bauzas which is that this periodic task can harm our dependent projects
17:49:31 sean-k-mooney well im not disagreeing that it could cause issue if you misconfigred things
17:49:37 sean-k-mooney yep
17:49:58 sean-k-mooney i think we just have differnet levels of concners as to the effect
17:50:10 bauzas anyway, I don't wanna hold this
17:50:34 bauzas but I was just taking my operator hat, being afraid of my internal APIs
17:50:51 sean-k-mooney moving the client out of the loop will at least stop the load on keystone

Earlier   Later