Earlier  
Posted Nick Remark
#openstack-nova - 2021-03-03
15:33:43 sean-k-mooney live_migration_inbound_addr is only for live migration
15:34:17 sean-k-mooney you can adres this via kernel routes
16:02:39 openstackgerrit Stephen Finucane proposed openstack/nova master: apidb: Compact Rocky database migrations https://review.opendev.org/c/openstack/nova/+/759405
16:03:48 openstackgerrit Stephen Finucane proposed openstack/nova master: apidb: Compact Stein database migrations https://review.opendev.org/c/openstack/nova/+/759406
16:05:08 openstackgerrit Stephen Finucane proposed openstack/nova master: apidb: Compact Train database migrations https://review.opendev.org/c/openstack/nova/+/771420
16:18:52 openstackgerrit Sylvain Bauza proposed openstack/nova master: Bump the Compute RPC API to version 6.0 https://review.opendev.org/c/openstack/nova/+/761452
16:19:20 bauzas dansmith: gibi: stephenfin: artom: updated based on your feedbacks ^
16:19:31 gibi bauzas: looking
16:20:13 bauzas tl;dr: removed the unused args from the RPC methods stephenfin told + changed the numa tests to pin to 5.max instead
16:23:26 openstackgerrit Merged openstack/python-novaclient master: Add support for microversion v2.88 https://review.opendev.org/c/openstack/python-novaclient/+/770573
16:25:34 belmoreira recently found this behaviour in Nova: https://bugs.launchpad.net/nova/+bug/1917645 and I'm not sleeping well since :) not sure if this should be Nova or Oslo. I would appreciate some guidance.
16:25:35 openstack Launchpad bug 1917645 in OpenStack Compute (nova) "Nova can't create instances if RabbitMQ notification cluster is down" [Undecided,New]
16:29:19 bauzas belmoreira: if the rabbit is down, how the conductor and scheduler could help the nova-api service to tell which host ?
16:29:43 bauzas that actually reminds me preemptible instances
16:29:53 belmoreira we have an independent rabbit for everything
16:29:58 bauzas some kind of instance that would be pre-created
16:30:46 belmoreira bauzas each cell has it's own rabbit, including one for the support conductor/scheduler. Then we have a rabbit for the notifications
16:31:15 belmoreira support/super
16:31:29 bauzas oh sorry, I missed the fact you were mentioning the notifications rabbit
16:31:42 bauzas and not the API MQ or the cell MQs
16:32:02 bauzas I guess we hold on emitting notifications
16:33:13 belmoreira I was expecting that behaviour (maybe with an error msg), but in reality instances can't be created
16:33:40 belmoreira was looking into the config options that I think this is not handled at all
16:33:51 gibi belmoreira: I think we need to warp the notification sending with some exception handler and log a WARNING if the notification sending failed but not block the actual work
16:35:06 belmoreira gibi +1
16:35:29 gibi belmoreira: I think we simply did not have those error handled properly in the current code
16:36:06 gibi one can argue that not delivering a notification could mean some external system become desynced
16:36:42 belmoreira gibi I agree, but in that case should be configurable
16:37:27 gibi belmoreira: hm, that could work. Something like notification_failure_is_fatal config option
16:38:14 gibi so if somebody use the notification interface for charging customers based on usage then that deployer would likely make this config True
16:38:57 belmoreira makes sense to me
16:41:47 gibi me to
16:41:48 gibi o
16:44:12 belmoreira I'm not familiar with the notifications code. Is this something that someone can have a look?
16:44:44 gibi I can take a look but my backlog is pretty long so it will take time to reach that bug
16:46:44 belmoreira thanks gibi, meanwhile I can have a look but definitely I will need some guidance
16:47:10 gibi belmoreira: sure, let me know if you have questions
16:48:16 belmoreira thanks a lot
16:49:11 bauzas gibi: belmoreira: sorry wrapped into a meeting, but the approach looks good to me
16:57:39 openstackgerrit Balazs Gibizer proposed openstack/nova master: Remove non-libguestfs file injection for libvirt https://review.opendev.org/c/openstack/nova/+/324720
16:58:51 openstackgerrit Balazs Gibizer proposed openstack/nova master: Remove VFSLocalFS https://review.opendev.org/c/openstack/nova/+/778506
16:59:47 kashyap gibi: Interesting that you revived it (I agree). Anything in particular that made you revive?
17:00:35 gibi kashyap: the security bug behind it become public a week ago
17:00:49 gibi kashyap: and also we removed Xen support so one less complication
17:00:50 kashyap Argh, rotting security bugs :-(
17:00:53 gibi yepp
17:01:23 kashyap gibi: Right; fair enough. That problem is real ...
17:03:53 bauzas gibi: urgent review needs on them, I guess ?
17:04:06 gibi bauzas: no, it is a really old security bug
17:04:18 gibi so no need to rush
17:04:51 gibi it just become stuck in private state until the secu team did a spring cleaning recently
17:04:55 bauzas gibi: okay, focusing on blueprints reviews, but I can take a look at them later
17:04:56 gibi and made the bug public
17:05:02 gibi bauzas: sure, thanks
17:10:09 sean-k-mooney thats the one we talked about 2 weeks ago in teh meeting right
17:10:30 sean-k-mooney i assume you revied the old patches
17:11:25 gibi sean-k-mooney: yes, I revived, rebased, and realized the we deleted Xen since so I put a cleanup top of it
17:11:32 gibi sean-k-mooney: but the basic idea is the same
17:11:49 gibi sean-k-mooney: fail to boot if file injection is requested but libguestfs is not available on the compute
17:11:51 sean-k-mooney cool ill try and review this this week
17:11:56 gibi sean-k-mooney: thanks
17:12:48 sean-k-mooney by the way i saw your question on the port numa patches. ill hopefully get time to rebase that tomorrow to address it i just need to fix that env i was using it for something else but ill do that in the morning
17:12:59 sean-k-mooney thanks for taking a look
17:13:03 gibi ack
17:19:49 bauzas sean-k-mooney: just a quick q, why do we need to pass the list of ARQs when shelving an instance ? I guess this is for the cyborg-agent to free up the resources ?
17:19:55 bauzas context : https://review.opendev.org/c/openstack/nova/+/778440/1/nova/compute/manager.py
17:25:25 sean-k-mooney bauzas: we need to free them yes
17:25:31 sean-k-mooney so its for unbinding them
17:25:46 sean-k-mooney technially its only needed for the shelve_offload part
17:26:30 bauzas yup, that's what I guessed
17:36:11 sean-k-mooney i commented in line but i dont think this is a ddos vector really
17:36:46 sean-k-mooney the new api query only happens if the instance has cyborg resoucs. and it will happen only once per shleved instance
17:38:30 sean-k-mooney bauzas: we also prefilter the list by the timeout and only do this for instance that have exceed the time out so we wont check this on every iteration of the perodic
17:39:14 sean-k-mooney pulling the client out of the loop is not a bad idea
17:39:43 bauzas sure, but I wonder whether some malicious user could create 10000 small instances by one and wait for 3600secs
17:40:07 sean-k-mooney i mean they would hit there instance quota right
17:40:16 sean-k-mooney shelved instances still count to that
17:40:27 bauzas the problem is that we call N times the cyborg api
17:40:31 bauzas at the same time
17:40:45 bauzas and you multiply by the periodic value
17:41:07 sean-k-mooney ya but you cant avoid that wihout caching the info in nova which we do not do intentionally
17:41:20 bauzas maybe not an attack vector but some performance impact for sure
17:41:31 sean-k-mooney i dont think it will be
17:42:10 bauzas on a large cloud with 10000 instances being shelved at the same time from the same tenant, cyborg will face 10000 times a connection roundtrip
17:42:15 sean-k-mooney if the nova api is beefy enought to handel the 10000 shleve api calls then the cyborg one should be able to handel 10000 arq lookups
17:42:19 bauzas from different tenants*
17:42:37 bauzas that's a periodic
17:42:41 bauzas not an API straight call
17:42:52 sean-k-mooney sure i know
17:42:54 bauzas during those 3600 secs, you can create and shelve as much instances as you want
17:43:11 sean-k-mooney right but we defualt to 0
17:43:16 bauzas but once the periodic runs, it will pick all the shelved instances during this window
17:43:19 sean-k-mooney e.g. offloading without a delay
17:44:24 sean-k-mooney so for it to be an issue the operator has to opt in to offloading after a period of time and increase it enouch for the shelved instance to build up enough to ddos the cyborg api
17:44:55 sean-k-mooney pragmatically i dont think we will enough user of cyborg+shelve +that non default config for this to realisticlly happen
17:45:29 sean-k-mooney it could but the instance.save() before this would propably ddos the db before the cyborg issue was hit
17:46:23 sean-k-mooney im not saying it not a valid concern i just dont think it makes it substantailly worse then it would be already
17:46:29 bauzas sean-k-mooney: I'm just saying "doc it"
17:46:43 sean-k-mooney well we should doc the instance.save then too right
17:46:54 bauzas because shelving has a very specific implication now if you use cyborg

Earlier   Later