| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-03-03 | |||
| 17:19:49 | bauzas | sean-k-mooney: just a quick q, why do we need to pass the list of ARQs when shelving an instance ? I guess this is for the cyborg-agent to free up the resources ? | |
| 17:19:55 | bauzas | context : https://review.opendev.org/c/openstack/nova/+/778440/1/nova/compute/manager.py | |
| 17:25:25 | sean-k-mooney | bauzas: we need to free them yes | |
| 17:25:31 | sean-k-mooney | so its for unbinding them | |
| 17:25:46 | sean-k-mooney | technially its only needed for the shelve_offload part | |
| 17:26:30 | bauzas | yup, that's what I guessed | |
| 17:36:11 | sean-k-mooney | i commented in line but i dont think this is a ddos vector really | |
| 17:36:46 | sean-k-mooney | the new api query only happens if the instance has cyborg resoucs. and it will happen only once per shleved instance | |
| 17:38:30 | sean-k-mooney | bauzas: we also prefilter the list by the timeout and only do this for instance that have exceed the time out so we wont check this on every iteration of the perodic | |
| 17:39:14 | sean-k-mooney | pulling the client out of the loop is not a bad idea | |
| 17:39:43 | bauzas | sure, but I wonder whether some malicious user could create 10000 small instances by one and wait for 3600secs | |
| 17:40:07 | sean-k-mooney | i mean they would hit there instance quota right | |
| 17:40:16 | sean-k-mooney | shelved instances still count to that | |
| 17:40:27 | bauzas | the problem is that we call N times the cyborg api | |
| 17:40:31 | bauzas | at the same time | |
| 17:40:45 | bauzas | and you multiply by the periodic value | |
| 17:41:07 | sean-k-mooney | ya but you cant avoid that wihout caching the info in nova which we do not do intentionally | |
| 17:41:20 | bauzas | maybe not an attack vector but some performance impact for sure | |
| 17:41:31 | sean-k-mooney | i dont think it will be | |
| 17:42:10 | bauzas | on a large cloud with 10000 instances being shelved at the same time from the same tenant, cyborg will face 10000 times a connection roundtrip | |
| 17:42:15 | sean-k-mooney | if the nova api is beefy enought to handel the 10000 shleve api calls then the cyborg one should be able to handel 10000 arq lookups | |
| 17:42:19 | bauzas | from different tenants* | |
| 17:42:37 | bauzas | that's a periodic | |
| 17:42:41 | bauzas | not an API straight call | |
| 17:42:52 | sean-k-mooney | sure i know | |
| 17:42:54 | bauzas | during those 3600 secs, you can create and shelve as much instances as you want | |
| 17:43:11 | sean-k-mooney | right but we defualt to 0 | |
| 17:43:16 | bauzas | but once the periodic runs, it will pick all the shelved instances during this window | |
| 17:43:19 | sean-k-mooney | e.g. offloading without a delay | |
| 17:44:24 | sean-k-mooney | so for it to be an issue the operator has to opt in to offloading after a period of time and increase it enouch for the shelved instance to build up enough to ddos the cyborg api | |
| 17:44:55 | sean-k-mooney | pragmatically i dont think we will enough user of cyborg+shelve +that non default config for this to realisticlly happen | |
| 17:45:29 | sean-k-mooney | it could but the instance.save() before this would propably ddos the db before the cyborg issue was hit | |
| 17:46:23 | sean-k-mooney | im not saying it not a valid concern i just dont think it makes it substantailly worse then it would be already | |
| 17:46:29 | bauzas | sean-k-mooney: I'm just saying "doc it" | |
| 17:46:43 | sean-k-mooney | well we should doc the instance.save then too right | |
| 17:46:54 | bauzas | because shelving has a very specific implication now if you use cyborg | |
| 17:47:23 | sean-k-mooney | well shelve offload already calls out to cinder | |
| 17:47:24 | bauzas | well, I'm pretty sure ops size the DB write performance on the nova database | |
| 17:47:33 | sean-k-mooney | to remove the volumn attaments right? | |
| 17:47:54 | sean-k-mooney | and it should be calling neutron to unbind the neutron ports | |
| 17:48:06 | sean-k-mooney | so im wonderign wny cyborg is special in this case | |
| 17:48:26 | bauzas | actually, you're making a point | |
| 17:48:57 | bauzas | but that doesn't infirm my theory | |
| 17:49:16 | bauzas | which is that this periodic task can harm our dependent projects | |
| 17:49:31 | sean-k-mooney | well im not disagreeing that it could cause issue if you misconfigred things | |
| 17:49:37 | sean-k-mooney | yep | |
| 17:49:58 | sean-k-mooney | i think we just have differnet levels of concners as to the effect | |
| 17:50:10 | bauzas | anyway, I don't wanna hold this | |
| 17:50:34 | bauzas | but I was just taking my operator hat, being afraid of my internal APIs | |
| 17:50:51 | sean-k-mooney | moving the client out of the loop will at least stop the load on keystone | |
| 17:50:58 | sean-k-mooney | so that proably shoudl be done in anycase | |
| 17:51:05 | bauzas | and while I trust cinder and neutron in terms of sizing, I'd certainly take much attention to cyborg API if I know this | |
| 17:51:32 | sean-k-mooney | ya there is less op experince with cyborg for sure | |
| 17:51:51 | bauzas | anyway, this is nearly 7pm for me and I really worked late yesterday night, so I'll bail out | |
| 17:52:02 | sean-k-mooney | that said at least with osp donwstream we tend to use one db server for all the dbs | |
| 17:52:15 | sean-k-mooney | enjoy your evnening o/ | |
| 17:52:47 | bauzas | like, hello keystone | |
| 17:53:17 | bauzas | bear with me | |
| 17:53:32 | bauzas | and give me a token that you'll store for a long period of time | |
| 17:53:56 | bauzas | glance ? be quiet and just give me the images | |
| 17:54:21 | bauzas | but meh, I'm diverging | |
| 17:54:22 | bauzas | \o | |
| 17:55:13 | sean-k-mooney | im sure cern use different dbs and rabbitmq instance for different services but many smaller clouds wont | |
| 20:56:22 | openstackgerrit | Merged openstack/nova master: requirements.txt: Bump os-brick to 4.2.0 https://review.opendev.org/c/openstack/nova/+/778177 | |
| 21:01:22 | openstackgerrit | Merged openstack/nova master: Remove unused Instance method https://review.opendev.org/c/openstack/nova/+/757615 | |
| 21:38:14 | openstackgerrit | Merged openstack/nova master: hyper-v rbd volume support https://review.opendev.org/c/openstack/nova/+/763550 | |
| #openstack-nova - 2021-03-04 | |||
| 02:51:09 | openstackgerrit | YuehuiLei proposed openstack/nova-specs master: Add xena directory for specs https://review.opendev.org/c/openstack/nova-specs/+/778604 | |
| 03:23:19 | openstackgerrit | Brin Zhang proposed openstack/nova master: Add missed accel_uuids for _poll_shelved_instances https://review.opendev.org/c/openstack/nova/+/778440 | |
| 07:41:14 | openstackgerrit | Yongli He proposed openstack/nova master: Smartnic support - cyborg drive https://review.opendev.org/c/openstack/nova/+/771362 | |
| 07:41:15 | openstackgerrit | Yongli He proposed openstack/nova master: smartnic support https://review.opendev.org/c/openstack/nova/+/758944 | |
| 08:31:17 | yonglihe | gibi: hope you have some bandwidth.. | |
| 09:20:17 | bauzas | lyarwood: morning | |
| 09:20:48 | openstack | Launchpad bug 1820007 in os-brick "Failed to attach encrypted volumes after detach: volume device not found at /dev/disk/by-id" [Undecided,Fix released] - Assigned to Lee Yarwood (lyarwood) | |
| 09:21:27 | bauzas | I'll recheck my change, but given the bug was fixed, I wonder whether it was related | |
| 09:33:14 | lyarwood | morning | |
| 09:34:58 | gibi | morning folks | |
| 09:35:13 | gibi | stephenfin: do we need https://review.opendev.org/c/openstack/nova/+/765798 for the hypervisor api bp? | |
| 09:36:02 | stephenfin | I think it's a nice-to-have rather than a necessity. gmann did say yesterday that he was going to take over that patch though, | |
| 09:37:00 | openstackgerrit | Lee Yarwood proposed openstack/nova master: WIP: nova-next: Start testing the 'q35' machine type https://review.opendev.org/c/openstack/nova/+/708701 | |
| 09:37:20 | lyarwood | weird, just had `error: remote unpack failed: error Missing blob 08f698bbe36d414eca19d0760a5acc3a714c404e` errors trying to push that ^ | |
| 09:37:34 | lyarwood | appears the git gc actually did something for a change and it's fixed after a git fetch -va | |
| 09:37:48 | lyarwood | bauzas: looking at that issue now sorry | |
| 09:38:11 | bauzas | lyarwood: no worries at all, rechecked meanwhile | |
| 09:38:20 | bauzas | just an open question | |
| 09:38:37 | bauzas | and I thought you could be interested in since you worked on the fix | |
| 09:41:17 | lyarwood | bauzas: yeah this is slightly different | |
| 09:41:19 | lyarwood | 43684 Mar 03 20:53:14.002343 ubuntu-focal-limestone-regionone-0023285105 nova-compute[53948]: ERROR oslo_messaging.rpc.server libvirt.libvirtError: internal error: unable to execute QEMU command 'blockdev-add': Could not open '/dev/disk/by-id/scsi-360000000000000000e00000 000010001': Operation not permitted | |
| 09:41:30 | lyarwood | we find the device but can't attach it | |
| 09:42:39 | lyarwood | weird, we even encrypt it | |
| 09:42:58 | lyarwood | ah but through /dev/sda | |
| 09:43:03 | lyarwood | I hope that's the same device | |
| 09:43:28 | gibi | stephenfin: ack, then I will not push hard on that policy patch | |
| 09:46:27 | bauzas | lyarwood: that looks to me a race finding the device, right? | |
| 09:50:14 | lyarwood | bauzas: no we've found the device fine, /dev/disk/by-id/scsi-360000000000000000e00000000010001 is connected correctly and we even format/encrypt it with LUKSv1 | |
| 09:50:29 | lyarwood | bauzas: QEMU just isn't happy with the passphrase we've provided AFAICT | |
| 09:50:34 | bauzas | ah | |
| 09:50:46 | lyarwood | bauzas: the only odd thing is that we format/encrypt /dev/sda | |
| 09:51:37 | lyarwood | https://github.com/openstack/os-brick/blame/bd629a3a4105f7f3f9f35b71350ea3c66f3690e9/os_brick/encryptors/cryptsetup.py#L80-L81 and https://github.com/openstack/os-brick/blob/bd629a3a4105f7f3f9f35b71350ea3c66f3690e9/os_brick/encryptors/luks.py#L97 cause that | |
| 09:51:47 | lyarwood | but I've never seen that be a problem before, it should be the same underlying block device | |
| 09:52:35 | lyarwood | unless theres some weirdness in the block layers and /dev/disk/by-id/scsi-360000000000000000e00000000010001 still doesn't look like it's encrypted by the time QEMU attempts to attach it | |