| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-27 | |||
| 11:22:15 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement encryption on backingStore https://review.opendev.org/c/openstack/nova/+/870012 | |
| 11:51:35 | opendevreview | Jorge San Emeterio proposed openstack/nova master: WIP: Dividing global privsep profile https://review.opendev.org/c/openstack/nova/+/871729 | |
| 11:51:35 | opendevreview | Jorge San Emeterio proposed openstack/nova master: WIP: Moving privsep profiles to nova/__init__.py https://review.opendev.org/c/openstack/nova/+/872010 | |
| 12:13:12 | opendevreview | Kashyap Chamarthy proposed openstack/nova stable/xena: libvirt: At start-up rework compareCPU() usage with a workaround https://review.opendev.org/c/openstack/nova/+/872011 | |
| 12:37:37 | elodilles | bauzas: speaking about the tons of things you are to review, i don't know whether you saw the 2023.2 Bobcat schedule plan: https://review.opendev.org/c/openstack/releases/+/869976 | |
| 12:37:41 | elodilles | o:) | |
| 12:38:21 | auniyal | Hello o/ | |
| 12:39:06 | auniyal | please review these Improving logging at '_allocate_mdevs'. | |
| 12:39:06 | auniyal | - https://review.opendev.org/q/topic:bug%252F1992451 | |
| 13:18:13 | bauzas | elodilles: I've seen it and I forgot to ask on Tuesday for folks to look at this change | |
| 13:18:21 | bauzas | elodilles: during the nova meeting | |
| 13:18:52 | bauzas | elodilles: I'm not bad about it fwiw, as we would have the Summit *before* Specfreeze which is nice | |
| 13:44:38 | opendevreview | Jorge San Emeterio proposed openstack/nova master: DNM: Testing check pipeline on master branch https://review.opendev.org/c/openstack/nova/+/872018 | |
| 13:52:24 | elodilles | bauzas: yepp, it's a fairly long cycle (28 weeks) so milestones have longer times | |
| 14:12:07 | darkhorse | Hi team, Is it possible to disallow a user from all compute resources using keystone policy or any other means? The use case is I want to have multiple roles to a project. Some users have compute permissions, some have network permissions, some have volume permissions etc. | |
| 14:12:36 | darkhorse | I am not sure if this question is more relevant to keystone channel. | |
| 15:54:47 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/victoria: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833436 | |
| 16:01:03 | opendevreview | Merged openstack/nova stable/yoga: Add a workaround to skip hypervisor version check on LM https://review.opendev.org/c/openstack/nova/+/851202 | |
| 16:02:42 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437 | |
| 16:02:58 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438 | |
| 16:03:45 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/ussuri: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833437 | |
| 16:04:03 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438 | |
| 16:04:21 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/train: reenable greendns in nova. https://review.opendev.org/c/openstack/nova/+/833438 | |
| 16:14:33 | sean-k-mooney | dansmith: can you take a look at https://review.opendev.org/c/openstack/nova/+/863919/12/nova/tests/unit/compute/test_resource_tracker.py#1553 when you have time | |
| 16:15:26 | dansmith | sean-k-mooney: yeah, it's on my list.. ya'll drug me into a call at 6:30am and I have another starting in a few ;P | |
| 16:15:48 | sean-k-mooney | sorry about that | |
| 16:15:49 | dansmith | thanks for hitting the ones below, that's some progress for sure | |
| 16:16:25 | sean-k-mooney | so ill try and test teh rest when i get devstack deploy so ill loop back to them again later today if that works out | |
| 16:16:39 | dansmith | cool | |
| 17:10:51 | sean-k-mooney | cool devstack stacked first time | |
| 17:11:24 | sean-k-mooney | now ot actully use your patches | |
| 17:12:26 | sean-k-mooney | so currently i have all in one so when i checkout your branch and restart the nova serices im expecting the compute to start up and writeh the uuid to a file and update the service version right | |
| 17:16:39 | sean-k-mooney | yep that worked https://paste.opendev.org/show/bLMvRfuUugp9iMcn4R60/ | |
| 17:17:47 | sean-k-mooney | im going to break this in a few differnt ways and document it in an ether pad and ill let you know if i find anything odd or broken | |
| 17:25:44 | dansmith | sean-k-mooney: yeah, cool | |
| 17:29:59 | sean-k-mooney | https://etherpad.opendev.org/p/Stable-compute-uuid-manual-testing ill be usign that to keep my notes | |
| 17:30:37 | sean-k-mooney | there is not really much there yet but ill kep adding to it as i go | |
| 17:59:27 | sean-k-mooney | dansmith: found a bug | |
| 17:59:36 | sean-k-mooney | getting logs now | |
| 18:00:02 | dansmith | in "test 3 restart with deleted file" ? | |
| 18:00:08 | sean-k-mooney | yep | |
| 18:00:38 | dansmith | does that mean deleted node id file? | |
| 18:00:53 | sean-k-mooney | i moved it to compute_id_old | |
| 18:01:11 | sean-k-mooney | the agent tried to add a duplicate row in the db which resulted in a key error | |
| 18:01:15 | sean-k-mooney | and it wrote a new file | |
| 18:01:22 | sean-k-mooney | with a diffent uuid | |
| 18:01:50 | dansmith | well, right, | |
| 18:02:08 | dansmith | because it knows it's not an upgrade (now) and it didn't find a local compute node uuid | |
| 18:02:14 | dansmith | what do you think it should do there? | |
| 18:02:45 | sean-k-mooney | well it shoudl not result in a traceback in the log for one | |
| 18:03:02 | sean-k-mooney | but it should see that the compute node exist in the db and not try to create a new one | |
| 18:03:12 | sean-k-mooney | and we should not write a new file to disk | |
| 18:03:26 | dansmith | but the whole point of this is that the file becomes the source of truth | |
| 18:04:00 | dansmith | so maybe we should see the failure to create with a keyerror as a "something's wrong abort" but I bet it happens too late for us to abort | |
| 18:04:28 | sean-k-mooney | well the file didnt exist correct | |
| 18:04:48 | sean-k-mooney | so before we try to create a new compute node record shoudl we not check if one exits the old way | |
| 18:04:55 | sean-k-mooney | and abort then | |
| 18:05:00 | dansmith | we do, specifically on upgrade only | |
| 18:05:05 | dansmith | which you saw work | |
| 18:05:24 | dansmith | but once you have gotten past upgrade, it knows you're not upgrading, and can only really assume that it's a greenfield deployment | |
| 18:05:36 | dansmith | if we're going to stop relying on the hostname, there's really no other way, right? | |
| 18:06:17 | sean-k-mooney | no if we have a compute node and the serivce is upgraed then the file shoudl exist | |
| 18:06:25 | sean-k-mooney | if it does not then we knwo somethign odd happened | |
| 18:07:11 | sean-k-mooney | basicly if compute service>=62? and not file -> error | |
| 18:07:31 | sean-k-mooney | for greanfiled thre wont be a compute node in the db | |
| 18:07:31 | dansmith | how do we know the difference between "already upgraded" and "greenfield" ? | |
| 18:07:53 | dansmith | but again, the only way you know the "comptue node in the db" is because you're relying on the hostname, which we should not be doing | |
| 18:08:01 | sean-k-mooney | or we can catch the duplicate key error form the db | |
| 18:08:16 | dansmith | depending on when the duplicate key happens, I'm fine aborting in that case for sure | |
| 18:08:35 | sean-k-mooney | we at least shoudl not write the file with the wrong uuid | |
| 18:08:43 | dansmith | but I think we don't create that record until far too late | |
| 18:08:43 | sean-k-mooney | its currently writing the one for the row that was rejected | |
| 18:08:53 | sean-k-mooney | proably because we set the singlton uuid | |
| 18:09:24 | dansmith | okay, I'm with you on not writing it if there's a keyerror, but I think that will be pretty hard to line those things up | |
| 18:09:35 | dansmith | so maybe delete it if we wrote it or something | |
| 18:10:01 | dansmith | although, hang on | |
| 18:10:22 | dansmith | let's say I deploy two new computes with the same hostname by accident, they both generate and write unique uuids, | |
| 18:10:30 | dansmith | one will fail to create their compute node because of the clash, | |
| 18:10:48 | dansmith | but if I just rename the offending duplicate one, then the uuid it generated is fine to use when I restart | |
| 18:10:55 | sean-k-mooney | yep | |
| 18:11:14 | dansmith | the "I deleted the id" case seems pretty edge-y to me, because there are tons of things I can randomly delete from a running system that will break stuff | |
| 18:11:25 | dansmith | if we can abort on key conflict at start, then I'm on board, but otherwise, I dunno | |
| 18:11:48 | dansmith | this is a bit like saying I deleted /etc/ssh/ssh_host_key and when I restarted it generated a new one and my clients all complain :) | |
| 18:12:05 | sean-k-mooney | honestly i did this as my first test because i tough it could be a trivial thing that might happen and we should prevent it | |
| 18:12:24 | sean-k-mooney | dansmith: i was more thinging what happens if you fail to bind mount this in a container | |
| 18:13:26 | sean-k-mooney | if the file is ever lost i think it defeats much of the utility of the feature if we allow the agent to start | |
| 18:13:43 | dansmith | okay, but if you do, there's not much harm, because the resolution is to restart with it, and no harm no foul right? | |
| 18:14:00 | dansmith | if you failed to bind mount this, you likely also failed to bind-mount the instance images no? | |
| 18:14:17 | sean-k-mooney | fair its in the state dir | |
| 18:14:18 | dansmith | in the pre-provisioned case, this goes in /etc/nova anyway | |
| 18:14:22 | sean-k-mooney | although it depend on the location | |
| 18:17:00 | sean-k-mooney | any way im goign to keep testing/breaking it for a while and see what else i find | |
| 18:17:09 | sean-k-mooney | i just added the traceback | |
| 18:17:29 | dansmith | ack, I say we punt on that for the moment and see what else you can find | |
| 18:17:47 | dansmith | maybe we can circle back with some extra stuff that makes that smarter | |
| 18:18:02 | dansmith | at least catching keyerror and logging something that says "okay, here's how you've screwed up..." | |
| 18:18:18 | dansmith | because we have that trace right now and it's not super obvious why | |
| 18:26:47 | sean-k-mooney | so i think the abort is broken in general | |
| 18:28:30 | dansmith | the abort on rename? | |
| 18:28:47 | sean-k-mooney | yep so i tried restarting it with the incorrect uuid | |