| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-30 | |||
| 22:15:18 | johnsom | Yep, the instance is owned by the Octavia service account. | |
| 22:15:49 | sean-k-mooney | no what i mean is when openstack server delete 6fb27c67-539c-4525-8747-e26487d15e75 is run | |
| 22:15:49 | atmark | yes, i have admin access to all tenants | |
| 22:15:52 | sean-k-mooney | what user is that | |
| 22:16:04 | sean-k-mooney | ok so thats an admin user | |
| 22:16:41 | sean-k-mooney | i think you should be able to delete the vm then even if your current token is not for the correct project due to admin right if you are not using new policy | |
| 22:17:30 | sean-k-mooney | the server show is implying that its not in the current project | |
| 22:17:46 | sean-k-mooney | so i was wondering if the delete was failing for the same reason | |
| 22:18:04 | johnsom | There is a --all-projects for the server delete command too, just like for list. | |
| 22:18:13 | sean-k-mooney | i tought admins could bypass that check | |
| 22:18:25 | johnsom | I thought so too honestly | |
| 22:19:04 | sean-k-mooney | its been a while since i have done that so cant recall if you need to set the project id somehow or not | |
| 22:20:17 | atmark | openstack server --os-project-id 848fe125a93a408ba8a8044fb87e9cdf delete 6fb27c67-539c-4525-8747-e26487d15e75 | |
| 22:20:48 | atmark | 848.. is tenant ID of octavia | |
| 22:22:19 | sean-k-mooney | that would require your current user to have the member or admin role on that project for keystone to be able to issue the token | |
| 22:22:40 | sean-k-mooney | the uuid shoudl be enough if your an admin | |
| 22:22:51 | sean-k-mooney | --all-project is only required to delete by name | |
| 22:23:06 | sean-k-mooney | on a differnt project | |
| 22:23:15 | atmark | i'm able to list VMs `openstack server list --os-project-id 848fe125a93a408ba8a8044fb87e9cdf` | |
| 22:23:27 | sean-k-mooney | ack | |
| 22:23:33 | sean-k-mooney | and i assume the delete didnt work | |
| 22:23:38 | atmark | it didn't | |
| 22:23:51 | sean-k-mooney | you could try --force | |
| 22:23:56 | sean-k-mooney | but that is not for this usecase | |
| 22:24:10 | sean-k-mooney | its for forceing the delete now if you have soft delete enabled | |
| 22:24:26 | sean-k-mooney | you could also try reset-state | |
| 22:24:28 | sean-k-mooney | then delete | |
| 22:24:41 | sean-k-mooney | so reset the state to error | |
| 22:24:44 | sean-k-mooney | then delete form error | |
| 22:25:46 | atmark | throws an no ID found on reset-state | |
| 22:26:42 | sean-k-mooney | ya so i think what is happenign is there is a record in teh api db | |
| 22:26:58 | sean-k-mooney | but there is no cell mapping for the instnace and no recored in teh cell db | |
| 22:27:18 | sean-k-mooney | reset-state woudl try and update the instance in the cell db | |
| 22:27:32 | sean-k-mooney | but that fails because it was not created | |
| 22:28:01 | atmark | this is what's in the db https://paste.openstack.org/show/bbS8V1bZK0VcFO7WBD7p/ | |
| 22:28:15 | sean-k-mooney | if you have db access your could confirm that by checkign the api db to see if the build request exists and/or instance_mapping | |
| 22:28:45 | sean-k-mooney | ok thats in which db | |
| 22:28:48 | sean-k-mooney | cell0 | |
| 22:28:51 | sean-k-mooney | or api | |
| 22:29:00 | sean-k-mooney | nova has 3 databases by default | |
| 22:29:02 | atmark | full columns https://paste.openstack.org/show/bsnbvSiQ7ieLLCW7zPyF/ | |
| 22:29:07 | atmark | nova | |
| 22:29:28 | sean-k-mooney | that does not tell me anything what installer did you use | |
| 22:29:35 | atmark | kolla-ansible | |
| 22:29:49 | sean-k-mooney | ok i have that deploy locally let me see which db that is | |
| 22:29:55 | sean-k-mooney | which version? | |
| 22:30:00 | atmark | ussuri | |
| 22:30:05 | sean-k-mooney | ack | |
| 22:30:27 | atmark | there 3 nova db, nova, nova_api and nova_cell0 | |
| 22:30:36 | atmark | the paste came from nova db | |
| 22:30:39 | sean-k-mooney | ack ok so nova i the nova cell1 db | |
| 22:31:08 | sean-k-mooney | so an instance should only end up in cell1 after it has been asigned a host | |
| 22:31:51 | sean-k-mooney | the only time after its booted that it can not have a host in cell 1 is if its shleved | |
| 22:36:14 | atmark | so it didn't end up in any cell since it doesn't have a host? | |
| 22:36:27 | atmark | host is NULL | |
| 22:36:42 | sean-k-mooney | select * from nova_api.instance_mappings where instance_uuid = "ded221d0-8410-4f89-b322-b9ff0207a0dd"; | |
| 22:36:50 | sean-k-mooney | can you run that but with your instance uuid | |
| 22:37:19 | sean-k-mooney | you should see somethign like this https://paste.openstack.org/show/bAUqQT5Al2Q9GoF9Nln2/ | |
| 22:38:15 | atmark | https://paste.openstack.org/show/b9tx6I2rQ3tUtsGBR98g/ | |
| 22:38:20 | sean-k-mooney | what im wondering is does that exist and is the cell_id set | |
| 22:38:32 | sean-k-mooney | and if so is that cell in your cell_mappings table | |
| 22:38:54 | sean-k-mooney | ok so cell_id is 6 | |
| 22:39:11 | sean-k-mooney | the cell mappings | |
| 22:39:16 | sean-k-mooney | has potitally password inf | |
| 22:39:22 | sean-k-mooney | so dont paste that | |
| 22:39:46 | atmark | 6 exist | |
| 22:39:49 | sean-k-mooney | but does the db connection string for cell id 6 end in nova | |
| 22:39:49 | atmark | https://paste.openstack.org/show/bh9PvsUNrxNuYSRsATga/ | |
| 22:40:08 | atmark | 6 and 9 | |
| 22:40:13 | sean-k-mooney | yep | |
| 22:40:20 | sean-k-mooney | does 6 point to the db where the instance iss | |
| 22:40:25 | sean-k-mooney | i.e. nova in your case | |
| 22:40:48 | sean-k-mooney | it shoudl end in :3306/nova | |
| 22:41:10 | sean-k-mooney | this is how we know which database to look in to delete the instace at the cell level | |
| 22:41:39 | atmark | https://paste.openstack.org/show/blXqFYFedaFGS9U7igPG/ | |
| 22:41:50 | atmark | cell0 | |
| 22:42:06 | atmark | it's pointing to nova_cell0 | |
| 22:42:24 | sean-k-mooney | ya so cell0 is where it should be if it fail | |
| 22:42:43 | sean-k-mooney | can you ceck cell 0 and see if its also in the instance table there | |
| 22:43:48 | atmark | yup it's here | |
| 22:44:05 | sean-k-mooney | ok so its in both nova and nova_cell0 | |
| 22:44:10 | sean-k-mooney | thats really strange | |
| 22:44:14 | atmark | correct | |
| 22:44:40 | sean-k-mooney | so the quickest way to fix this is just to delete that one record | |
| 22:44:50 | sean-k-mooney | but it shoudl only ever end up in one of those two dbs | |
| 22:45:07 | sean-k-mooney | any instnace that failed before it was schduled to a host should end up in nova_cell0 | |
| 22:45:14 | atmark | delete in both db? | |
| 22:45:38 | atmark | probably safe to delete in both right | |
| 22:45:44 | sean-k-mooney | you could try just removing it in nova | |
| 22:45:57 | sean-k-mooney | and then delete again but ya it should be safe to delete in both | |
| 22:46:27 | sean-k-mooney | i dont know of a failrue mode like this unless you change you cell mapping at some point | |
| 22:46:44 | sean-k-mooney | normally cell 0 get id 1 an cell_1 gets id 2 | |
| 22:47:05 | sean-k-mooney | although there is no significance to the speicic id | |
| 22:47:13 | atmark | iirc, we did played with mappings because something failed | |
| 22:47:40 | sean-k-mooney | ack so maybe this change at some point | |
| 22:48:07 | sean-k-mooney | i have never seen this specific issue before | |
| 22:48:23 | atmark | yep. anywany, this is for monday. don't wanna cause something bad | |
| 22:48:40 | atmark | thanks for the help | |
| 22:49:18 | atmark | we have 3 prod envs, only 1 env is exhibiting this issue | |
| 22:49:25 | sean-k-mooney | ack | |