| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-21 | |||
| 12:37:53 | admin1 | i see that the UUID appears in only 1 filed in the compute_nodes tables | |
| 12:46:00 | admin1 | sean-k-mooney, i think the hostname changed from fqdn -> non-fqdn | |
| 12:46:14 | admin1 | hostname remained the same | |
| 13:11:15 | admin1 | sean-k-mooney, how to check own uuid ? | |
| 13:11:19 | admin1 | from the hypervisor | |
| 13:19:43 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: compute: enhance compute evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858383 | |
| 13:19:44 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384 | |
| 13:21:55 | sahid | o/ gibi sean-k-mooney I have added you change that you were looking for, hope that makes sense | |
| 13:21:58 | sahid | https://review.opendev.org/c/openstack/nova/+/858384/20/nova/api/openstack/compute/evacuate.py#104 | |
| 13:22:09 | sahid | s/you/the | |
| 13:39:07 | sean-k-mooney | sahid: admin1 sorry was on a call downstream. sahid ill try and take a look at yyour change in general later in the week but that section looks like i was expecting so i think that will be fine | |
| 13:39:49 | sean-k-mooney | admin1: that is unforgunete the base way to resolve this issue would be to set teh hostname back to the fqdn | |
| 13:39:49 | admin1 | i have one vm in this i need to migrate .. after that i can just delete /re-initialize it | |
| 13:40:16 | sahid | sean-k-mooney: no worries, thanks a lot for your return | |
| 13:40:16 | sean-k-mooney | can you check the instance.host value for that vm | |
| 13:41:00 | sean-k-mooney | admin1: the instance.host value is ment to match the host value in the nova.conf | |
| 13:41:39 | sean-k-mooney | if you have just one vm the simpleist fix woudl be to set the nova.conf host value on that node to match the instance.host on the vm | |
| 13:41:51 | sean-k-mooney | then you should be able to cold migrate teh vm | |
| 13:42:51 | sean-k-mooney | live migrtate might also work depending on the vm. e.g. if you are using any special feature like sriov or cpu pinning then cold migration has a higher proablity of working | |
| 13:43:14 | admin1 | i was not able to find hostname or name value in nova.conf | |
| 13:43:35 | sean-k-mooney | admin1: if its not set the defautl is socket.gethostname() | |
| 13:44:24 | sean-k-mooney | admin1: https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.host | |
| 13:45:31 | sean-k-mooney | admin1: nova does not support changing the hostname because it currpts our db. we have had a bad expirince with customer doing this acidentally of late to the point that we are now working on detecting it and prevent the compute agent form starting when it happens https://review.opendev.org/q/topic:bp%252Fstable-compute-uuids | |
| 13:46:26 | admin1 | Failed to create resource provider record in placement API for UUID 88b9b395-784f-4d78-8497-3d674f7dff64 .. Conflicting resource provider name: h20 already exists .. this is what I have | |
| 13:46:38 | admin1 | so question is where does this UUID come from ? | |
| 13:48:04 | sean-k-mooney | ah yes that makes sesne | |
| 13:48:19 | sean-k-mooney | ok remove the host value | |
| 13:48:28 | sean-k-mooney | and upstea the instnace.host for that one instnace | |
| 13:49:05 | sean-k-mooney | they way the uuid is calulated today is we use the nova.conf host value to look for a compute service record with the same host value | |
| 13:49:37 | sean-k-mooney | *we look for a compute node record with the same host value not comptue service | |
| 13:50:32 | admin1 | so wherever in database, h20 with old UUID appears, i need to just updated it with the new 88b9b395-784f-4d78-8497-3d674f7dff64 uuid ? | |
| 13:55:23 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/wallaby: [stable-only] Use os-brick from source in wallaby https://review.opendev.org/c/openstack/nova/+/865134 | |
| 14:04:08 | sean-k-mooney | admin1: no | |
| 14:04:32 | sean-k-mooney | you should leave teh comptue node alone and update the host value on the one instnace that is affected | |
| 14:04:48 | sean-k-mooney | admin1: presumable its the full fqdn corrently right | |
| 14:04:58 | sean-k-mooney | and the host is not just the hostname not fqdn | |
| 14:05:05 | sean-k-mooney | on the compute node | |
| 14:05:17 | sean-k-mooney | so you need to make them match then migrate it | |
| 14:05:41 | sean-k-mooney | we use the instance.host to determin the rpc endpoint of the compute service that manages it | |
| 14:06:25 | sean-k-mooney | admin1: so if the compute service name change and you have just one vm the shortest way to fix it is update that one instnace and then migrate it | |
| 14:06:30 | admin1 | right .. its full fqdn, but the issue is the current hostname is also not able to register into placement .. it saysFailed to create resource provider record in placement API for UUID 88b9b395-784f-4d78-8497-3d674f7dff64 .. Conflicting resource provider name: h20 already exists | |
| 14:06:54 | admin1 | so just update for this one instance, node to h20 instead of h20.fqdn | |
| 14:07:32 | admin1 | host does show h20 .. node shows h20.fqdn | |
| 14:08:00 | sean-k-mooney | ok so i need you to check a few things all of which shoudl be the same | |
| 14:08:24 | sean-k-mooney | we need to check that the instance.host value and serivce.host value are the same. | |
| 14:08:38 | sean-k-mooney | the hyperviour hostname and placement RP name need to be the same | |
| 14:08:48 | sean-k-mooney | and the compute node uuid and placment uuid need to match | |
| 14:09:09 | sean-k-mooney | and the compute node host value must match the instace.host and service.host values | |
| 14:09:42 | sean-k-mooney | those are the 4 things that need to align. | |
| 14:10:25 | sean-k-mooney | nova does not support changign the hostname or the [DEFAULT]/host value after the agent is first started on a physical server | |
| 14:10:56 | sean-k-mooney | chaning either will currpt both the nova db and create issues in placemnt | |
| 14:11:44 | admin1 | " compute node uuid and placment uuid need to match" - where/how would I see those values ? | |
| 14:11:47 | admin1 | from the db ? | |
| 14:12:09 | sean-k-mooney | yep although you can actuly get them form the api too | |
| 14:12:17 | sean-k-mooney | the placment uuid is jsut in the placement show output | |
| 14:12:31 | sean-k-mooney | the compute node uuid is in the hypervior api if you use a new enough verion | |
| 14:12:47 | sean-k-mooney | admin1: but yes you can get it in the cell db compute_nodes table | |
| 14:14:29 | admin1 | in placement, i already have h20 as 59cc8a37-cee4-4dbc-84bf-18f56366bb2d, and h20.fqdn as 7bf78a2d-ce88-4ba2-a5b5-27fa4479f887 .. but in the h20 nova-compute logs, it tries to register itself as 5fecf61b-feb6-4af4-82d9-e5f5245e6ae9 | |
| 14:17:21 | admin1 | a grep of the whole database dump shows that that UUID ... 59c is only in 2 places ..... resource_providers and compute_nodes | |
| 14:17:52 | admin1 | so if I update those 2 tables with the new uuid 5f that the node is trying to register as instead of of the 59 in the db, would it fix ? | |
| 14:20:24 | sean-k-mooney | admin1: thats because you have presumable already deleted teh old compute service entry for the host | |
| 14:21:46 | sean-k-mooney | a safer approch might be to remove the resouce providers for that host in placment | |
| 14:22:03 | sean-k-mooney | allow the compute service to start up and regeister its self | |
| 14:22:17 | sean-k-mooney | then make sure the instance alines and them migrate it | |
| 14:23:25 | sean-k-mooney | admin1: i want to make it very clear however that the hostname changing is one of the most distructive things that can happen to the nova/placment dbs and is very non trivial to recover form | |
| 14:23:58 | sean-k-mooney | if you remove the placment rp with/without the fqdn | |
| 14:24:05 | sean-k-mooney | it will allow the compute service to start | |
| 14:24:28 | sean-k-mooney | if you ensure the instnace.host matches the running compute service you should then be able to manage it and migrate it | |
| 14:24:43 | admin1 | is there an api way to delete the entry from placement | |
| 14:24:46 | admin1 | instead from db | |
| 14:24:50 | admin1 | cli way | |
| 14:24:56 | sean-k-mooney | yes if it has no allocations | |
| 14:25:20 | sean-k-mooney | https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-delete | |
| 14:25:58 | sean-k-mooney | you might have to do https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-allocation-delete to delete the allcotion for the vm first | |
| 14:26:38 | sean-k-mooney | once the compute service is running and regesterd in placment again | |
| 14:27:04 | sean-k-mooney | update the instanace.host value to match the running service | |
| 14:27:25 | sean-k-mooney | optionally run https://docs.openstack.org/nova/latest/cli/nova-manage.html#placement-heal-allocations | |
| 14:27:33 | sean-k-mooney | for the singel instnace | |
| 14:27:37 | sean-k-mooney | and then migrate it | |
| 14:28:07 | sean-k-mooney | heal allocations will restorte the allocatiosn in palcment that you deleted to allow you to delete the resouce provider | |
| 14:28:27 | sean-k-mooney | cold migration will fix the alloction on the destination in either case when you confirm the migration | |
| 14:29:24 | admin1 | UUID of the consumer -- is the UUID of the vm ? | |
| 14:29:31 | sean-k-mooney | yep | |
| 14:29:41 | sean-k-mooney | in this case at least | |
| 14:30:02 | sean-k-mooney | if you are cold migrating a vm it will also have a second allocation using the migration uuid | |
| 14:32:02 | admin1 | resource provider allocation show UUID ( of h20 ) shows blank, but delete gives Resource provider has allocations | |
| 14:33:36 | admin1 | so there could be some more allocations in the old uuid .. but not present int he virsh list that i can see | |
| 14:34:07 | sean-k-mooney | were you able to delete the fqdn version | |
| 14:34:14 | admin1 | yes | |
| 14:34:18 | admin1 | fqdn one is gone | |
| 14:34:23 | sean-k-mooney | and the vm has the fqdn currently | |
| 14:34:51 | sean-k-mooney | if so can you check what virsh hostname outputs | |
| 14:35:01 | admin1 | in the instances.node , its set to fqdn | |
| 14:36:06 | sean-k-mooney | is it the hostname or hostname.fqdn | |
| 14:36:23 | sean-k-mooney | or i guess hostname.domain | |
| 14:36:57 | admin1 | i was able to delete both now | |
| 14:37:33 | sean-k-mooney | oh ok good | |
| 14:37:45 | sean-k-mooney | so the compute agent should now be able to start | |
| 14:37:50 | admin1 | the GUI hypervisors showed the instances .. | |
| 14:38:17 | admin1 | it registered itself now .. | |