| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-21 | |||
| 14:08:00 | sean-k-mooney | ok so i need you to check a few things all of which shoudl be the same | |
| 14:08:24 | sean-k-mooney | we need to check that the instance.host value and serivce.host value are the same. | |
| 14:08:38 | sean-k-mooney | the hyperviour hostname and placement RP name need to be the same | |
| 14:08:48 | sean-k-mooney | and the compute node uuid and placment uuid need to match | |
| 14:09:09 | sean-k-mooney | and the compute node host value must match the instace.host and service.host values | |
| 14:09:42 | sean-k-mooney | those are the 4 things that need to align. | |
| 14:10:25 | sean-k-mooney | nova does not support changign the hostname or the [DEFAULT]/host value after the agent is first started on a physical server | |
| 14:10:56 | sean-k-mooney | chaning either will currpt both the nova db and create issues in placemnt | |
| 14:11:44 | admin1 | " compute node uuid and placment uuid need to match" - where/how would I see those values ? | |
| 14:11:47 | admin1 | from the db ? | |
| 14:12:09 | sean-k-mooney | yep although you can actuly get them form the api too | |
| 14:12:17 | sean-k-mooney | the placment uuid is jsut in the placement show output | |
| 14:12:31 | sean-k-mooney | the compute node uuid is in the hypervior api if you use a new enough verion | |
| 14:12:47 | sean-k-mooney | admin1: but yes you can get it in the cell db compute_nodes table | |
| 14:14:29 | admin1 | in placement, i already have h20 as 59cc8a37-cee4-4dbc-84bf-18f56366bb2d, and h20.fqdn as 7bf78a2d-ce88-4ba2-a5b5-27fa4479f887 .. but in the h20 nova-compute logs, it tries to register itself as 5fecf61b-feb6-4af4-82d9-e5f5245e6ae9 | |
| 14:17:21 | admin1 | a grep of the whole database dump shows that that UUID ... 59c is only in 2 places ..... resource_providers and compute_nodes | |
| 14:17:52 | admin1 | so if I update those 2 tables with the new uuid 5f that the node is trying to register as instead of of the 59 in the db, would it fix ? | |
| 14:20:24 | sean-k-mooney | admin1: thats because you have presumable already deleted teh old compute service entry for the host | |
| 14:21:46 | sean-k-mooney | a safer approch might be to remove the resouce providers for that host in placment | |
| 14:22:03 | sean-k-mooney | allow the compute service to start up and regeister its self | |
| 14:22:17 | sean-k-mooney | then make sure the instance alines and them migrate it | |
| 14:23:25 | sean-k-mooney | admin1: i want to make it very clear however that the hostname changing is one of the most distructive things that can happen to the nova/placment dbs and is very non trivial to recover form | |
| 14:23:58 | sean-k-mooney | if you remove the placment rp with/without the fqdn | |
| 14:24:05 | sean-k-mooney | it will allow the compute service to start | |
| 14:24:28 | sean-k-mooney | if you ensure the instnace.host matches the running compute service you should then be able to manage it and migrate it | |
| 14:24:43 | admin1 | is there an api way to delete the entry from placement | |
| 14:24:46 | admin1 | instead from db | |
| 14:24:50 | admin1 | cli way | |
| 14:24:56 | sean-k-mooney | yes if it has no allocations | |
| 14:25:20 | sean-k-mooney | https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-delete | |
| 14:25:58 | sean-k-mooney | you might have to do https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-allocation-delete to delete the allcotion for the vm first | |
| 14:26:38 | sean-k-mooney | once the compute service is running and regesterd in placment again | |
| 14:27:04 | sean-k-mooney | update the instanace.host value to match the running service | |
| 14:27:25 | sean-k-mooney | optionally run https://docs.openstack.org/nova/latest/cli/nova-manage.html#placement-heal-allocations | |
| 14:27:33 | sean-k-mooney | for the singel instnace | |
| 14:27:37 | sean-k-mooney | and then migrate it | |
| 14:28:07 | sean-k-mooney | heal allocations will restorte the allocatiosn in palcment that you deleted to allow you to delete the resouce provider | |
| 14:28:27 | sean-k-mooney | cold migration will fix the alloction on the destination in either case when you confirm the migration | |
| 14:29:24 | admin1 | UUID of the consumer -- is the UUID of the vm ? | |
| 14:29:31 | sean-k-mooney | yep | |
| 14:29:41 | sean-k-mooney | in this case at least | |
| 14:30:02 | sean-k-mooney | if you are cold migrating a vm it will also have a second allocation using the migration uuid | |
| 14:32:02 | admin1 | resource provider allocation show UUID ( of h20 ) shows blank, but delete gives Resource provider has allocations | |
| 14:33:36 | admin1 | so there could be some more allocations in the old uuid .. but not present int he virsh list that i can see | |
| 14:34:07 | sean-k-mooney | were you able to delete the fqdn version | |
| 14:34:14 | admin1 | yes | |
| 14:34:18 | admin1 | fqdn one is gone | |
| 14:34:23 | sean-k-mooney | and the vm has the fqdn currently | |
| 14:34:51 | sean-k-mooney | if so can you check what virsh hostname outputs | |
| 14:35:01 | admin1 | in the instances.node , its set to fqdn | |
| 14:36:06 | sean-k-mooney | is it the hostname or hostname.fqdn | |
| 14:36:23 | sean-k-mooney | or i guess hostname.domain | |
| 14:36:57 | admin1 | i was able to delete both now | |
| 14:37:33 | sean-k-mooney | oh ok good | |
| 14:37:45 | sean-k-mooney | so the compute agent should now be able to start | |
| 14:37:50 | admin1 | the GUI hypervisors showed the instances .. | |
| 14:38:17 | admin1 | it registered itself now .. | |
| 14:38:25 | sean-k-mooney | ya if you have db currption like this its hard to resolve | |
| 14:38:58 | sean-k-mooney | so if its regeisterd its self you just need to ensure the instance.host and service.host agree and you shoudl be able to migrate | |
| 14:39:12 | admin1 | now when i try to migrate, using openstack server migrate, it says compute host h20 could not be found | |
| 14:39:20 | admin1 | so i guess its trying to refer to some other h20 | |
| 14:40:06 | sean-k-mooney | you see the compute service in openstack compute service list right and its up | |
| 14:40:30 | sean-k-mooney | oh did you update the compute service mappings in the api deb | |
| 14:40:45 | admin1 | not the last part .. | |
| 14:40:45 | sean-k-mooney | you need to run nova-mange cell_v2 discover_hosts i think | |
| 14:41:06 | admin1 | that would be from inside the nova venv ? | |
| 14:41:06 | sean-k-mooney | the new comptue service recorred need to be mapped to the correct cell | |
| 14:41:19 | sean-k-mooney | ideally form one of the contoller with db access | |
| 14:41:23 | admin1 | ok | |
| 14:41:39 | admin1 | from the os utils as admin, or as nova in the nova venv | |
| 14:41:58 | sean-k-mooney | nova-mange uses the credential in the nova.conf | |
| 14:42:04 | admin1 | got it | |
| 14:42:11 | sean-k-mooney | so you should use the config the conductor uses | |
| 14:42:19 | sean-k-mooney | or ap | |
| 14:42:23 | sean-k-mooney | *or api | |
| 14:42:51 | sean-k-mooney | https://docs.openstack.org/nova/latest/cli/nova-manage.html#cell-v2-discover-hosts | |
| 14:50:09 | admin1 | sean-k-mooney , thank you .. finally its migrating to another host | |
| 14:50:33 | sean-k-mooney | admin1: once you have completed that and confirmed the migration | |
| 14:50:41 | admin1 | i read the spec .. having a uuid that is associated with the server and not associated with hostname will fix issues like this in future | |
| 14:50:44 | sean-k-mooney | i woudl advise checking if any other computes have had a host name change | |
| 14:51:15 | sean-k-mooney | admin1: that not really what the sepc is goign to do | |
| 14:51:36 | admin1 | in my case, grafana showed fqdn while others were non-fqdn, so another collegue decided to change the fqdn to just hostname only and not full fqdn to make the graphs sane | |
| 14:51:37 | sean-k-mooney | admin1: the spec will record the compute node uuid in a file so we can detech if the hostname changes | |
| 14:52:10 | sean-k-mooney | admin1: ya if they did that everywhere they woudl have severly currpted the db | |
| 14:52:32 | sean-k-mooney | you woudl need to cold migrate every workload on teh affected hosts to resolve it | |
| 14:52:58 | sean-k-mooney | its much much better to correct the hostname if no new instnace have been created | |
| 14:54:21 | admin1 | yeah .. we normally ensure all is what is needed before deploying the first vm | |
| 14:54:43 | admin1 | but in this case, since the server was replaced and pxe put the fqdn, it slipped | |
| 14:54:55 | sean-k-mooney | ack | |
| 14:54:58 | admin1 | and it was corrected only after the first vm was deployed | |
| 14:55:19 | sean-k-mooney | i would advise autiting the rest to make sure there are no other in this state | |
| 14:55:40 | sean-k-mooney | the longer you hosts like this the harder it is to fix | |
| 15:13:23 | opendevreview | Merged openstack/os-vif stable/zed: Move mtu update request into ovsdb transaction https://review.opendev.org/c/openstack/os-vif/+/863993 | |
| 16:23:16 | opendevreview | sean mooney proposed openstack/nova-specs master: add spec for fqdn in hostname https://review.opendev.org/c/openstack/nova-specs/+/862626 | |
| 16:24:19 | sean-k-mooney | gibi: dansmith melwitt ^ that hopefully has adressed all the outstanding comments | |
| 16:26:10 | admin1 | as an operator, i would prefer to not have fqdn but just hostnames .. because when it comes to monitoring and graphs, with fqdn it clutters the whole graph .. h20.location.dev.cloud.domain.com where the location.dev.cloud.domain.com is redundant in all | |
| 16:26:58 | sean-k-mooney | admin1: that spec is for vms | |
| 16:27:12 | sean-k-mooney | and on the compute nodes you are free to use either | |
| 16:27:21 | sean-k-mooney | i prefer using hostnames on the computes nodes too | |
| 16:30:32 | sean-k-mooney | admin1: for what its worth nova did not test and recommended against using fqdns for the compute node hostname for a very long time | |