| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-09-09 | |||
| 08:52:24 | gibi | bauzas: if you feel the power, then could you please review https://review.opendev.org/c/openstack/placement/+/807014 I think (and CI thinks) it is good and fixing the transaction issue | |
| 08:52:31 | bauzas | hopefully next year, we'll change the superb to a new electric vehicle, because we tested long-range trips with intermediary recharges, and this works | |
| 08:52:47 | bauzas | gibi: excellent point, I need to look at this one | |
| 08:52:56 | bauzas | gibi: I also need to amend the vgpu doc | |
| 08:53:56 | gibi | and if you are at placement land then https://review.opendev.org/c/openstack/placement/+/807155 is simple and fixes a bug in consumer_types | |
| 08:53:59 | bauzas | gibi: I guess melwitt addressed your excellent concern ? | |
| 08:54:06 | songwenping_ | hi,team, is any entry to delete compute node except nova service-delete? | |
| 08:54:28 | gibi | bauzas: yes, she added a independent transaction for re-reading the rp data in the retry loop, and it seems to work well | |
| 08:54:34 | bauzas | songwenping_: you shouldn't delete the compute node entries | |
| 08:54:55 | bauzas | songwenping_: either the virt driver or the service deletion can do this | |
| 08:55:19 | songwenping_ | how virt driver works? | |
| 08:57:09 | kashyap | songwenping_: The bird's-eye view is this: | |
| 08:57:42 | kashyap | nova-api (in coordination with nova-scheduler) --> nova-compute (virt driver) --> launches libvirtd --> launches QEMU | |
| 08:57:55 | kashyap | But you have to be more specific for people to answer :) | |
| 08:59:01 | songwenping_ | no, i means which scenes virt driver delete the compute node? | |
| 09:16:15 | bauzas | songwenping_: sorry, I need to jump off for 30 mins | |
| 09:16:38 | bauzas | but basically, the virt driver gives the inventories and the compute nodes to the RT which creates the necessary records | |
| 09:16:43 | bauzas | RT : ResourceTracker | |
| 09:17:13 | bauzas | as the RT is run by the nova-compute service, you need to delete the service | |
| 09:17:50 | bauzas | there is a tight relationship between an RPC service (the nova-compute manager) and the compute node record | |
| 11:15:56 | songwenping_ | hi, team, anybody knows why placement check source node resource when evacuate VM? | |
| 11:19:19 | sean-k-mooney | how do you mean | |
| 11:19:29 | sean-k-mooney | as it it can include the source host in the set of host returned | |
| 11:21:03 | sean-k-mooney | that would be because we do not currently filter host but just up hosts | |
| 11:21:08 | sean-k-mooney | https://github.com/openstack/nova/blob/master/nova/scheduler/request_filter.py#L241-L254 | |
| 11:21:16 | sean-k-mooney | we do filter hosts by there disabeld status however | |
| 11:21:39 | sean-k-mooney | so fi you had dissbled the host you are evacuatating form then it would not be included in the placment query | |
| 11:22:30 | sean-k-mooney | songwenping_: the source node will be elimiated by the scudler after the placment query so it does not really have a negitive impact to include it | |
| 11:34:05 | opendevreview | Merged openstack/nova master: Support Cpu Compararion on Aarch64 Platform https://review.opendev.org/c/openstack/nova/+/763928 | |
| 11:41:48 | songwenping_ | sean-k-mooney: wait for a min, i am finding the placement code. | |
| 11:44:14 | songwenping_ | when we evacuate vm, placement will check_capacity_exceeded, https://github.com/openstack/placement/blob/master/placement/objects/allocation.py#L73 | |
| 11:45:47 | songwenping_ | https://github.com/openstack/placement/blob/master/placement/objects/allocation.py#L120 contains source node provider id and dest node provider id. | |
| 12:05:33 | sean-k-mooney | yes | |
| 12:06:11 | sean-k-mooney | when we make the allocation candiate quest we do not exclude the host we are evacuating form | |
| 12:06:40 | sean-k-mooney | ah i see | |
| 12:07:45 | sean-k-mooney | this should not have any negitive effect | |
| 12:08:53 | sean-k-mooney | we are technicall checkign allocation for one addtional host that we dont need placment to condiser | |
| 12:10:00 | songwenping_ | sometimes there are some rubbish data at allocation table, this lead to evacuate failed due to this check. | |
| 12:12:05 | sean-k-mooney | that shoudl just eliminate the host as an allcoation candiate no? | |
| 12:12:20 | sean-k-mooney | have you filed a bug for this | |
| 12:12:39 | sean-k-mooney | placment should not be made aware of evuacation or other lifecycle operations explcitly | |
| 12:13:20 | songwenping_ | havenot filed bug. | |
| 12:13:25 | sean-k-mooney | so im not sure we could proceed with any apprcoh that required modifciation of placment to make it expcitly aware of evacuation | |
| 12:13:48 | sean-k-mooney | but we might be able to handel the error condition with corupt data | |
| 12:13:57 | sean-k-mooney | so that it woudl not fail | |
| 12:17:56 | gibi | songwenping_: during evacuation the source node allocation of the VM is kept and the dest node allocation is added to it | |
| 12:18:18 | gibi | so for an evacuated VM you will see allocation on the source and dest nodes _until_ the source compute is down | |
| 12:18:40 | gibi | when the source compute node is recovered it will delete the allocation on the source node for the already evacuated VMs | |
| 12:19:27 | sean-k-mooney | gibi: for evacuate should we not be using the migration uuid for the evaucation | |
| 12:19:32 | sean-k-mooney | alocaltions | |
| 12:19:36 | sean-k-mooney | like we do for rezise | |
| 12:20:24 | songwenping_ | gibi: yeah, this is right workflow, but placement check the source node capacity before evacuate. | |
| 12:21:12 | gibi | sean-k-mooney: we never switched the evac workflow to migration allocations | |
| 12:21:15 | gibi | there is todo in the code | |
| 12:21:40 | sean-k-mooney | ah ok | |
| 12:22:02 | sean-k-mooney | i guess that would be the correct way to fix this then | |
| 12:23:08 | gibi | songwenping_: do you have a pointer where nova checks the source node capacity during evac? | |
| 12:23:35 | sean-k-mooney | its in placment https://github.com/openstack/placement/blob/master/placement/objects/allocation.py#L120 | |
| 12:24:43 | songwenping_ | yes, it's placement checking. | |
| 12:25:41 | gibi | what placement does during any kind of allocation update, is to replace the existing allocation of the consumer (VM in this case) with the new requested allocation | |
| 12:26:01 | gibi | so if the node is not overallocated then the replace_all should not fail | |
| 12:26:18 | gibi | if you got your node overallocated then I guess the allocation update can fail | |
| 12:27:07 | gibi | as placement will not allow you to overallocate | |
| 12:27:10 | songwenping_ | yes, i means why we check the source node overallocated as the vm will be evacuated to the dest node. | |
| 12:27:38 | gibi | placement does not know these things, what placement sees is a request to update the allocation of a consumer | |
| 12:27:54 | gibi | and the old allocation has resources on the source node, the new allocation has resource on the source node and on the dest node | |
| 12:28:06 | gibi | placement goes and replaces the whole old allocation with the new allocation | |
| 12:28:33 | gibi | if you got your compute overallocated then simply removing the source allocation and adding it back at this step fails | |
| 12:29:07 | gibi | sean-k-mooney: btw, moving to migration allocation will not solve this as there we move the VM allocation to the migration allocation and that move will fail for the same reason | |
| 12:29:27 | gibi | in short, if you overallocated your compute then placement will reject allocation updates on that compute | |
| 12:29:41 | gibi | you need to resolve the overallocation | |
| 12:29:42 | sean-k-mooney | it should reject new allcoations | |
| 12:29:50 | gibi | no | |
| 12:29:59 | sean-k-mooney | but if we dont change the reosuce in a currnt oen it shoudl work no | |
| 12:30:12 | gibi | placement implements replace_all for allocation update | |
| 12:30:23 | gibi | so it is a delete + create in the same transaction | |
| 12:30:32 | gibi | but after the delte the compute is full | |
| 12:30:36 | gibi | so create fails | |
| 12:30:43 | songwenping_ | on our product env, there are some unknow reasons that vm has allocations on two resource provider. | |
| 12:33:36 | gibi | sean-k-mooney: I mean I got that it would be nice to detect that the source node allocation does not change during the update and dont delete + re-create it, but still placement does not do that logic | |
| 12:34:08 | sean-k-mooney | well i was hoping that a simple uuid update woudl not triger this check | |
| 12:34:28 | gibi | sean-k-mooney: there is no way to update a consumer uuid | |
| 12:35:09 | gibi | sean-k-mooney: you update the allocation of a consumer or you create / delete consumers | |
| 12:35:13 | gibi | there is no rename consumer | |
| 12:35:25 | gibi | and there is no partial allocation update | |
| 12:35:28 | gibi | just total one | |
| 12:35:58 | sean-k-mooney | ack | |
| 12:36:27 | gibi | probably the easyest thing is to implement rename consumer, the partial allocation update feels hard | |
| 12:37:10 | sean-k-mooney | ya so we would move the source allocation to the migration uuid create a new allocation for the vm using its uuid | |
| 12:37:18 | gibi | then we can reimplement the allocation move from VM -> migration with the rename | |
| 12:37:24 | sean-k-mooney | then have the dest delete the migration allocation after evac | |
| 12:38:02 | sean-k-mooney | which will avoid leaking the allcoation in placment if the source compute never comes back | |
| 12:38:13 | sean-k-mooney | well actully no | |
| 12:38:30 | sean-k-mooney | we have to be careful | |
| 12:38:42 | sean-k-mooney | to make sure if the evac fails we can jsut evac again | |
| 12:38:53 | gibi | yeah, I feel there is a reason why we kept the source allocation for the source compute to clean up | |
| 12:38:53 | sean-k-mooney | so the dest vm need to have the migration uuid | |
| 12:39:14 | sean-k-mooney | untill it succeed then we cna remvoe the source vm allcoation and rename the migration allcotion | |
| 12:39:36 | gibi | having the migration uuid to allocate on the test is a surgery as today we just call the scheduler and that always uses the instance uuid to allocate | |
| 12:39:43 | gibi | s/test/dest/ | |