| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-03-22 | |||
| 17:03:30 | EugenMayer | Is there any 'good way' to set the task-state of an instance that has been stuck in 'image backup' due to an issue in glance? so the field OS-EXT-STS:task_state is on "image_backup" | |
| 17:05:07 | EugenMayer | i see there is 'nova set --state' or 'nova reset-state' but both seeem to operate on the instance-power-state (OS-EXT-STS:power_state) or OS-EXT-STS:vm_state - but not the task-state | |
| 17:07:20 | zigo | sean-k-mooney: Yeah, this was an evacuate operation. | |
| 17:07:47 | dansmith | zigo: I thought you said live migrate? | |
| 17:07:50 | sean-k-mooney | zigo: ok the reason this breaks is for evacuation we only have 1 allocation in placemnt against both hosts | |
| 17:08:17 | sean-k-mooney | and since the souce host is over capstiy because you reduce the allocate ration the entire allcoation is considered invlaid | |
| 17:08:34 | sean-k-mooney | we disussed this at the ptg 1 or 2 ptgs ago | |
| 17:09:13 | sean-k-mooney | i cant recall if we said we should fix this after consumer types but i dont think we had a workaround other then tempoarly increase the allcoation ratio so its nolonger over commited | |
| 17:09:55 | dansmith | sean-k-mooney: we could also solve it the way we do for cold migration, which is hold the allocation on the source with the migration uuid right? | |
| 17:10:13 | sean-k-mooney | dansmith: yes we could that was on eof the options | |
| 17:10:50 | sean-k-mooney | im trying to find the launchpad bug | |
| 17:11:44 | bauzas | dansmith: sean-k-mooney: yeah, the Migration uuid for evacuate seems the better and cleaner approach | |
| 17:12:04 | sean-k-mooney | bauzas: that is what we were proposing doing | |
| 17:12:26 | sean-k-mooney | but i dont think anyone has worked on it since | |
| 17:13:39 | bauzas | :-) | |
| 17:13:47 | sean-k-mooney | https://bugs.launchpad.net/nova/+bug/1943191 | |
| 17:13:49 | sean-k-mooney | that might be it | |
| 17:13:58 | EugenMayer | I'am looking on https://wiki.openstack.org/wiki/CrashUp/Recover_From_Nova_Uncontrolled_Operations to understand how to recover from the crashed task state 'image_backup' but i'am not sure how to actual act upon that. Should i use the nova api? | |
| 17:14:03 | sean-k-mooney | and https://bugs.launchpad.net/nova/+bug/1924123 | |
| 17:14:37 | bauzas | sean-k-mooney: some people expect bugs to be fixed automatically :) | |
| 17:14:55 | bauzas | we don't have yet AI bots smart enough to close the gaps | |
| 17:14:58 | sean-k-mooney | EugenMayer: the wiki is basicaly unmaintained | |
| 17:15:09 | EugenMayer | i see. Thank you | |
| 17:16:03 | sean-k-mooney | in the early days of openstack we used the wiki for sepc and project created docs(docs not by the docs team) | |
| 17:16:14 | EugenMayer | I'am really not sure hot to again recover from the failed task the proper way. The only way i yet know, which is huge is: reset the state, then restart the compute the vm is hosted so thee state is somewhat recovered | |
| 17:17:13 | sean-k-mooney | there is not way to recover form it really beyond that | |
| 17:17:29 | sean-k-mooney | we dont provide a api to allow taskt to be restarted | |
| 17:17:53 | dansmith | reset state and reboot the vm is what I'd try first, | |
| 17:18:00 | sean-k-mooney | yep same | |
| 17:18:03 | dansmith | not restarting the compute I'd hope | |
| 17:18:16 | sean-k-mooney | ya that normally shoudl not be required | |
| 17:18:25 | sean-k-mooney | i guess it woudl depend on why it failed | |
| 17:18:43 | dansmith | definitely not expected for anything like a glance thing | |
| 17:18:50 | EugenMayer | trying that. AFAIR i had to restart the entire compute last time. Anyway, trying that | |
| 17:19:17 | sean-k-mooney | do you recall way? | |
| 17:19:20 | sean-k-mooney | *why | |
| 17:19:23 | EugenMayer | dansmith well this happens the 4th time. A stuck glance image backup task leaves the task_state of the instance in a broken state | |
| 17:19:26 | dansmith | honestly restarting the compute shouldn't even do anything, AFAIK | |
| 17:19:59 | sean-k-mooney | i wonder if the main thread of the compute agent was blocked on an io operations | |
| 17:20:16 | sean-k-mooney | that is the only thing i can think of that would be fixed by an agent restart | |
| 17:20:41 | sean-k-mooney | we were not using a thread pool on some of the older release for those | |
| 17:20:53 | dansmith | sean-k-mooney: compute is the thing that "consumes" the task_state and turns it into a vm_state, so to speak, so maybe we clear task_state in init_host in some cases? | |
| 17:20:58 | EugenMayer | well i'am on xena, so not really old | |
| 17:21:13 | dansmith | but either way, reset_state to error is supposed to let you clear everything by enabling force reboot I think | |
| 17:21:17 | dansmith | or that's the intent | |
| 17:21:18 | sean-k-mooney | dansmith: i think we do yes but not sure about this case | |
| 17:22:01 | EugenMayer | dansmith it is clear, swt wise, that there is more then one misconception in the microservice and task callstack. I'am not sure if glance is required to call a webhook on success or error (not sure how the result is propagated) but this is simply not the right design. | |
| 17:22:30 | EugenMayer | should the task crash on glance, neither success nor error is called (ever) and there seems nothing to recover from that | |
| 17:22:42 | dansmith | EugenMayer: none of that :) | |
| 17:22:58 | dansmith | everything is nova->glance | |
| 17:23:05 | sean-k-mooney | i belive this is a blocking call to do the upload to glace | |
| 17:23:15 | sean-k-mooney | if its async then either nova would poll | |
| 17:23:22 | sean-k-mooney | or we woudl get an external event form glance | |
| 17:23:26 | dansmith | so depending on the failure, nova should clean up whatever it can.. an upload to glance for sure should be recoverable on our end, so that's likely it's own bug if we're missing something | |
| 17:23:28 | sean-k-mooney | but i think image upload if blocking | |
| 17:23:35 | dansmith | sean-k-mooney: none of that with glance | |
| 17:24:00 | sean-k-mooney | right we dont do polling or external event right | |
| 17:24:07 | sean-k-mooney | we just do two blocking calls | |
| 17:24:09 | EugenMayer | if it is a blocking task, well the blocking should cleanup - which it seem to not do | |
| 17:24:15 | sean-k-mooney | one for creating the image and the second for the data upload | |
| 17:24:34 | dansmith | EugenMayer: if you can repro the problem that's definitely a bug candidate | |
| 17:24:42 | sean-k-mooney | EugenMayer: yes it should clean up if we get an error form glance | |
| 17:24:56 | dansmith | there are some situations where it might not make sense to clean up, but I would think a glance thing would always be something we can handle | |
| 17:25:04 | EugenMayer | dansmith i can reproduce this the 4th time. If you tell me what to gather, i will grab the logs you need the 5th time - which will happen | |
| 17:25:13 | dansmith | EugenMayer: logs | |
| 17:25:18 | EugenMayer | which logs to get? | |
| 17:25:26 | dansmith | all of them? :) | |
| 17:25:26 | sean-k-mooney | dansmith: i would expect the vm to go back to active or error if we dont clean up right | |
| 17:25:32 | dansmith | nova-compute, nova-api at least | |
| 17:25:36 | dansmith | sean-k-mooney: error, yeah | |
| 17:25:56 | EugenMayer | vm is in active state, power is on, task_state is image_backup | |
| 17:26:07 | dansmith | that said, reset_state resets task_state so that should be the way to get out here | |
| 17:26:23 | sean-k-mooney | you can reset state to active | |
| 17:26:39 | EugenMayer | reset-state --active + reboot seems to recover just right. Also viewing the console works (which is one of the problems with a partial state recovery) | |
| 17:26:40 | dansmith | EugenMayer: we're saying that what we would expect is vm_state=ERROR,task_state=None | |
| 17:26:44 | sean-k-mooney | rather then error and potentaly just trigger the backup/snapthot again | |
| 17:26:55 | EugenMayer | dansmith that never happened yet | |
| 17:27:07 | dansmith | EugenMayer: I know, I'm saying that's what we expect nova should be doing | |
| 17:27:17 | sean-k-mooney | EugenMayer: do you know why the glance operation is failing. | |
| 17:28:02 | sean-k-mooney | dansmith: i could see an argument to be made that we woudl have vm_state=Active task_state=None but the snapshot action was marked as error in the server event log | |
| 17:28:20 | sean-k-mooney | if the vm was indeed still runing proberly depending on how it failed | |
| 17:28:42 | EugenMayer | there is so much one can break right now. e.g. a other topic is using terraform and rescale a flavor. In 2 of 5 cases the following happens (i cannot tell you exactly). The old flavor is delete (too early), the new one is created, then the instance is fetched, this fails since the flavor_id of the old flavor is still set and cannot be found. TF | |
| 17:28:42 | EugenMayer | cancles and that's it | |
| 17:28:43 | dansmith | sean-k-mooney: the problem is one of signaling, which is why we (originally as designed) went to error,None for everything and then you do a start (which does nothing) to reset back to active as sort of "ack" | |
| 17:29:07 | EugenMayer | stuck again - stuck that one now needs to shelve the instance and restore it from glance using the 'new flavor' | |
| 17:30:01 | EugenMayer | i did not yet check the tf openstack provider implementation to see what they have implemented and how that is a timing issue in the first place (since it does not happen every timee) .. but if i look at the openstack rest api / nova api .. swapping flavors is not designed at all. | |
| 17:30:19 | sean-k-mooney | EugenMayer: well flavor are intenede to be imuatble so you idealy woudl not delete them until all instance using them are resized | |
| 17:30:28 | sean-k-mooney | we do cache the flavor | |
| 17:30:30 | sean-k-mooney | in the insntace | |
| 17:30:42 | dansmith | EugenMayer: are you describing two issues or one? if the former, then let's not complicate diagnosing this one | |
| 17:30:47 | sean-k-mooney | but really you shoudl try to avoid removing flavor or image that are in use | |
| 17:31:08 | EugenMayer | well i cannot tell why tf openstack providere deletes the flav too early or whatever happens in detail (i did not check the sequence in the code yet) | |
| 17:31:23 | sean-k-mooney | EugenMayer: it should not delete it at all | |
| 17:31:31 | EugenMayer | dansmith sorry, my bad. second issue (the latter one with the flav) | |
| 17:31:44 | sean-k-mooney | it sould like they are implementing the hacky workflow that horizon use to have | |
| 17:31:48 | dansmith | EugenMayer: yeah, not helping :) | |
| 17:32:00 | EugenMayer | dansmith sorry. my bad. | |
| 17:32:11 | sean-k-mooney | where they allowed you to update a flavor by deletign and recating it but ya lets not talk about that issue now | |
| 17:35:16 | EugenMayer | well if you ask me to the state error - one should not mark the instance as 'error' if a image_backup task failed - there is no reason for that. Creating a glance image does not required the instance to shutdown or similar, this said, i assume both task (the instance running) and the creation of the image can work in parallel and are independent | |
| 17:35:54 | dansmith | EugenMayer: going to error state is just the nova convention (in most places) | |