| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-12 | |||
| 09:03:06 | sean-k-mooney | where as shelve/unshleve is moving form on a host to not and vise versa | |
| 09:03:20 | gibi | yeah | |
| 09:03:40 | sean-k-mooney | its kind of a pendantic distinction but i dont quite consider them equal | |
| 09:03:47 | sean-k-mooney | you could argue it either way | |
| 09:04:15 | sean-k-mooney | so i would not be agaisn allowing stop ot work in a host dwonstate by the way | |
| 09:04:30 | sean-k-mooney | you woudl update the db and treat it kind of like local delete | |
| 09:04:43 | sean-k-mooney | when the compute agent comes back up it woudl reconsile the vm state | |
| 09:05:13 | sean-k-mooney | if you stoped it then evacuated that would solve sahid's case | |
| 09:05:20 | opendevreview | Amit Uniyal proposed openstack/nova master: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/854499 | |
| 09:05:21 | opendevreview | Amit Uniyal proposed openstack/nova master: [compute] always set instnace.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/791135 | |
| 09:05:43 | sahid | sean-k-mooney: yes it's also a possibility | |
| 09:06:56 | sean-k-mooney | the one thing to keep in mind i guess is that even if we allow stop | |
| 09:07:08 | sean-k-mooney | it doen not chnage the responsiblity for the admin | |
| 09:07:28 | sean-k-mooney | that is you are requried as an admin to ensure a host is fenced or all vms are stopped before you evacuate | |
| 09:08:00 | sean-k-mooney | if we allow stop in a down host state the admin still need to ensure it is stoped to prevent data currpption | |
| 09:08:17 | sean-k-mooney | but if they can then that woudl allwo them to evacuate without start the vm again | |
| 09:08:24 | gibi | this is why I would connect the stopping to the evacuation action, that way it is clear that on the source host it is not stopped | |
| 09:08:51 | sean-k-mooney | ack ya that cleaner | |
| 09:09:04 | sean-k-mooney | and the existing check for is it safe to evacute woudl also be checked | |
| 09:09:13 | gibi | yes | |
| 09:09:15 | sean-k-mooney | e.g. the heatbeat has been missed or you set force_down | |
| 09:09:54 | sean-k-mooney | i think johnthetubaguy expressed interest in this in the past | |
| 09:10:17 | sean-k-mooney | or at least supprot for the people that were askign for it in the past | |
| 09:11:09 | sean-k-mooney | oh that reminds me i guess we are not merging my default change? | |
| 09:11:34 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/830829 | |
| 09:12:36 | sean-k-mooney | i would either like to merge that before RC1 or after we create the stable branch | |
| 09:17:38 | sahid | thank you guys - are we agree to extend evacuate with target state (AsBefore, Stopped) ? Can I report a bug with destailled description or should I share a spec? | |
| 09:21:17 | sean-k-mooney | sahid: all api changes require a spec regardless of how trivial | |
| 09:21:25 | sean-k-mooney | this would need a spec and a new microverion | |
| 09:22:24 | sahid | ack I make this happen for A. | |
| 09:22:24 | sean-k-mooney | it can be a pretty short spec but there will at least be a conductor rpc change and likely a compute one too pass the target state | |
| 09:22:43 | sahid | sure no worries | |
| 09:23:52 | sean-k-mooney | sahid: the repo is open for sepc reviews so whenever you have time feel free to submit one | |
| 09:24:44 | sahid | +1 | |
| 09:29:37 | bauzas | sahid, gibi, sean-k-mooney: tbh, I think we discussed about the host-evacuate support before, and we said this was close to orchestration as a client could be doing it | |
| 09:29:58 | bauzas | so the consensus is to tell that some client or script could be doing it | |
| 09:30:10 | bauzas | (or Heat, or whatever else) | |
| 09:31:22 | sean-k-mooney | bauzas: yep | |
| 09:31:43 | sean-k-mooney | bauzas: but what we were discussiong regaring a spec wa allwoing the api to accept a target state to evacuate too | |
| 09:32:03 | sean-k-mooney | oneof:active,poweredoff or current | |
| 09:32:20 | sean-k-mooney | where current or not specified means what we do today | |
| 10:10:31 | zigo | gibi: oslo.concurrency 5.0.1 hangs when I try to run its unit tests, which is probably related to your patch (at: https://review.opendev.org/c/openstack/oslo.concurrency/+/855714 ). Any idea what's going on? Do I need the latest Eventlet (I know we're lagging one minor version behind)? | |
| 10:12:17 | zigo | Ah no, I'm even with 0.30.2, maybe that's why... | |
| 10:14:24 | gibi | zigo: do you know where it is hanging? which test acse? | |
| 10:14:25 | gibi | case | |
| 10:15:13 | zigo | Hard to tell... | |
| 10:16:05 | zigo | Last output was: https://paste.opendev.org/show/bsDsX7d0gw1SDAnLcRMA/ | |
| 10:16:20 | zigo | Then it hangged ... | |
| 10:16:36 | sean-k-mooney | whats that form? | |
| 10:16:49 | zigo | sean-k-mooney: building oslo.concurrency. | |
| 10:17:23 | sean-k-mooney | oh i disconnected and reconnected for a bit | |
| 10:17:32 | sean-k-mooney | missed the start of your conversation with gibi i think | |
| 10:18:11 | zigo | Problem is: I have autopkgtest issues in Eventlet ... 0.33.x :/ | |
| 10:18:47 | sean-k-mooney | https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L184-L201 | |
| 10:20:10 | sean-k-mooney | i see | |
| 10:20:28 | sean-k-mooney | so thats failing whiel building in some cases | |
| 10:20:58 | sean-k-mooney | im not really sure who/why | |
| 10:21:01 | gibi | zigo: could you point to how you run the unti tests? | |
| 10:21:31 | sean-k-mooney | the est is just concorting an instnace of an exeption class | |
| 10:21:58 | sean-k-mooney | asserting when you call str() on it that it contians the message | |
| 10:22:10 | gibi | there are unit tests in the repo that can be run with and without eventlet monkey patching but there are tests that can only run in eventlet | |
| 10:22:13 | gibi | hence the https://github.com/openstack/oslo.concurrency/blob/01cf2ffdf48c21f886b2aa3f766be5d268248c18/tox.ini#L14-L15 | |
| 10:22:30 | sean-k-mooney | maybe this print is the issue https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L203 | |
| 10:22:49 | zigo | I simply do this: | |
| 10:22:49 | zigo | PYTHON=python3 stestr run --subunit | subunit2pyunit | |
| 10:23:11 | gibi | I can imagine that if you run the eventlet aware test without eventlet monkey patching then the eventlet only test might missbehave | |
| 10:23:38 | zigo | I'll try further and let you know where it leads me. | |
| 10:24:07 | sean-k-mooney | if you do things like eventlet.spwawn directly without monkeypatching | |
| 10:24:17 | sean-k-mooney | you need to manually invoke the event loop to have it run | |
| 10:24:33 | sean-k-mooney | we saw that in the nova-api when we added scater gather | |
| 10:25:20 | gibi | zigo: based on that command line you run without eventlet monkey patching but you still run tests from https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/unit/test_lockutils_eventlet.py | |
| 10:26:07 | gibi | zigo: you can try not running those test to see if that resolve the hang | |
| 10:26:30 | sean-k-mooney | you could fix them by adding eventlet.sleep(seconds=0) | |
| 10:26:48 | sean-k-mooney | i think that will make it work if not monkeypatched | |
| 10:27:05 | sean-k-mooney | but ya not runnign them would be better | |
| 10:27:24 | gibi | there is the place where the test monkey patches selectively https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/__init__.py | |
| 10:28:10 | sean-k-mooney | https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_lockutils_eventlet.py#L49 | |
| 10:28:23 | sean-k-mooney | so if we also do that on lin 51 before the pool.waitall() | |
| 10:28:29 | sean-k-mooney | i think that might actully work in either case | |
| 10:28:55 | sean-k-mooney | but i might make sense for use to use the skip funciton in the classs | |
| 10:29:02 | sean-k-mooney | to have it skip in not monkey patched | |
| 10:35:57 | sean-k-mooney | gibi: is there any reason not to check if we are monkey patched in the setup function and call skipTest | |
| 10:40:21 | gibi | I don't know we might need to consult with oslo cores | |
| 10:40:59 | gibi | because there was eventlet specific tests before I added mine I thought it is handled centrally not to run them in a non patched env | |
| 10:41:06 | gibi | this might not be the case | |
| 10:41:31 | sean-k-mooney | i dont see that logic genericlly | |
| 10:41:44 | sean-k-mooney | its certenly possible to add | |
| 11:20:24 | noonedeadpunk | hello folks! I was wondering - does reverting resize failure rings anybody a bell? Ie - create server, resize server, revert resize -> VM is "stuck" in REVERT_RESIZE until message timeouts, then it goes back to VERIFY_RESIZE but with original flavor, and then nova-compute shutdown VM on hypervisor at all. The only way to recover is to reset state | |
| 11:21:51 | sean-k-mooney | not that i recall but i can see that poteilaly happening if we raise an excption in the revert path and cant proceed | |
| 11:22:17 | sean-k-mooney | like if the souce host was down or soemthing like that we would not be able to revert | |
| 11:22:58 | noonedeadpunk | paste: https://paste.openstack.org/show/bh3kML89sPFYDN9HkDSd/ | |
| 11:23:54 | noonedeadpunk | sean-k-mooney: to have that said, out of 100 tempest runs of tempest.api.compute.servers.test_server_actions.ServerActionsTestJSON.test_resize_server_revert 57 has failed | |
| 11:24:47 | noonedeadpunk | the stack trace I've spotted: https://paste.openstack.org/show/bytEWO0CHe8cSVktE3tv/ | |
| 11:25:43 | noonedeadpunk | I assumed it can be related to the heartbeat_in_pthread thing, as in the region it was "default" setting on Xena (which is enabled), but disabling it didn't fix that | |
| 11:27:33 | sean-k-mooney | do you have any timeouts form ovsdbapp | |
| 11:27:52 | sean-k-mooney | in the nova-compute logs | |
| 11:30:01 | noonedeadpunk | sean-k-mooney: um, nope | |
| 11:30:10 | noonedeadpunk | it's ovs, not ovn fwiw | |
| 11:30:22 | sean-k-mooney | ya it would be the same either way | |
| 11:30:48 | sean-k-mooney | https://bugzilla.redhat.com/show_bug.cgi?id=2085583 | |