Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-12
09:03:40 sean-k-mooney its kind of a pendantic distinction but i dont quite consider them equal
09:03:47 sean-k-mooney you could argue it either way
09:04:15 sean-k-mooney so i would not be agaisn allowing stop ot work in a host dwonstate by the way
09:04:30 sean-k-mooney you woudl update the db and treat it kind of like local delete
09:04:43 sean-k-mooney when the compute agent comes back up it woudl reconsile the vm state
09:05:13 sean-k-mooney if you stoped it then evacuated that would solve sahid's case
09:05:20 opendevreview Amit Uniyal proposed openstack/nova master: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/854499
09:05:21 opendevreview Amit Uniyal proposed openstack/nova master: [compute] always set instnace.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/791135
09:05:43 sahid sean-k-mooney: yes it's also a possibility
09:06:56 sean-k-mooney the one thing to keep in mind i guess is that even if we allow stop
09:07:08 sean-k-mooney it doen not chnage the responsiblity for the admin
09:07:28 sean-k-mooney that is you are requried as an admin to ensure a host is fenced or all vms are stopped before you evacuate
09:08:00 sean-k-mooney if we allow stop in a down host state the admin still need to ensure it is stoped to prevent data currpption
09:08:17 sean-k-mooney but if they can then that woudl allwo them to evacuate without start the vm again
09:08:24 gibi this is why I would connect the stopping to the evacuation action, that way it is clear that on the source host it is not stopped
09:08:51 sean-k-mooney ack ya that cleaner
09:09:04 sean-k-mooney and the existing check for is it safe to evacute woudl also be checked
09:09:13 gibi yes
09:09:15 sean-k-mooney e.g. the heatbeat has been missed or you set force_down
09:09:54 sean-k-mooney i think johnthetubaguy expressed interest in this in the past
09:10:17 sean-k-mooney or at least supprot for the people that were askign for it in the past
09:11:09 sean-k-mooney oh that reminds me i guess we are not merging my default change?
09:11:34 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/830829
09:12:36 sean-k-mooney i would either like to merge that before RC1 or after we create the stable branch
09:17:38 sahid thank you guys - are we agree to extend evacuate with target state (AsBefore, Stopped) ? Can I report a bug with destailled description or should I share a spec?
09:21:17 sean-k-mooney sahid: all api changes require a spec regardless of how trivial
09:21:25 sean-k-mooney this would need a spec and a new microverion
09:22:24 sean-k-mooney it can be a pretty short spec but there will at least be a conductor rpc change and likely a compute one too pass the target state
09:22:24 sahid ack I make this happen for A.
09:22:43 sahid sure no worries
09:23:52 sean-k-mooney sahid: the repo is open for sepc reviews so whenever you have time feel free to submit one
09:24:44 sahid +1
09:29:37 bauzas sahid, gibi, sean-k-mooney: tbh, I think we discussed about the host-evacuate support before, and we said this was close to orchestration as a client could be doing it
09:29:58 bauzas so the consensus is to tell that some client or script could be doing it
09:30:10 bauzas (or Heat, or whatever else)
09:31:22 sean-k-mooney bauzas: yep
09:31:43 sean-k-mooney bauzas: but what we were discussiong regaring a spec wa allwoing the api to accept a target state to evacuate too
09:32:03 sean-k-mooney oneof:active,poweredoff or current
09:32:20 sean-k-mooney where current or not specified means what we do today
10:10:31 zigo gibi: oslo.concurrency 5.0.1 hangs when I try to run its unit tests, which is probably related to your patch (at: https://review.opendev.org/c/openstack/oslo.concurrency/+/855714 ). Any idea what's going on? Do I need the latest Eventlet (I know we're lagging one minor version behind)?
10:12:17 zigo Ah no, I'm even with 0.30.2, maybe that's why...
10:14:24 gibi zigo: do you know where it is hanging? which test acse?
10:14:25 gibi case
10:15:13 zigo Hard to tell...
10:16:05 zigo Last output was: https://paste.opendev.org/show/bsDsX7d0gw1SDAnLcRMA/
10:16:20 zigo Then it hangged ...
10:16:36 sean-k-mooney whats that form?
10:16:49 zigo sean-k-mooney: building oslo.concurrency.
10:17:23 sean-k-mooney oh i disconnected and reconnected for a bit
10:17:32 sean-k-mooney missed the start of your conversation with gibi i think
10:18:11 zigo Problem is: I have autopkgtest issues in Eventlet ... 0.33.x :/
10:18:47 sean-k-mooney https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L184-L201
10:20:10 sean-k-mooney i see
10:20:28 sean-k-mooney so thats failing whiel building in some cases
10:20:58 sean-k-mooney im not really sure who/why
10:21:01 gibi zigo: could you point to how you run the unti tests?
10:21:31 sean-k-mooney the est is just concorting an instnace of an exeption class
10:21:58 sean-k-mooney asserting when you call str() on it that it contians the message
10:22:10 gibi there are unit tests in the repo that can be run with and without eventlet monkey patching but there are tests that can only run in eventlet
10:22:13 gibi hence the https://github.com/openstack/oslo.concurrency/blob/01cf2ffdf48c21f886b2aa3f766be5d268248c18/tox.ini#L14-L15
10:22:30 sean-k-mooney maybe this print is the issue https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_processutils.py#L203
10:22:49 zigo PYTHON=python3 stestr run --subunit | subunit2pyunit
10:22:49 zigo I simply do this:
10:23:11 gibi I can imagine that if you run the eventlet aware test without eventlet monkey patching then the eventlet only test might missbehave
10:23:38 zigo I'll try further and let you know where it leads me.
10:24:07 sean-k-mooney if you do things like eventlet.spwawn directly without monkeypatching
10:24:17 sean-k-mooney you need to manually invoke the event loop to have it run
10:24:33 sean-k-mooney we saw that in the nova-api when we added scater gather
10:25:20 gibi zigo: based on that command line you run without eventlet monkey patching but you still run tests from https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/unit/test_lockutils_eventlet.py
10:26:07 gibi zigo: you can try not running those test to see if that resolve the hang
10:26:30 sean-k-mooney you could fix them by adding eventlet.sleep(seconds=0)
10:26:48 sean-k-mooney i think that will make it work if not monkeypatched
10:27:05 sean-k-mooney but ya not runnign them would be better
10:27:24 gibi there is the place where the test monkey patches selectively https://github.com/openstack/oslo.concurrency/blob/master/oslo_concurrency/tests/__init__.py
10:28:10 sean-k-mooney https://github.com/openstack/oslo.concurrency/blob/5397838f4117300a509bff474dfcdd60b5993677/oslo_concurrency/tests/unit/test_lockutils_eventlet.py#L49
10:28:23 sean-k-mooney so if we also do that on lin 51 before the pool.waitall()
10:28:29 sean-k-mooney i think that might actully work in either case
10:28:55 sean-k-mooney but i might make sense for use to use the skip funciton in the classs
10:29:02 sean-k-mooney to have it skip in not monkey patched
10:35:57 sean-k-mooney gibi: is there any reason not to check if we are monkey patched in the setup function and call skipTest
10:40:21 gibi I don't know we might need to consult with oslo cores
10:40:59 gibi because there was eventlet specific tests before I added mine I thought it is handled centrally not to run them in a non patched env
10:41:06 gibi this might not be the case
10:41:31 sean-k-mooney i dont see that logic genericlly
10:41:44 sean-k-mooney its certenly possible to add
11:20:24 noonedeadpunk hello folks! I was wondering - does reverting resize failure rings anybody a bell? Ie - create server, resize server, revert resize -> VM is "stuck" in REVERT_RESIZE until message timeouts, then it goes back to VERIFY_RESIZE but with original flavor, and then nova-compute shutdown VM on hypervisor at all. The only way to recover is to reset state
11:21:51 sean-k-mooney not that i recall but i can see that poteilaly happening if we raise an excption in the revert path and cant proceed
11:22:17 sean-k-mooney like if the souce host was down or soemthing like that we would not be able to revert
11:22:58 noonedeadpunk paste: https://paste.openstack.org/show/bh3kML89sPFYDN9HkDSd/
11:23:54 noonedeadpunk sean-k-mooney: to have that said, out of 100 tempest runs of tempest.api.compute.servers.test_server_actions.ServerActionsTestJSON.test_resize_server_revert 57 has failed
11:24:47 noonedeadpunk the stack trace I've spotted: https://paste.openstack.org/show/bytEWO0CHe8cSVktE3tv/
11:25:43 noonedeadpunk I assumed it can be related to the heartbeat_in_pthread thing, as in the region it was "default" setting on Xena (which is enabled), but disabling it didn't fix that
11:27:33 sean-k-mooney do you have any timeouts form ovsdbapp
11:27:52 sean-k-mooney in the nova-compute logs
11:30:01 noonedeadpunk sean-k-mooney: um, nope
11:30:10 noonedeadpunk it's ovs, not ovn fwiw
11:30:22 sean-k-mooney ya it would be the same either way
11:30:48 sean-k-mooney https://bugzilla.redhat.com/show_bug.cgi?id=2085583
11:30:56 sean-k-mooney i was wondering if it was related to that
11:32:33 sean-k-mooney those are still pending backport upstream https://review.opendev.org/c/openstack/os-vif/+/841771

Earlier   Later