| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-12-08 | |||
| 17:07:47 | mriedem | now that volume backed live migration is fixed.... | |
| 17:07:50 | mriedem | it is? | |
| 17:08:09 | mriedem | you mean https://review.openstack.org/#/c/507202/ ? | |
| 17:08:35 | melwitt | yeah, I thought that fixed it. you had those DNM patches depend on it and the unskipped tests passed | |
| 17:08:46 | melwitt | did I miss something? | |
| 17:08:52 | mriedem | melwitt: i believe those tests will intermittently pass regardless of that patch | |
| 17:09:06 | melwitt | okay, that's what I was missing then | |
| 17:09:09 | mriedem | the volume-backed and iscsi live migration tests were skipped in tempest long ago due to random failures | |
| 17:09:17 | mriedem | that no one could sleuth because qemu | |
| 17:09:29 | mriedem | using UCA now might have changed that | |
| 17:09:31 | melwitt | I see | |
| 17:09:43 | mriedem | bug 1524898 | |
| 17:09:44 | openstack | bug 1524898 in OpenStack Compute (nova) "Volume based live migration aborted unexpectedly" [High,In progress] https://launchpad.net/bugs/1524898 | |
| 17:10:11 | melwitt | okay, I thought maybe that was related to the recent fix we had | |
| 17:10:20 | mriedem | even if we enable the test in tempest again, | |
| 17:10:33 | mriedem | tempest is branchless so it would run on stable branches and could fail there if we're using older packages that had a problem | |
| 17:10:37 | mriedem | hard to tell | |
| 17:10:53 | mriedem | i would very much like to get volume-backed live migration testing enabled in the gate though yes | |
| 17:10:54 | melwitt | oh, right. | |
| 17:11:08 | mriedem | we should probably change the global skip on the test in tempest to a config option, | |
| 17:11:11 | mriedem | so that we can enable the test per-branch | |
| 17:11:24 | melwitt | I can't recall if we skip certain tests for stable-only, or if that's even possible | |
| 17:11:36 | mriedem | we can configure tempest per-branch, per-job | |
| 17:11:43 | mriedem | we can configure tempest in our in-tree live migration scripts | |
| 17:11:46 | mriedem | the post_test_hook | |
| 17:12:07 | mriedem | https://github.com/openstack/nova/blob/master/nova/tests/live_migration/hooks/run_tests.sh#L33 | |
| 17:12:19 | mriedem | with the magic of ansible | |
| 17:13:18 | melwitt | okay, so we'd add another option volume_backed_live_migration or such | |
| 17:13:25 | mriedem | so if you wanted to be cool, propose a change to tempest to add a config option for volume-backed live migration to replace that skip line, default to false for stable, and then set it to true in our live migration job post_test_hook script | |
| 17:13:44 | mriedem | and then be available when it melts the gate | |
| 17:14:09 | melwitt | can do | |
| 17:14:56 | melwitt | thanks | |
| 17:16:18 | mriedem | it is my pleasure | |
| 17:16:27 | mriedem | that's a weird phrase | |
| 17:18:18 | mriedem | dmsimard: interesting, even when it appears that we do set force_tcg mode, guestfs returns None rather than an int | |
| 17:18:19 | mriedem | http://logs.openstack.org/50/524750/1/check/legacy-tempest-dsvm-neutron-full-centos-7/a7f051e/logs/screen-n-cpu.txt.gz#_Dec_04_13_43_19_923098 | |
| 17:18:21 | melwitt | lots of zvm:recheck spam going on | |
| 17:18:25 | mriedem | Failed to force guestfs TCG mode. guestfs_set_backend_settings returned: None | |
| 17:18:45 | dmsimard | weird | |
| 17:19:43 | dmsimard | mriedem: it seems like there's two set_backend_settings happening there | |
| 17:19:57 | dmsimard | event=trace eh=0 buf='set_backend_settings "force_tcg"' array=[] | |
| 17:20:05 | dmsimard | and event=trace eh=0 buf='set_backend_settings = 0' array=[] ? | |
| 17:21:21 | mriedem | i think one is the value being set? | |
| 17:21:32 | mriedem | or the return value? | |
| 17:21:41 | mriedem | i don't know what the array is | |
| 17:21:46 | mriedem | but like "add_drive = 0" | |
| 17:24:27 | mriedem | http://libguestfs.org/guestfs.3.html#guestfs_set_event_callback | |
| 17:33:32 | mriedem | ok this is where we have the eventlet switch that blows up http://logs.openstack.org/50/524750/1/check/legacy-tempest-dsvm-neutron-full-centos-7/a7f051e/logs/screen-n-cpu.txt.gz#_Dec_04_13_43_19_941869 | |
| 17:33:39 | mriedem | that's right in the middle of the guestfs.launch call | |
| 17:33:48 | mriedem | error: cannot switch to a different thread | |
| 17:34:44 | mriedem | heh http://logs.openstack.org/50/524750/1/check/legacy-tempest-dsvm-neutron-full-centos-7/a7f051e/logs/screen-n-cpu.txt.gz#_Dec_04_15_25_29_320753 | |
| 17:34:54 | mriedem | so, i think the logging is what's maybe making it switch | |
| 17:35:13 | mriedem | so you enable guestfs debug to get callbacks from guestfs to log them, | |
| 17:35:19 | mriedem | meanwhile, guestfs is running in a thread | |
| 17:36:18 | mriedem | that doesn't explain why we were hanging before we enabled debug though | |
| 17:36:32 | mriedem | but i can push a patch to only launch in a thread if guestfs.debug is False | |
| 17:38:34 | fried_rice | cdent You going to be around in a couple hours? For a "warm handoff"? | |
| 17:38:41 | mriedem | gross | |
| 17:39:25 | cdent | fried_rice: yessir, I should be around, or at least nearby, make a ping and I’ll (eventually) come running | |
| 17:40:35 | fried_rice | cdent ack | |
| 17:40:59 | fried_rice | cdent Leaving a bit more of a mess than anticipated, I'm afraid. But nothing insurmountable. | |
| 17:49:26 | mnaser | is there anyone here who worked on glance image signature verification? | |
| 17:50:03 | mnaser | we have a change in puppet-nova to add the ability to configure it, but i'd like to know if its a global nova config (aka, api/etc need it) or nova-compute only | |
| 17:50:20 | mnaser | from my understanding of code, the option seems to be used in the glance image download code which probably only happens on computes? | |
| 17:52:09 | mriedem | peter-hamilton: ^ | |
| 17:52:24 | mriedem | mnaser: dane-fichter is who i'd usually go to but haven't seem him in channel in a long time | |
| 17:52:28 | alaski | mriedem: in case this didn't already get answered, instance actions are not synced between a cell and the api level in v1. There are spurious records at the API level that never get used, but it would have taken a lot of effort to not write them there. Requests for instance actions are proxied to the cell and results are returned from there. | |
| 17:52:54 | mnaser | mriedem: ok cool, thanks for the info.. i'll research the code a tad bit more | |
| 17:53:03 | mriedem | alaski: hmm, i'm not sure how @wrap_instance_event in the compute service would ever work then | |
| 17:53:08 | mriedem | alaski: because the api creates the action record | |
| 17:53:15 | mriedem | and the compute tries to look it up to record events against it | |
| 17:53:24 | mriedem | but if the action isn't in the cell db, how are they ever found? | |
| 17:53:34 | superdan | mriedem: they get replayed in the cell I imagine | |
| 17:53:42 | mriedem | superdan: where/how? | |
| 17:53:51 | superdan | so they're down there, but the api never looks at them, only at the stale ones that are left by the start | |
| 17:54:04 | superdan | mriedem: because we replay calls into compute/api down in the cell | |
| 17:54:11 | superdan | which would create them there too | |
| 17:54:37 | mriedem | so if a compute service fails to find an action created in the api, it could just be because of a timing issue? | |
| 17:54:55 | superdan | no, I don't think so, | |
| 17:55:02 | mriedem | ok, was going to say - i find that hard to believe | |
| 17:55:07 | mriedem | because we'd be randomly failing all the time | |
| 17:55:09 | superdan | because it should still get created in the cell before you kick off anything | |
| 17:56:08 | mriedem | ok that doesn't seem to be happening | |
| 17:56:23 | mriedem | i dug through the cellsv1 code for awhile yesterday and don't see anything special about instance actions in there thoguh | |
| 17:56:27 | mriedem | *though, like how they are handled | |
| 17:56:28 | superdan | right | |
| 17:56:31 | superdan | because they're not | |
| 17:57:21 | superdan | mriedem: for example, this one: https://github.com/openstack/nova/blob/master/nova/compute/api.py#L3098-L3099 | |
| 17:57:33 | superdan | mriedem: we run that in the api node as part of the cells_api indirection | |
| 17:57:46 | superdan | mriedem: then we pass it down to the cell and run that same code again in the cell, | |
| 17:57:46 | openstackgerrit | Merged openstack/nova-specs master: Document Queens review priorities https://review.openstack.org/526573 | |
| 17:57:57 | superdan | then we call to compute, which would finish it, but only in the cell | |
| 17:58:01 | superdan | and then you're done | |
| 17:58:07 | mriedem | and that's all synchronous? | |
| 17:58:24 | superdan | if you go look at it from the api, you'd only see the start, because you never finished it in the api cell, and nothing ever sync'd the finish | |
| 17:58:52 | superdan | mriedem: is what synchronous? the api cell before the child cell? sure, but they don't know about each other really anyway, so it doesn't matter | |
| 17:59:58 | mriedem | is the replay of the instance action create from the top level to the child cell synchronous | |
| 18:00:00 | mriedem | was my question | |
| 18:00:10 | mriedem | trying to figure out what is special that makes this fail the cellsv1 job https://review.openstack.org/#/c/523676 | |
| 18:00:15 | mriedem | http://logs.openstack.org/76/523676/11/check/legacy-tempest-dsvm-cells/1d155c6/logs/screen-n-cpu.txt.gz?level=TRACE#_Dec_07_09_02_09_579991 | |