Earlier  
Posted Nick Remark
#openstack-nova - 2020-01-22
19:10:57 melwitt efried: this is stable/pike so I doubt it right? the PyYAML is a new thing?
19:11:11 efried oh
19:11:17 efried I don't know. Is devstack branchless or something?
19:11:25 sean-k-mooney efried: here is an example job i need to go update https://review.opendev.org/#/c/679656/12 but it show you how to create a custom job with custom flavor and run standard tempest test to validate thing
19:11:25 melwitt gah gzipped logs again
19:11:33 sean-k-mooney efried: no devstack is branched
19:11:43 melwitt efried: devstack has branches but tempest is branchless
19:11:53 sean-k-mooney yes ^
19:12:21 efried melwitt: okay, that failure looks different, ignore me.
19:12:40 melwitt how are these old crusty jobs having the friggin gzipped logs, dang
19:12:59 efried thanks sean-k-mooney
19:13:29 efried oh, yeah melwitt, that's going to continue to be a problem I guess, unless we plan to migrate legacy jobs to zv3 on stable. Much sad face.
19:13:44 efried we may have to figure out how to get that turned off in the job itself for stable.
19:13:45 melwitt oh geeez :''''''(
19:13:51 sean-k-mooney am well it should not be
19:13:59 sean-k-mooney oh actully
19:14:05 efried I thought it was an infra-level thing.
19:14:11 sean-k-mooney ya it is
19:14:37 efried so, if we're going to have to figure that out anyway, I guess we might as well do it on master, asap, so we can stop being frustrated at least partially.
19:14:39 sean-k-mooney so we need to modify devstack-gate to stop compressing the logs internally
19:14:46 sean-k-mooney and backport that to stable branches
19:15:18 melwitt isn't that what you did though? it didn't stop the compression for legacy jobs on master
19:15:30 melwitt so even if we backport it, it won't help right?
19:15:33 sean-k-mooney my fix fixed non legacy jobs
19:16:31 sean-k-mooney i intended it to fix both but it seams that devstack gate is compressing them somewhere else too
19:16:36 sean-k-mooney or the legacy jobs are
19:16:43 sean-k-mooney it might not be in devstack gate
19:16:54 melwitt got it
19:17:38 melwitt ok, now to find where this thing is failing LM
19:18:41 sean-k-mooney i have been doing " curl <log url> | zcat | lnav -q"
19:19:45 melwitt that's helpful, thanks
19:21:56 sean-k-mooney melwitt: artom was say you can also to "curl <url> | zless" but i like using lnav to browse logs
19:23:13 melwitt I still haven't gotten around to trying lnav so I'll use zless for now
19:23:47 sean-k-mooney ya if you do try it by default lnav dumps all the logs to std out so the -q is there to stop that
19:24:16 melwitt \:| [instance: 5bca844b-4ca6-4c63-a579-8f00316da019] Instance spawn was interrupted before instance_claim, setting instance to ERROR state
19:24:51 sean-k-mooney am that is new
19:24:59 melwitt yeah, it's new to me
19:25:18 sean-k-mooney does that mean qemu crashed or something?
19:25:25 melwitt no idea
19:25:51 sean-k-mooney well that makes two of us
19:26:34 melwitt I don't expect it's related to the patch, and the patch has been sitting around forever so I wouldn't be surprised if it needs a rebase by now. but the nova-live-migration job is failing pretty consistently on it and so far, for unknown reasons
19:26:58 melwitt rechecked it like 3 times and always fail on nova-live-migration
19:27:59 sean-k-mooney i can take a look at it a bit later. what is the review?
19:28:27 sean-k-mooney once i fix the osp 15 backport im doing ill take a quick look before moveing to osp 13
19:28:30 melwitt sean-k-mooney: it's this one https://review.opendev.org/683008
19:28:56 sean-k-mooney cool i have it open.
19:28:58 melwitt note that it is very old, not rebased since september
19:29:00 artom sean-k-mooney, huh, that's the same thing we saw in whitebox
19:29:25 artom We thought it was because we were restarting nova-compute to change configs
19:29:31 sean-k-mooney melwitt: if you do a rebase via the ui i should not hurt
19:29:39 sean-k-mooney but i dont think that is the issue
19:29:53 sean-k-mooney zuul will do a merge with master so it should basicaly be the same
19:29:56 melwitt yeah, probably worth it to try. it's gonna have to be rebased to merge anyway
19:30:06 melwitt let me do it now
19:30:39 sean-k-mooney well its more or less the same as a recheck and will avoid the merge commit so it wont hurt
19:30:49 sean-k-mooney artom: did you fix it in whitebox?
19:31:07 artom sean-k-mooney, by extending the service restart wait delay
19:31:21 artom Which apparently has nothing to do with it, as we're now seeing the same thing in Nova?
19:31:27 sean-k-mooney we are not restart services in this job
19:31:31 sean-k-mooney actully
19:31:35 sean-k-mooney we might be
19:31:39 openstackgerrit melanie witt proposed openstack/nova stable/pike: Avoid redundant initialize_connection on source post live migration https://review.opendev.org/683008
19:31:44 artom Wait, n-l-m?
19:31:46 sean-k-mooney i think we do
19:31:52 sean-k-mooney ya
19:32:03 artom We're changing some configs IIRC
19:32:04 sean-k-mooney i think there is a hack that swaps the storage to ceph or something
19:32:05 artom So yeah, we are
19:32:41 sean-k-mooney i think its in the post run playbook
19:33:07 sean-k-mooney ok its not there
19:33:15 sean-k-mooney its proably in the test hook
19:34:56 sean-k-mooney ya so here https://github.com/openstack/nova/blob/master/gate/live_migration/hooks/run_tests.sh#L55-L67
19:35:14 artom Ah, yeah, configure_and_start_nova
19:35:53 artom Are... are we just going to hax a sleep in there? Before the run_tempest call?
19:35:54 sean-k-mooney https://github.com/openstack/nova/blob/master/gate/live_migration/hooks/ceph.sh#L71-L100
19:36:33 sean-k-mooney we could hack a sleep here https://github.com/openstack/nova/blob/master/gate/live_migration/hooks/ceph.sh#L114
19:37:26 artom Why in ceph?
19:37:48 sean-k-mooney this is where we restart nova after reconfiguring to use ceph
19:38:36 sean-k-mooney so in the ligration job we deploy withour ceph intially run tempest then after that we reconfivure the deploymetn for ceph and do more testing
19:40:11 sean-k-mooney so we only restart the serivce in the ceph teting at the end here https://github.com/openstack/nova/blob/master/gate/live_migration/hooks/run_tests.sh#L64
19:40:43 sean-k-mooney that function get loaded form the ceph.sh file here https://github.com/openstack/nova/blob/master/gate/live_migration/hooks/run_tests.sh#L21
19:53:26 stephenfin sean-k-mooney: Seeing as you're still here, does my comment here make sense https://review.opendev.org/#/c/662522/13/nova/compute/api.py@3637
19:53:40 stephenfin checking it tomorrow is fine too. I am outta here
19:56:18 sean-k-mooney ill check
19:58:06 sean-k-mooney am the conducor calls the schduer
19:58:30 sean-k-mooney so it could modify the numa toplogy before it calls the select destination function i think
19:59:51 sean-k-mooney i would have to check but you are right in that you definetly need to recalulate the numa toployg based on the new flavor before the numa toplogy filter runs
20:00:20 sean-k-mooney meaining either in the api or in the conductor if the conductor si what is invoking the scheduler.
20:01:27 sean-k-mooney i dont have that part of code loaded up in my brain right now hence vague answer
21:54:31 lucidguy I'll be very very appreciative if someone can assist me with this, i've spent hours googling etc. https://paste.ubuntu.com/p/6wqhRX68tr/
21:56:01 sean-k-mooney you are booting vms with 1TB of ram
21:56:30 sean-k-mooney that is fun
21:57:12 sean-k-mooney this looks like a kvm kernel bug
21:59:35 lucidguy sean-k-mooney: Suggestions?
22:01:11 sean-k-mooney ubutun 16.04 is getting near the end of its life have you updated to the latest 16.04 kernel available on the host.
22:01:18 sean-k-mooney its possible this is already fixed if not
22:01:26 sean-k-mooney this is not an openstack related bug
22:02:00 sean-k-mooney so your best bet is to ralk to the ubuntu kernel folks or reach our to the kvm comunity on irc
22:02:53 sean-k-mooney personlly if you are stuck on 16.04 but can upgrade your kernel i would consider using there hardware enabling kernel which tends to be more up to date
22:05:39 lucidguy sean-k-mooney: We are already using the hwe kernel

Earlier   Later