Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-03
14:42:26 mriedem we shouldn't flat out expose system metadata
14:42:43 mriedem "if you want to query the point in time properties that where inherited from an image during the launch."
14:42:53 mriedem expose those as some other field then
14:43:16 mriedem we don't need to expose all of the garbage we carry in system metadata
14:43:20 mriedem like old_vm_state during a resize
14:44:03 mriedem bauzas: i backported the placement fix to overwrite allocations https://review.openstack.org/#/c/490231/
14:46:24 sdague ok, I'm good with Won't Fix
14:47:27 openstackgerrit Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954
14:47:34 mriedem well, alternatively the fix is a microversion to expose some specific part of system metadata and what that entails
14:47:51 mriedem sdague: ever thought about indexing qemu instance logs in our ci runs?
14:48:05 mriedem when live migration jobs fail, a lot of the time it's due to
14:48:05 mriedem http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz
14:48:11 mriedem /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed
14:48:17 mriedem but ^ isn't exposed in anything we index
14:48:25 jaypipes dansmith: did you see cdent just pushed a revision on the confirm resize patch?
14:48:59 dansmith jaypipes: a bit ago while we were talking, yeah. he said so and that's when I pulled to start working
14:49:14 jaypipes gotcha. just making sure you noticed. carry on.
14:49:48 mriedem what i'd really love is if the libvirt / qemu job had some way to get those details from the guest
14:50:08 mriedem kashyap: mdbooth: you know how during a live migration we're checking the domain job status to see when it completes, or if it fails?
14:50:19 mdbooth mriedem: Yep
14:50:23 mriedem is there any way to get the qemu guest logs when that fails, like http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz
14:50:38 mriedem i really want: /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed
14:51:13 mriedem when ^ happens, the only failure we get in the n-cpu logs is that on the destination when we're doing post-live migration at destination, the instance (guest domain) isn't found
14:51:18 mriedem because it blew up on the source side
14:52:21 mdbooth mriedem: Is ^^^ from dest?
14:53:52 mriedem no that's source
14:53:58 mriedem here is another one http://logs.openstack.org/66/483566/10/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/0437fbe/logs/subnode-2/libvirt/qemu/instance-00000011.txt.gz
14:54:07 mriedem different error, but results in the same kind of thing in the dest n-cpu logs
14:54:13 mriedem InstanceNotFound during post live migration at destination
14:54:15 mriedem b/c it failed on the source
14:54:27 cfriesen mriedem: what's the complication with getting that file from the dest?
14:55:56 mriedem https://bugs.launchpad.net/nova/+bug/1706377
14:55:56 openstack Launchpad bug 1706377 in OpenStack Compute (nova) "(libvirt) live migration fails on source host due to "Assertion `!(bs->open_flags & BDRV_O_INACTIVE)' failed."" [Undecided,Confirmed]
14:55:58 mdbooth mriedem: Why are we calling post if the migration failed?
14:56:09 mriedem mdbooth: because libvirt told us the job was complete
14:56:15 mriedem see my notes in https://bugs.launchpad.net/nova/+bug/1706377
14:56:23 mdbooth mriedem: *That's* the bug
14:56:31 mdbooth mriedem: And we already kinda knew about that, right?
14:56:55 mdbooth Didn't I leave a comment in there to that effect?
14:57:01 mriedem in where?
14:57:10 mdbooth libvirt/drive
14:57:11 mdbooth r
14:57:44 mdbooth mriedem: Sorry, libvirt/guest.py
14:57:46 mdbooth is_job_complete
14:58:08 dansmith jaypipes: mriedem: okay I got the resource override stuff working in jaypipes' patch and some unified code between them for doubling/undoubling resources, so now I'm going to look at the peripheral test failures
14:58:14 mdbooth mriedem: It's there in one of my trademark big blocks of comment
14:58:19 dansmith I have about 30 minutes until my next meeting so I will push ahead of that regardless of my progress
14:58:34 mdbooth # Secondly, with the current method we only know that 'no job'
14:58:34 mdbooth # indicates completion. It does not necessarily indicate successful
14:58:34 mdbooth # completion: the job could have failed, or been cancelled. When
14:58:34 mdbooth # polling for block job info we have no way to detect this, so we
14:58:34 mdbooth # assume success.
14:58:52 jaypipes dansmith: k. I'm happy to take the baton on fixing periphery tests when you go to your meeting.
14:58:52 mriedem ah ok,
14:59:00 mriedem that was written around the time of the great swap volume rewrite
14:59:29 dansmith jaypipes: ack
14:59:39 mriedem cfriesen: i don't understand your question
14:59:48 mriedem cfriesen: the migration completes but actually fails on the source,
15:00:16 mriedem but we don't know it fails, we just know the job is 'complete' so we tell dest to do post live migration stuff, and when it does, the guest never made it to dest (or it was deleted by libvirt when it found that the source failed)
15:00:35 mriedem so i'm trying to figure out a way to get the qemu instance logs into the n-cpu logs for debug
15:00:39 mdbooth mriedem: I think we should rewrite that polling block to consume events instead. It's also less buggy.
15:00:45 cfriesen mriedem: I was just thinking that we had all the info needed to get the file, so didn't see what the problem was....but it's not "can we get the file", but "can we determine there was a failure so we know to go get the file"
15:00:50 mdbooth As in, it was designed for this in the first place.
15:01:27 mriedem cfriesen: maybe, i don't know how configurable that path is
15:01:53 mriedem seems pretty hacky though, i'd think you could get qemu guest logs from libvirt apis
15:02:12 mdbooth mriedem: I don't think so, btw.
15:02:44 cfriesen mriedem: ah, right, we don't control all the clouds this runs on. I think it is configurable where those logs go.
15:03:05 mriedem right
15:03:13 mriedem that's why i'd need an api
15:04:10 mdbooth Basically we should switch to using libvirt events api. Extensive documentation here: http://libvirt.org/docs/libvirt-appdev-guide/en-US/html/Application_Development_Guide-Guest_Domains-Event_Not.html
15:04:33 cfriesen with a big TBD on that page?
15:04:35 mriedem mdbooth: yeah i suppose virConnectDomainEventJobCompletedCallback
15:04:50 mdbooth cfriesen: You need more? Pshaw
15:05:52 mdbooth cfriesen: It's a small TBD, anyway. Classier that way.
15:07:19 mdbooth mriedem: I wonder if we could register a libvirt error handler, and dump errors into nova compute logs as a matter of course:http://libvirt.org/docs/libvirt-appdev-guide-python/en-US/html/libvirt_application_development_guide_using_python-Error_Handling-Registering_Error_Handler.html
15:07:31 mdbooth That might achieve what you want in practise.
15:11:14 mriedem where is the error array defined?
15:11:33 mdbooth mriedem: rtfs
15:11:47 cfriesen mriedem: mdbooth: is there a libvirt bug here? I mean the source is running _live_migration_monitor() and calling guest.get_job_info(). shouldn't libvirt detect a failure?
15:11:48 mriedem ha, we already register an error handler
15:11:49 mriedem def _libvirt_error_handler(context, err):
15:11:49 mriedem # Just ignore instead of default outputting to stderr.
15:11:49 mriedem pass
15:12:02 mdbooth mriedem: hehe
15:12:07 cfriesen and if it doesn't, are we going to get an error in the callback?
15:12:20 mdbooth cfriesen: No
15:12:30 mriedem "with error being a list of information about the error being raised. "
15:12:59 mriedem i suppose it's similar to a libvirtError
15:13:39 mdbooth cfriesen: I don't recall the detail now, but at the time kashyap and I went over the libvirt and libvirt python binding code very carefully
15:13:52 mdbooth cfriesen: We're extracting everything from it which can be extracted
15:14:06 mriedem ah yup it's just the libvirtError.err list
15:14:09 mdbooth Hence my big comment explaining what we're not getting
15:15:05 mdbooth It's not designed to be used this way
15:15:38 mdbooth The intention was that you'd consume events instead. That api has been stable for much longer.
15:17:14 sdague mriedem: you'd need another grok parser
15:19:51 cfriesen mdbooth: okay, I think I got it. On another note, currently with block live migration if the guest is dirtying disk quickly the current logs don't show information about the initial block transfer.
15:20:15 mriedem would be nice if the libvirtError python binding class just had a nice __repr__
15:20:22 mriedem maybe it does already...
15:21:48 openstackgerrit Dan Smith proposed openstack/nova master: remove provider allocs in confirm/revert resize https://review.openstack.org/488510
15:21:49 openstackgerrit Dan Smith proposed openstack/nova master: Sum allocations in the scheduler when resizing to the same host https://review.openstack.org/490085

Earlier   Later