Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-03
14:32:50 sdague bauzas: why did you mark https://bugs.launchpad.net/nova/+bug/1707160 as critical even though you didn't think it was a nova bug?
14:32:50 openstack Launchpad bug 1707160 in neutron "test_create_port_in_allowed_allocation_pools test fails on ironic grenade" [Critical,Confirmed] - Assigned to Ihar Hrachyshka (ihar-hrachyshka)
14:32:53 edleafe In Pike, it should just have the ironic custom RC
14:33:29 bauzas sdague: just for getting traction
14:33:38 bauzas because it's a gate issue
14:34:00 bauzas but anyway
14:34:18 mdbooth jaypipes: Thanks, also for the attaboy ;)
14:34:31 jaypipes mdbooth: heh :)
14:34:53 sdague bauzas: ok, I thought we save critical for must fix rc bugs
14:35:14 bauzas sdague: np, your modification is good to me
14:37:06 mriedem melwitt: some suggestions in https://review.openstack.org/#/c/470578/
14:38:12 melwitt mriedem: cool, thanks
14:39:56 sdague I'm assuming this would need a spec - https://bugs.launchpad.net/nova/+bug/1708458 ?
14:39:56 openstack Launchpad bug 1708458 in OpenStack Compute (nova) "Expose instance system_metadata in compute API" [Undecided,New]
14:42:01 mriedem sdague: jesus yes
14:42:26 mriedem we shouldn't flat out expose system metadata
14:42:43 mriedem "if you want to query the point in time properties that where inherited from an image during the launch."
14:42:53 mriedem expose those as some other field then
14:43:16 mriedem we don't need to expose all of the garbage we carry in system metadata
14:43:20 mriedem like old_vm_state during a resize
14:44:03 mriedem bauzas: i backported the placement fix to overwrite allocations https://review.openstack.org/#/c/490231/
14:46:24 sdague ok, I'm good with Won't Fix
14:47:27 openstackgerrit Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954
14:47:34 mriedem well, alternatively the fix is a microversion to expose some specific part of system metadata and what that entails
14:47:51 mriedem sdague: ever thought about indexing qemu instance logs in our ci runs?
14:48:05 mriedem when live migration jobs fail, a lot of the time it's due to
14:48:05 mriedem http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz
14:48:11 mriedem /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed
14:48:17 mriedem but ^ isn't exposed in anything we index
14:48:25 jaypipes dansmith: did you see cdent just pushed a revision on the confirm resize patch?
14:48:59 dansmith jaypipes: a bit ago while we were talking, yeah. he said so and that's when I pulled to start working
14:49:14 jaypipes gotcha. just making sure you noticed. carry on.
14:49:48 mriedem what i'd really love is if the libvirt / qemu job had some way to get those details from the guest
14:50:08 mriedem kashyap: mdbooth: you know how during a live migration we're checking the domain job status to see when it completes, or if it fails?
14:50:19 mdbooth mriedem: Yep
14:50:23 mriedem is there any way to get the qemu guest logs when that fails, like http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz
14:50:38 mriedem i really want: /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed
14:51:13 mriedem when ^ happens, the only failure we get in the n-cpu logs is that on the destination when we're doing post-live migration at destination, the instance (guest domain) isn't found
14:51:18 mriedem because it blew up on the source side
14:52:21 mdbooth mriedem: Is ^^^ from dest?
14:53:52 mriedem no that's source
14:53:58 mriedem here is another one http://logs.openstack.org/66/483566/10/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/0437fbe/logs/subnode-2/libvirt/qemu/instance-00000011.txt.gz
14:54:07 mriedem different error, but results in the same kind of thing in the dest n-cpu logs
14:54:13 mriedem InstanceNotFound during post live migration at destination
14:54:15 mriedem b/c it failed on the source
14:54:27 cfriesen mriedem: what's the complication with getting that file from the dest?
14:55:56 mriedem https://bugs.launchpad.net/nova/+bug/1706377
14:55:56 openstack Launchpad bug 1706377 in OpenStack Compute (nova) "(libvirt) live migration fails on source host due to "Assertion `!(bs->open_flags & BDRV_O_INACTIVE)' failed."" [Undecided,Confirmed]
14:55:58 mdbooth mriedem: Why are we calling post if the migration failed?
14:56:09 mriedem mdbooth: because libvirt told us the job was complete
14:56:15 mriedem see my notes in https://bugs.launchpad.net/nova/+bug/1706377
14:56:23 mdbooth mriedem: *That's* the bug
14:56:31 mdbooth mriedem: And we already kinda knew about that, right?
14:56:55 mdbooth Didn't I leave a comment in there to that effect?
14:57:01 mriedem in where?
14:57:10 mdbooth libvirt/drive
14:57:11 mdbooth r
14:57:44 mdbooth mriedem: Sorry, libvirt/guest.py
14:57:46 mdbooth is_job_complete
14:58:08 dansmith jaypipes: mriedem: okay I got the resource override stuff working in jaypipes' patch and some unified code between them for doubling/undoubling resources, so now I'm going to look at the peripheral test failures
14:58:14 mdbooth mriedem: It's there in one of my trademark big blocks of comment
14:58:19 dansmith I have about 30 minutes until my next meeting so I will push ahead of that regardless of my progress
14:58:34 mdbooth # Secondly, with the current method we only know that 'no job'
14:58:34 mdbooth # indicates completion. It does not necessarily indicate successful
14:58:34 mdbooth # completion: the job could have failed, or been cancelled. When
14:58:34 mdbooth # polling for block job info we have no way to detect this, so we
14:58:34 mdbooth # assume success.
14:58:52 jaypipes dansmith: k. I'm happy to take the baton on fixing periphery tests when you go to your meeting.
14:58:52 mriedem ah ok,
14:59:00 mriedem that was written around the time of the great swap volume rewrite
14:59:29 dansmith jaypipes: ack
14:59:39 mriedem cfriesen: i don't understand your question
14:59:48 mriedem cfriesen: the migration completes but actually fails on the source,
15:00:16 mriedem but we don't know it fails, we just know the job is 'complete' so we tell dest to do post live migration stuff, and when it does, the guest never made it to dest (or it was deleted by libvirt when it found that the source failed)
15:00:35 mriedem so i'm trying to figure out a way to get the qemu instance logs into the n-cpu logs for debug
15:00:39 mdbooth mriedem: I think we should rewrite that polling block to consume events instead. It's also less buggy.
15:00:45 cfriesen mriedem: I was just thinking that we had all the info needed to get the file, so didn't see what the problem was....but it's not "can we get the file", but "can we determine there was a failure so we know to go get the file"
15:00:50 mdbooth As in, it was designed for this in the first place.
15:01:27 mriedem cfriesen: maybe, i don't know how configurable that path is
15:01:53 mriedem seems pretty hacky though, i'd think you could get qemu guest logs from libvirt apis
15:02:12 mdbooth mriedem: I don't think so, btw.
15:02:44 cfriesen mriedem: ah, right, we don't control all the clouds this runs on. I think it is configurable where those logs go.
15:03:05 mriedem right
15:03:13 mriedem that's why i'd need an api
15:04:10 mdbooth Basically we should switch to using libvirt events api. Extensive documentation here: http://libvirt.org/docs/libvirt-appdev-guide/en-US/html/Application_Development_Guide-Guest_Domains-Event_Not.html
15:04:33 cfriesen with a big TBD on that page?
15:04:35 mriedem mdbooth: yeah i suppose virConnectDomainEventJobCompletedCallback
15:04:50 mdbooth cfriesen: You need more? Pshaw
15:05:52 mdbooth cfriesen: It's a small TBD, anyway. Classier that way.
15:07:19 mdbooth mriedem: I wonder if we could register a libvirt error handler, and dump errors into nova compute logs as a matter of course:http://libvirt.org/docs/libvirt-appdev-guide-python/en-US/html/libvirt_application_development_guide_using_python-Error_Handling-Registering_Error_Handler.html
15:07:31 mdbooth That might achieve what you want in practise.
15:11:14 mriedem where is the error array defined?
15:11:33 mdbooth mriedem: rtfs
15:11:47 cfriesen mriedem: mdbooth: is there a libvirt bug here? I mean the source is running _live_migration_monitor() and calling guest.get_job_info(). shouldn't libvirt detect a failure?
15:11:48 mriedem ha, we already register an error handler
15:11:49 mriedem def _libvirt_error_handler(context, err):
15:11:49 mriedem # Just ignore instead of default outputting to stderr.
15:11:49 mriedem pass
15:12:02 mdbooth mriedem: hehe
15:12:07 cfriesen and if it doesn't, are we going to get an error in the callback?

Earlier   Later