| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-03 | |||
| 14:47:27 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954 | |
| 14:47:34 | mriedem | well, alternatively the fix is a microversion to expose some specific part of system metadata and what that entails | |
| 14:47:51 | mriedem | sdague: ever thought about indexing qemu instance logs in our ci runs? | |
| 14:48:05 | mriedem | when live migration jobs fail, a lot of the time it's due to | |
| 14:48:05 | mriedem | http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz | |
| 14:48:11 | mriedem | /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed | |
| 14:48:17 | mriedem | but ^ isn't exposed in anything we index | |
| 14:48:25 | jaypipes | dansmith: did you see cdent just pushed a revision on the confirm resize patch? | |
| 14:48:59 | dansmith | jaypipes: a bit ago while we were talking, yeah. he said so and that's when I pulled to start working | |
| 14:49:14 | jaypipes | gotcha. just making sure you noticed. carry on. | |
| 14:49:48 | mriedem | what i'd really love is if the libvirt / qemu job had some way to get those details from the guest | |
| 14:50:08 | mriedem | kashyap: mdbooth: you know how during a live migration we're checking the domain job status to see when it completes, or if it fails? | |
| 14:50:19 | mdbooth | mriedem: Yep | |
| 14:50:23 | mriedem | is there any way to get the qemu guest logs when that fails, like http://logs.openstack.org/10/490110/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/6c1da1c/logs/subnode-2/libvirt/qemu/instance-00000003.txt.gz | |
| 14:50:38 | mriedem | i really want: /build/qemu-orucB6/qemu-2.8+dfsg/nbd/server.c:nbd_co_receive_request():L1135: reading from socket failed | |
| 14:51:13 | mriedem | when ^ happens, the only failure we get in the n-cpu logs is that on the destination when we're doing post-live migration at destination, the instance (guest domain) isn't found | |
| 14:51:18 | mriedem | because it blew up on the source side | |
| 14:52:21 | mdbooth | mriedem: Is ^^^ from dest? | |
| 14:53:52 | mriedem | no that's source | |
| 14:53:58 | mriedem | here is another one http://logs.openstack.org/66/483566/10/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/0437fbe/logs/subnode-2/libvirt/qemu/instance-00000011.txt.gz | |
| 14:54:07 | mriedem | different error, but results in the same kind of thing in the dest n-cpu logs | |
| 14:54:13 | mriedem | InstanceNotFound during post live migration at destination | |
| 14:54:15 | mriedem | b/c it failed on the source | |
| 14:54:27 | cfriesen | mriedem: what's the complication with getting that file from the dest? | |
| 14:55:56 | mriedem | https://bugs.launchpad.net/nova/+bug/1706377 | |
| 14:55:56 | openstack | Launchpad bug 1706377 in OpenStack Compute (nova) "(libvirt) live migration fails on source host due to "Assertion `!(bs->open_flags & BDRV_O_INACTIVE)' failed."" [Undecided,Confirmed] | |
| 14:55:58 | mdbooth | mriedem: Why are we calling post if the migration failed? | |
| 14:56:09 | mriedem | mdbooth: because libvirt told us the job was complete | |
| 14:56:15 | mriedem | see my notes in https://bugs.launchpad.net/nova/+bug/1706377 | |
| 14:56:23 | mdbooth | mriedem: *That's* the bug | |
| 14:56:31 | mdbooth | mriedem: And we already kinda knew about that, right? | |
| 14:56:55 | mdbooth | Didn't I leave a comment in there to that effect? | |
| 14:57:01 | mriedem | in where? | |
| 14:57:10 | mdbooth | libvirt/drive | |
| 14:57:11 | mdbooth | r | |
| 14:57:44 | mdbooth | mriedem: Sorry, libvirt/guest.py | |
| 14:57:46 | mdbooth | is_job_complete | |
| 14:58:08 | dansmith | jaypipes: mriedem: okay I got the resource override stuff working in jaypipes' patch and some unified code between them for doubling/undoubling resources, so now I'm going to look at the peripheral test failures | |
| 14:58:14 | mdbooth | mriedem: It's there in one of my trademark big blocks of comment | |
| 14:58:19 | dansmith | I have about 30 minutes until my next meeting so I will push ahead of that regardless of my progress | |
| 14:58:34 | mdbooth | # Secondly, with the current method we only know that 'no job' | |
| 14:58:34 | mdbooth | # indicates completion. It does not necessarily indicate successful | |
| 14:58:34 | mdbooth | # completion: the job could have failed, or been cancelled. When | |
| 14:58:34 | mdbooth | # polling for block job info we have no way to detect this, so we | |
| 14:58:34 | mdbooth | # assume success. | |
| 14:58:52 | jaypipes | dansmith: k. I'm happy to take the baton on fixing periphery tests when you go to your meeting. | |
| 14:58:52 | mriedem | ah ok, | |
| 14:59:00 | mriedem | that was written around the time of the great swap volume rewrite | |
| 14:59:29 | dansmith | jaypipes: ack | |
| 14:59:39 | mriedem | cfriesen: i don't understand your question | |
| 14:59:48 | mriedem | cfriesen: the migration completes but actually fails on the source, | |
| 15:00:16 | mriedem | but we don't know it fails, we just know the job is 'complete' so we tell dest to do post live migration stuff, and when it does, the guest never made it to dest (or it was deleted by libvirt when it found that the source failed) | |
| 15:00:35 | mriedem | so i'm trying to figure out a way to get the qemu instance logs into the n-cpu logs for debug | |
| 15:00:39 | mdbooth | mriedem: I think we should rewrite that polling block to consume events instead. It's also less buggy. | |
| 15:00:45 | cfriesen | mriedem: I was just thinking that we had all the info needed to get the file, so didn't see what the problem was....but it's not "can we get the file", but "can we determine there was a failure so we know to go get the file" | |
| 15:00:50 | mdbooth | As in, it was designed for this in the first place. | |
| 15:01:27 | mriedem | cfriesen: maybe, i don't know how configurable that path is | |
| 15:01:53 | mriedem | seems pretty hacky though, i'd think you could get qemu guest logs from libvirt apis | |
| 15:02:12 | mdbooth | mriedem: I don't think so, btw. | |
| 15:02:44 | cfriesen | mriedem: ah, right, we don't control all the clouds this runs on. I think it is configurable where those logs go. | |
| 15:03:05 | mriedem | right | |
| 15:03:13 | mriedem | that's why i'd need an api | |
| 15:04:10 | mdbooth | Basically we should switch to using libvirt events api. Extensive documentation here: http://libvirt.org/docs/libvirt-appdev-guide/en-US/html/Application_Development_Guide-Guest_Domains-Event_Not.html | |
| 15:04:33 | cfriesen | with a big TBD on that page? | |
| 15:04:35 | mriedem | mdbooth: yeah i suppose virConnectDomainEventJobCompletedCallback | |
| 15:04:50 | mdbooth | cfriesen: You need more? Pshaw | |
| 15:05:52 | mdbooth | cfriesen: It's a small TBD, anyway. Classier that way. | |
| 15:07:19 | mdbooth | mriedem: I wonder if we could register a libvirt error handler, and dump errors into nova compute logs as a matter of course:http://libvirt.org/docs/libvirt-appdev-guide-python/en-US/html/libvirt_application_development_guide_using_python-Error_Handling-Registering_Error_Handler.html | |
| 15:07:31 | mdbooth | That might achieve what you want in practise. | |
| 15:11:14 | mriedem | where is the error array defined? | |
| 15:11:33 | mdbooth | mriedem: rtfs | |
| 15:11:47 | cfriesen | mriedem: mdbooth: is there a libvirt bug here? I mean the source is running _live_migration_monitor() and calling guest.get_job_info(). shouldn't libvirt detect a failure? | |
| 15:11:48 | mriedem | ha, we already register an error handler | |
| 15:11:49 | mriedem | def _libvirt_error_handler(context, err): | |
| 15:11:49 | mriedem | # Just ignore instead of default outputting to stderr. | |
| 15:11:49 | mriedem | pass | |
| 15:12:02 | mdbooth | mriedem: hehe | |
| 15:12:07 | cfriesen | and if it doesn't, are we going to get an error in the callback? | |
| 15:12:20 | mdbooth | cfriesen: No | |
| 15:12:30 | mriedem | "with error being a list of information about the error being raised. " | |
| 15:12:59 | mriedem | i suppose it's similar to a libvirtError | |
| 15:13:39 | mdbooth | cfriesen: I don't recall the detail now, but at the time kashyap and I went over the libvirt and libvirt python binding code very carefully | |
| 15:13:52 | mdbooth | cfriesen: We're extracting everything from it which can be extracted | |
| 15:14:06 | mriedem | ah yup it's just the libvirtError.err list | |
| 15:14:09 | mdbooth | Hence my big comment explaining what we're not getting | |
| 15:15:05 | mdbooth | It's not designed to be used this way | |
| 15:15:38 | mdbooth | The intention was that you'd consume events instead. That api has been stable for much longer. | |
| 15:17:14 | sdague | mriedem: you'd need another grok parser | |
| 15:19:51 | cfriesen | mdbooth: okay, I think I got it. On another note, currently with block live migration if the guest is dirtying disk quickly the current logs don't show information about the initial block transfer. | |
| 15:20:15 | mriedem | would be nice if the libvirtError python binding class just had a nice __repr__ | |
| 15:20:22 | mriedem | maybe it does already... | |
| 15:21:48 | openstackgerrit | Dan Smith proposed openstack/nova master: remove provider allocs in confirm/revert resize https://review.openstack.org/488510 | |
| 15:21:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Sum allocations in the scheduler when resizing to the same host https://review.openstack.org/490085 | |
| 15:21:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Add resource utilities to scheduler utils https://review.openstack.org/490514 | |
| 15:22:13 | dansmith | jaypipes: I gotta start getting ready for my call, so I'm pushing.. the last patch is the only one that needs attention, AFAIK, a few fails in unit tests at least | |
| 15:22:31 | jaypipes | dansmith: got it. will take the ball. | |
| 15:24:08 | dansmith | jaypipes: note the new patch in the middle that adds a couple of utils and generalizes something in mriedem's patch | |
| 15:24:38 | jaypipes | dansmith: noted | |
| 15:25:16 | jaypipes | dansmith: I presume that's not the patch with test failures, though, yes? the top is the one with failures? | |
| 15:25:28 | dansmith | jaypipes: just your last confirm/revert one yeah | |