Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-03
14:56:31 mdbooth mriedem: And we already kinda knew about that, right?
14:56:55 mdbooth Didn't I leave a comment in there to that effect?
14:57:01 mriedem in where?
14:57:10 mdbooth libvirt/drive
14:57:11 mdbooth r
14:57:44 mdbooth mriedem: Sorry, libvirt/guest.py
14:57:46 mdbooth is_job_complete
14:58:08 dansmith jaypipes: mriedem: okay I got the resource override stuff working in jaypipes' patch and some unified code between them for doubling/undoubling resources, so now I'm going to look at the peripheral test failures
14:58:14 mdbooth mriedem: It's there in one of my trademark big blocks of comment
14:58:19 dansmith I have about 30 minutes until my next meeting so I will push ahead of that regardless of my progress
14:58:34 mdbooth # assume success.
14:58:34 mdbooth # polling for block job info we have no way to detect this, so we
14:58:34 mdbooth # completion: the job could have failed, or been cancelled. When
14:58:34 mdbooth # indicates completion. It does not necessarily indicate successful
14:58:34 mdbooth # Secondly, with the current method we only know that 'no job'
14:58:52 mriedem ah ok,
14:58:52 jaypipes dansmith: k. I'm happy to take the baton on fixing periphery tests when you go to your meeting.
14:59:00 mriedem that was written around the time of the great swap volume rewrite
14:59:29 dansmith jaypipes: ack
14:59:39 mriedem cfriesen: i don't understand your question
14:59:48 mriedem cfriesen: the migration completes but actually fails on the source,
15:00:16 mriedem but we don't know it fails, we just know the job is 'complete' so we tell dest to do post live migration stuff, and when it does, the guest never made it to dest (or it was deleted by libvirt when it found that the source failed)
15:00:35 mriedem so i'm trying to figure out a way to get the qemu instance logs into the n-cpu logs for debug
15:00:39 mdbooth mriedem: I think we should rewrite that polling block to consume events instead. It's also less buggy.
15:00:45 cfriesen mriedem: I was just thinking that we had all the info needed to get the file, so didn't see what the problem was....but it's not "can we get the file", but "can we determine there was a failure so we know to go get the file"
15:00:50 mdbooth As in, it was designed for this in the first place.
15:01:27 mriedem cfriesen: maybe, i don't know how configurable that path is
15:01:53 mriedem seems pretty hacky though, i'd think you could get qemu guest logs from libvirt apis
15:02:12 mdbooth mriedem: I don't think so, btw.
15:02:44 cfriesen mriedem: ah, right, we don't control all the clouds this runs on. I think it is configurable where those logs go.
15:03:05 mriedem right
15:03:13 mriedem that's why i'd need an api
15:04:10 mdbooth Basically we should switch to using libvirt events api. Extensive documentation here: http://libvirt.org/docs/libvirt-appdev-guide/en-US/html/Application_Development_Guide-Guest_Domains-Event_Not.html
15:04:33 cfriesen with a big TBD on that page?
15:04:35 mriedem mdbooth: yeah i suppose virConnectDomainEventJobCompletedCallback
15:04:50 mdbooth cfriesen: You need more? Pshaw
15:05:52 mdbooth cfriesen: It's a small TBD, anyway. Classier that way.
15:07:19 mdbooth mriedem: I wonder if we could register a libvirt error handler, and dump errors into nova compute logs as a matter of course:http://libvirt.org/docs/libvirt-appdev-guide-python/en-US/html/libvirt_application_development_guide_using_python-Error_Handling-Registering_Error_Handler.html
15:07:31 mdbooth That might achieve what you want in practise.
15:11:14 mriedem where is the error array defined?
15:11:33 mdbooth mriedem: rtfs
15:11:47 cfriesen mriedem: mdbooth: is there a libvirt bug here? I mean the source is running _live_migration_monitor() and calling guest.get_job_info(). shouldn't libvirt detect a failure?
15:11:48 mriedem ha, we already register an error handler
15:11:49 mriedem pass
15:11:49 mriedem # Just ignore instead of default outputting to stderr.
15:11:49 mriedem def _libvirt_error_handler(context, err):
15:12:02 mdbooth mriedem: hehe
15:12:07 cfriesen and if it doesn't, are we going to get an error in the callback?
15:12:20 mdbooth cfriesen: No
15:12:30 mriedem "with error being a list of information about the error being raised. "
15:12:59 mriedem i suppose it's similar to a libvirtError
15:13:39 mdbooth cfriesen: I don't recall the detail now, but at the time kashyap and I went over the libvirt and libvirt python binding code very carefully
15:13:52 mdbooth cfriesen: We're extracting everything from it which can be extracted
15:14:06 mriedem ah yup it's just the libvirtError.err list
15:14:09 mdbooth Hence my big comment explaining what we're not getting
15:15:05 mdbooth It's not designed to be used this way
15:15:38 mdbooth The intention was that you'd consume events instead. That api has been stable for much longer.
15:17:14 sdague mriedem: you'd need another grok parser
15:19:51 cfriesen mdbooth: okay, I think I got it. On another note, currently with block live migration if the guest is dirtying disk quickly the current logs don't show information about the initial block transfer.
15:20:15 mriedem would be nice if the libvirtError python binding class just had a nice __repr__
15:20:22 mriedem maybe it does already...
15:21:48 openstackgerrit Dan Smith proposed openstack/nova master: remove provider allocs in confirm/revert resize https://review.openstack.org/488510
15:21:49 openstackgerrit Dan Smith proposed openstack/nova master: Add resource utilities to scheduler utils https://review.openstack.org/490514
15:21:49 openstackgerrit Dan Smith proposed openstack/nova master: Sum allocations in the scheduler when resizing to the same host https://review.openstack.org/490085
15:22:13 dansmith jaypipes: I gotta start getting ready for my call, so I'm pushing.. the last patch is the only one that needs attention, AFAIK, a few fails in unit tests at least
15:22:31 jaypipes dansmith: got it. will take the ball.
15:24:08 dansmith jaypipes: note the new patch in the middle that adds a couple of utils and generalizes something in mriedem's patch
15:24:38 jaypipes dansmith: noted
15:25:16 jaypipes dansmith: I presume that's not the patch with test failures, though, yes? the top is the one with failures?
15:25:28 dansmith jaypipes: just your last confirm/revert one yeah
15:25:34 jaypipes got it thx
15:31:09 mriedem comments in the utils patch
15:58:33 mriedem cdent: where would one find the current placement api-ref?
15:58:42 mriedem or is that still just in builds that touch it?
15:59:15 cdent mriedem: there’s no publishing job yet as we were waiting for the stack to merge so that we could then move it to the new location and _then_ publish it
16:01:39 ildikov mriedem: hey
16:01:52 ildikov mriedem: coming to the meeting?
16:03:28 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Start using oslo_config.sphinxext https://review.openstack.org/482961
16:03:29 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Start using oslo_policy.sphinxext https://review.openstack.org/479358
16:03:29 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Rework README to reflect new doc URLs https://review.openstack.org/480074
16:03:30 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Rework index page per new sections https://review.openstack.org/478485
16:03:30 openstackgerrit Stephen Finucane proposed openstack/nova master: policies: Fix Sphinx issues https://review.openstack.org/480516
16:03:31 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Remove dead files https://review.openstack.org/478470
16:04:12 mriedem cdent: ok i was going to see if we needed any changes to the api-ref for the allocations minItems:1 thing
16:05:15 cdent mriedem: i think this one is allocations: https://review.openstack.org/#/c/470933/
16:05:17 cdent so not merged yet
16:08:48 cdent mriedem: when you use the term “latent issue” what does that actually mean?
16:09:35 mriedem not introduced in pike
16:09:37 mriedem not a regression
16:10:02 cdent thanks
16:15:06 mriedem cdent: ok questions in https://review.openstack.org/#/c/470933/
16:16:18 cdent mriedem: cool, I’ll hope andrey can look at those soon. If not, I can, I’m in the api-wg meeting now and then after that am gone (officially) until monday
16:31:11 mriedem dtantsur: jlvillal: did ironic ever go ahead with raising minimum required microversions?
16:31:51 dtantsur mriedem: nope, we haven't got to it
16:32:04 dtantsur I think we have an api-wg guideline for that, resulting from our discussions
16:33:17 mriedem ok
16:33:24 mriedem edleafe: i think the unit test failures in https://review.openstack.org/#/c/487925/ are probably related
16:33:42 mriedem it's super rare to have a random unit test failure with the ironic stuff in nova, and it's suspect when you're touching that driver
16:35:00 mriedem need some cores to look at this https://review.openstack.org/#/c/489763/ - it's a regression in pike, pretty simple fix
16:35:56 edleafe mriedem: those failures do seem to be more or less random. They don't happen locally

Earlier   Later