| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-03 | |||
| 14:57:46 | mdbooth | is_job_complete | |
| 14:58:08 | dansmith | jaypipes: mriedem: okay I got the resource override stuff working in jaypipes' patch and some unified code between them for doubling/undoubling resources, so now I'm going to look at the peripheral test failures | |
| 14:58:14 | mdbooth | mriedem: It's there in one of my trademark big blocks of comment | |
| 14:58:19 | dansmith | I have about 30 minutes until my next meeting so I will push ahead of that regardless of my progress | |
| 14:58:34 | mdbooth | # Secondly, with the current method we only know that 'no job' | |
| 14:58:34 | mdbooth | # indicates completion. It does not necessarily indicate successful | |
| 14:58:34 | mdbooth | # completion: the job could have failed, or been cancelled. When | |
| 14:58:34 | mdbooth | # polling for block job info we have no way to detect this, so we | |
| 14:58:34 | mdbooth | # assume success. | |
| 14:58:52 | jaypipes | dansmith: k. I'm happy to take the baton on fixing periphery tests when you go to your meeting. | |
| 14:58:52 | mriedem | ah ok, | |
| 14:59:00 | mriedem | that was written around the time of the great swap volume rewrite | |
| 14:59:29 | dansmith | jaypipes: ack | |
| 14:59:39 | mriedem | cfriesen: i don't understand your question | |
| 14:59:48 | mriedem | cfriesen: the migration completes but actually fails on the source, | |
| 15:00:16 | mriedem | but we don't know it fails, we just know the job is 'complete' so we tell dest to do post live migration stuff, and when it does, the guest never made it to dest (or it was deleted by libvirt when it found that the source failed) | |
| 15:00:35 | mriedem | so i'm trying to figure out a way to get the qemu instance logs into the n-cpu logs for debug | |
| 15:00:39 | mdbooth | mriedem: I think we should rewrite that polling block to consume events instead. It's also less buggy. | |
| 15:00:45 | cfriesen | mriedem: I was just thinking that we had all the info needed to get the file, so didn't see what the problem was....but it's not "can we get the file", but "can we determine there was a failure so we know to go get the file" | |
| 15:00:50 | mdbooth | As in, it was designed for this in the first place. | |
| 15:01:27 | mriedem | cfriesen: maybe, i don't know how configurable that path is | |
| 15:01:53 | mriedem | seems pretty hacky though, i'd think you could get qemu guest logs from libvirt apis | |
| 15:02:12 | mdbooth | mriedem: I don't think so, btw. | |
| 15:02:44 | cfriesen | mriedem: ah, right, we don't control all the clouds this runs on. I think it is configurable where those logs go. | |
| 15:03:05 | mriedem | right | |
| 15:03:13 | mriedem | that's why i'd need an api | |
| 15:04:10 | mdbooth | Basically we should switch to using libvirt events api. Extensive documentation here: http://libvirt.org/docs/libvirt-appdev-guide/en-US/html/Application_Development_Guide-Guest_Domains-Event_Not.html | |
| 15:04:33 | cfriesen | with a big TBD on that page? | |
| 15:04:35 | mriedem | mdbooth: yeah i suppose virConnectDomainEventJobCompletedCallback | |
| 15:04:50 | mdbooth | cfriesen: You need more? Pshaw | |
| 15:05:52 | mdbooth | cfriesen: It's a small TBD, anyway. Classier that way. | |
| 15:07:19 | mdbooth | mriedem: I wonder if we could register a libvirt error handler, and dump errors into nova compute logs as a matter of course:http://libvirt.org/docs/libvirt-appdev-guide-python/en-US/html/libvirt_application_development_guide_using_python-Error_Handling-Registering_Error_Handler.html | |
| 15:07:31 | mdbooth | That might achieve what you want in practise. | |
| 15:11:14 | mriedem | where is the error array defined? | |
| 15:11:33 | mdbooth | mriedem: rtfs | |
| 15:11:47 | cfriesen | mriedem: mdbooth: is there a libvirt bug here? I mean the source is running _live_migration_monitor() and calling guest.get_job_info(). shouldn't libvirt detect a failure? | |
| 15:11:48 | mriedem | ha, we already register an error handler | |
| 15:11:49 | mriedem | def _libvirt_error_handler(context, err): | |
| 15:11:49 | mriedem | # Just ignore instead of default outputting to stderr. | |
| 15:11:49 | mriedem | pass | |
| 15:12:02 | mdbooth | mriedem: hehe | |
| 15:12:07 | cfriesen | and if it doesn't, are we going to get an error in the callback? | |
| 15:12:20 | mdbooth | cfriesen: No | |
| 15:12:30 | mriedem | "with error being a list of information about the error being raised. " | |
| 15:12:59 | mriedem | i suppose it's similar to a libvirtError | |
| 15:13:39 | mdbooth | cfriesen: I don't recall the detail now, but at the time kashyap and I went over the libvirt and libvirt python binding code very carefully | |
| 15:13:52 | mdbooth | cfriesen: We're extracting everything from it which can be extracted | |
| 15:14:06 | mriedem | ah yup it's just the libvirtError.err list | |
| 15:14:09 | mdbooth | Hence my big comment explaining what we're not getting | |
| 15:15:05 | mdbooth | It's not designed to be used this way | |
| 15:15:38 | mdbooth | The intention was that you'd consume events instead. That api has been stable for much longer. | |
| 15:17:14 | sdague | mriedem: you'd need another grok parser | |
| 15:19:51 | cfriesen | mdbooth: okay, I think I got it. On another note, currently with block live migration if the guest is dirtying disk quickly the current logs don't show information about the initial block transfer. | |
| 15:20:15 | mriedem | would be nice if the libvirtError python binding class just had a nice __repr__ | |
| 15:20:22 | mriedem | maybe it does already... | |
| 15:21:48 | openstackgerrit | Dan Smith proposed openstack/nova master: remove provider allocs in confirm/revert resize https://review.openstack.org/488510 | |
| 15:21:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Sum allocations in the scheduler when resizing to the same host https://review.openstack.org/490085 | |
| 15:21:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Add resource utilities to scheduler utils https://review.openstack.org/490514 | |
| 15:22:13 | dansmith | jaypipes: I gotta start getting ready for my call, so I'm pushing.. the last patch is the only one that needs attention, AFAIK, a few fails in unit tests at least | |
| 15:22:31 | jaypipes | dansmith: got it. will take the ball. | |
| 15:24:08 | dansmith | jaypipes: note the new patch in the middle that adds a couple of utils and generalizes something in mriedem's patch | |
| 15:24:38 | jaypipes | dansmith: noted | |
| 15:25:16 | jaypipes | dansmith: I presume that's not the patch with test failures, though, yes? the top is the one with failures? | |
| 15:25:28 | dansmith | jaypipes: just your last confirm/revert one yeah | |
| 15:25:34 | jaypipes | got it thx | |
| 15:31:09 | mriedem | comments in the utils patch | |
| 15:58:33 | mriedem | cdent: where would one find the current placement api-ref? | |
| 15:58:42 | mriedem | or is that still just in builds that touch it? | |
| 15:59:15 | cdent | mriedem: there’s no publishing job yet as we were waiting for the stack to merge so that we could then move it to the new location and _then_ publish it | |
| 16:01:39 | ildikov | mriedem: hey | |
| 16:01:52 | ildikov | mriedem: coming to the meeting? | |
| 16:03:28 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Start using oslo_config.sphinxext https://review.openstack.org/482961 | |
| 16:03:29 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Rework README to reflect new doc URLs https://review.openstack.org/480074 | |
| 16:03:29 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Start using oslo_policy.sphinxext https://review.openstack.org/479358 | |
| 16:03:30 | openstackgerrit | Stephen Finucane proposed openstack/nova master: policies: Fix Sphinx issues https://review.openstack.org/480516 | |
| 16:03:30 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Rework index page per new sections https://review.openstack.org/478485 | |
| 16:03:31 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Remove dead files https://review.openstack.org/478470 | |
| 16:04:12 | mriedem | cdent: ok i was going to see if we needed any changes to the api-ref for the allocations minItems:1 thing | |
| 16:05:15 | cdent | mriedem: i think this one is allocations: https://review.openstack.org/#/c/470933/ | |
| 16:05:17 | cdent | so not merged yet | |
| 16:08:48 | cdent | mriedem: when you use the term “latent issue” what does that actually mean? | |
| 16:09:35 | mriedem | not introduced in pike | |
| 16:09:37 | mriedem | not a regression | |
| 16:10:02 | cdent | thanks | |
| 16:15:06 | mriedem | cdent: ok questions in https://review.openstack.org/#/c/470933/ | |
| 16:16:18 | cdent | mriedem: cool, I’ll hope andrey can look at those soon. If not, I can, I’m in the api-wg meeting now and then after that am gone (officially) until monday | |
| 16:31:11 | mriedem | dtantsur: jlvillal: did ironic ever go ahead with raising minimum required microversions? | |
| 16:31:51 | dtantsur | mriedem: nope, we haven't got to it | |
| 16:32:04 | dtantsur | I think we have an api-wg guideline for that, resulting from our discussions | |
| 16:33:17 | mriedem | ok | |
| 16:33:24 | mriedem | edleafe: i think the unit test failures in https://review.openstack.org/#/c/487925/ are probably related | |
| 16:33:42 | mriedem | it's super rare to have a random unit test failure with the ironic stuff in nova, and it's suspect when you're touching that driver | |
| 16:35:00 | mriedem | need some cores to look at this https://review.openstack.org/#/c/489763/ - it's a regression in pike, pretty simple fix | |
| 16:35:56 | edleafe | mriedem: those failures do seem to be more or less random. They don't happen locally | |
| 16:37:56 | mriedem | edleafe: i've never seen either of those happen in our ci though | |
| 16:39:13 | mriedem | http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22AssertionError%3A%20Expected%20call%3A%20call('node.list'%2C%20associated%3DTrue%2C%20limit%3D0)%5C%22%20AND%20tags%3A%5C%22console%5C%22&from=7d | |
| 16:39:28 | mriedem | that only shows up in that change | |
| 16:40:22 | openstackgerrit | Merged openstack/nova master: reflow rpc doc to 80 columns https://review.openstack.org/490455 | |
| 16:41:10 | openstackgerrit | Merged openstack/nova master: doc: Make use of definition lists, literals https://review.openstack.org/490481 | |
| 16:44:22 | mriedem | bauzas: any idea why the ServerGroupAntiAffinityFilter would stop working if you aren't tracking instance changes between the compute and scheduler service? Like i know there is a race possibility, but is that filter completely dependent on those? shouldn't it fallback to check the db or something if it's not tracking instance updates? | |