| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 13:59:57 | mriedem | avolkov: you might like to take a crack at this https://bugs.launchpad.net/nova/+bug/1719460 | |
| 13:59:58 | openstack | Launchpad bug 1719460 in OpenStack Compute (nova) "(perf) Unnecessarily joining instance.services when listing instances regardless of microversion" [Medium,Triaged] | |
| 13:59:59 | mriedem | should be pretty simple | |
| 14:03:06 | dansmith | mriedem: any outcome from your testing yesterday? | |
| 14:03:43 | mriedem | dansmith: i've got the clean slate, just getting setup to start the 2nd scenario with the 500 ACTIVE and 500 ERROR instances | |
| 14:03:49 | mriedem | for the cell0 and cell1 listing | |
| 14:03:55 | dansmith | okay | |
| 14:04:11 | mriedem | going to need to do that flavor thing because otherwise it's a 60 second rpc timeout per call to select_destinations | |
| 14:04:27 | dansmith | yeah | |
| 14:06:53 | gibi | mriedem, sdague: this also means that the problems we see with the rpc tests in bug 1685333 is not beacuase of the lack of locking | |
| 14:06:54 | sdague | mriedem: is there a reset on rpc variables that is needed that's not happening? | |
| 14:06:54 | openstack | bug 1685333 in OpenStack Compute (nova) "Fatal Python error: Cannot recover from stack overflow. - in py35 unit test job" [High,Confirmed] https://launchpad.net/bugs/1685333 | |
| 14:07:19 | mriedem | sdague: the TestRPC class does a reset per test method | |
| 14:07:39 | sdague | gibi: I don't see how it could be. It might be a deadlock | |
| 14:07:48 | mriedem | dansmith: also came across this last night https://bugs.launchpad.net/nova/+bug/1719487 | |
| 14:07:50 | openstack | Launchpad bug 1719487 in OpenStack Compute (nova) "nova-manage db archive_deleted_rows is not multi-cell aware" [Wishlist,Triaged] - Assigned to Zhenyu Zheng (zhengzhenyu) | |
| 14:07:57 | sdague | the biggest issue though is it doesn't have the timeout bits in place, so it's hard to see what's going on | |
| 14:08:10 | sdague | I think if we trigger the timeout we get a stack trace | |
| 14:08:18 | gibi | sdague: that would be nice | |
| 14:08:22 | dansmith | mriedem: meh | |
| 14:08:50 | mriedem | meh?! | |
| 14:08:55 | mriedem | it's wishlist, sure | |
| 14:09:02 | dansmith | MEH | |
| 14:09:05 | mriedem | gdi | |
| 14:09:13 | gibi | sdague, mriedem: we are at the start of the cycle so I'm brave enough to try to remove the whole locking code and see what happens | |
| 14:09:48 | mriedem | dansmith: oh yeah, also came across this last night https://review.openstack.org/#/c/502236/ | |
| 14:09:49 | mriedem | derp | |
| 14:11:10 | dansmith | ack yeah | |
| 14:11:48 | mriedem | gibi: you can be brave locally to start :) | |
| 14:11:53 | sdague | gibi: yeh, well we should at least get the test_rpc under timeout control, regardless of the rest of it | |
| 14:13:10 | gibi | mriedem: I can definitly do that | |
| 14:13:20 | avolkov | mriedem: ack | |
| 14:13:56 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Make TestRPC inherit from the base nova TestCase https://review.openstack.org/507239 | |
| 14:14:11 | mriedem | ^ removes the lock thing | |
| 14:14:15 | mriedem | so we'll have the timeout fixture | |
| 14:15:29 | gibi | sdague: agree. I found one more place where we use the testtools.TestCase directly. I left a comment in your review | |
| 14:15:46 | sdague | gibi: in the fixtures? | |
| 14:16:00 | gibi | sdague: here https://github.com/openstack/nova/blob/62c4535a85f7d37f1c9da1e8a747f25ec63dc785/nova/tests/unit/api/openstack/test_requestlog.py#L38 | |
| 14:16:18 | sdague | ah, cool, good catch | |
| 14:16:39 | mriedem | i thought ^ was intentional | |
| 14:16:45 | mriedem | for the placement split or something | |
| 14:17:06 | gibi | sdague: I think fixtures are OK to derive from testtools.TestCase as we use fixtures like mixins | |
| 14:17:23 | sdague | gibi: yeh, some of the more advanced ones should see the timeout | |
| 14:17:23 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Fix --max-count handling for nova-manage cell_v2 map_instances https://review.openstack.org/507552 | |
| 14:17:29 | sdague | but I think that's follow on | |
| 14:17:33 | sdague | mriedem: it's a good question | |
| 14:18:30 | gibi | sdague, mriedem: at least this request_log test should be also under timeout control | |
| 14:18:36 | sdague | gibi: so, I'd actually rather handle nova/tests/unit/api/openstack/test_requestlog.py as follow on, because those actually do most of the fixture setup (except the timeout one) manually | |
| 14:18:44 | sdague | so it's going to be a bit more extensive change there | |
| 14:18:52 | sdague | I do agree that we should get that under timeout control | |
| 14:19:02 | sdague | but test_rpc is failing a lot now | |
| 14:19:16 | gibi | sdague: I'm OK with that approach. Then I'm +2 on your patch introducing BasicTestCase | |
| 14:19:17 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/ocata: Fix --max-count handling for nova-manage cell_v2 map_instances https://review.openstack.org/507556 | |
| 14:20:00 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/newton: Fix --max-count handling for nova-manage cell_v2 map_instances https://review.openstack.org/507557 | |
| 14:20:57 | manasm | bauzas: here is the exception I saw with the resize -2017-09-22 07:49:16.377 13573 ERROR nova.api.openstack.extensions File "/usr/lib/python2.7/site-packages/nova/scheduler/utils.py", line 567, in setup_instance_group | |
| 14:20:58 | manasm | 2017-09-22 07:49:16.377 13573 ERROR nova.api.openstack.extensions request_spec.instance_group.hosts = list(group_info.hosts) | |
| 14:20:59 | manasm | 2017-09-22 07:49:16.377 13573 ERROR nova.api.openstack.extensions | |
| 14:21:01 | manasm | 2017-09-22 07:49:16.377 13573 ERROR nova.api.openstack.extensions AttributeError: 'NoneType' object has no attribute 'hosts' | |
| 14:21:03 | manasm | 2017-09-22 07:49:16.377 13573 ERROR nova.api.openstack.extensions | |
| 14:21:07 | jaypipes | mriedem, dansmith, gibi, sdague, bauzas: any of you noticed weird glitches in the new Gerrit web UI where the screen blinks and flashes when you open up long in-page comments? | |
| 14:21:18 | dansmith | no | |
| 14:21:23 | jaypipes | hmmm | |
| 14:21:32 | gibi | at least not yet | |
| 14:21:37 | jaypipes | it's a good thing I don't have Tourettes. | |
| 14:21:57 | jaypipes | or epilepsy I gues | |
| 14:22:49 | jaypipes | efried: around? want to chat about "trait inheritance"... | |
| 14:23:01 | efried | jaypipes I thought you'd never ask :* | |
| 14:23:06 | jaypipes | lol | |
| 14:23:18 | efried | jaypipes I have also experienced the gerrit UI glitchiness. | |
| 14:23:31 | jaypipes | efried: oh, good (or bad...) at least I'm not the only one | |
| 14:23:46 | efried | So yeah, trait inheritance... | |
| 14:24:10 | efried | Did you see my long-winded comment with example based on (or at least attributed to) your response to my response etc. etc.? | |
| 14:24:26 | jaypipes | efried: yeah, so it's absolutely correct that whatever is constructing the provider tree will need to attach traits at the appropriate provider leel | |
| 14:24:27 | jaypipes | level | |
| 14:24:55 | efried | Yuh. And the spec (ultimately the docs) will need to dictate what level(s) is/are "appropriate". | |
| 14:25:12 | efried | Because the code is gonna hafta do some work to percolate 'em around, if that's supported. | |
| 14:25:56 | jaypipes | efried: no, there's no percolating around... | |
| 14:26:27 | sdague | jaypipes: url? | |
| 14:26:42 | efried | jaypipes sdague Talking about this 'un: https://review.openstack.org/#/c/497713/6/specs/queens/approved/add-trait-support-in-allocation-candidates.rst@42 | |
| 14:27:00 | jaypipes | sdague: are you talking about the gerrit thing or the nested providers thing? :) | |
| 14:27:06 | efried | (oh, sdague unless you were... yeah...) | |
| 14:27:09 | sdague | jaypipes: gerrit thing | |
| 14:27:34 | jaypipes | sdague: mostly seen it happen on specs with long (>8 replies) inline comment "threads" | |
| 14:27:44 | jaypipes | sdague: next time it happens I'll ping you a link | |
| 14:27:49 | efried | For me, the gerrit thing is intermittent, happens when I'm expanding comments on a long page with lots of comments | |
| 14:27:53 | jaypipes | ya | |
| 14:28:06 | sdague | gerrit sends back a lot of ajax calls to get all those bits | |
| 14:28:16 | efried | But not reproducible, cause I pop up to the review and back down and do the same thing and it doesn't happen the second time. | |
| 14:28:24 | sdague | if it's gone slow, or your connection is weird, it might take a while for them to pile in and render | |
| 14:28:43 | efried | I don't think it's ajax. Seems like client-side js focus() calls. | |
| 14:28:47 | jaypipes | sdague: nah, it's more like a loop in the UI that happens. | |
| 14:28:55 | jaypipes | sdague: ya, what efried said :) | |
| 14:29:06 | sdague | jaypipes: well, web console in chrome might help explain things | |
| 14:29:22 | jaypipes | like it can't decide which comment to align to the top of the screen canvas | |
| 14:29:35 | jaypipes | sdague: when it happens again I'll ping ya | |
| 14:29:41 | efried | I noticed focus bugs before the upgrade too, usually when composing a comment on a long page, it would jump around (shoving my comment box off the visible screen) | |
| 14:29:56 | jaypipes | efried: yeah, that's happened for a long time | |
| 14:30:46 | sdague | note, we also inject a lot of our own custom client side js to do the CI rollup, so it's entirely possible that is related to the issue | |
| 14:31:31 | sdague | regardless seeing if you can get an inspect console on the issue would be handy | |
| 14:32:20 | jaypipes | sdague: will do | |