| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-21 | |||
| 16:59:48 | cdent | leakypipes: I think I actually prefer option b for that url. you? | |
| 17:01:40 | leakypipes | cdent: not sure I follow you... GET /resource_providers is intended to return a list of the resource providers that can accomodate the requested resources, and in this case, if one of the requested resources is a shared resource, we naturally need to account for sharing providers (we just don't need to return information that indicates a sharing provider would be used) | |
| 17:04:21 | cdent | What I’m saying is that maybe (now that we have /allocation_candidates) listing /resource_providers should just do that: list resource providers that match queries, where resources is just another type of query. the concept of a shared resource isn’t necessarily a concept that is “natural” when request a listing of resource provider for no specific purpose | |
| 17:04:42 | cdent | (i’m not sold on this, just throwing it out there for thinking) | |
| 17:06:42 | leakypipes | cdent: I'm not actually sure what you're advocating :) | |
| 17:07:04 | cdent | i’m advocating leaving /resource_providers as it is right now | |
| 17:08:12 | cdent | leakypipes: or rather not advocating, wondering if it makes sense to do so | |
| 17:08:33 | leakypipes | cdent: I see. I don't suppose it would hurt to leave it as-is for a little while as we decide on that question. | |
| 17:11:04 | cdent | figleaf: you have any thoughts on the last few lines? | |
| 17:27:13 | cdent | have fun stephenfin | |
| 17:35:11 | mriedem | maybe we should put a dumb retry in the functional osapi client fixture if we get a 401 or 403 response for now? | |
| 17:35:20 | mriedem | as a hack until we're to RC1? | |
| 17:38:18 | melwitt | yeah, I forgot about that idea | |
| 17:41:27 | melwitt | I don't understand what the right way to fix this would be | |
| 17:42:14 | mriedem | rm -rf nova/tests/functional | |
| 17:42:25 | melwitt | hah | |
| 17:42:38 | cdent | is the auth error the main offender at this point? | |
| 17:42:54 | mriedem | it's what i've been noticing after the global keepalive=False change | |
| 17:43:14 | cdent | common url, or multiple urls? | |
| 17:43:15 | mriedem | which is odd since we use the noauth middleware i thought in all of the osapi fixture tests | |
| 17:43:27 | melwitt | yeah. I don't get it | |
| 17:43:35 | mriedem | i'll find the last one i just rechecked | |
| 17:44:03 | mriedem | http://logs.openstack.org/11/485011/4/gate/gate-nova-tox-functional-py35-ubuntu-xenial/42d69de/console.html#_2017-07-21_14_55_31_532907 | |
| 17:45:25 | mriedem | i've been seeing these in the functional jobs too | |
| 17:45:26 | mriedem | sys:1: ResourceWarning: unclosed file <_io.FileIO name=1 mode='wb' closefd=True> | |
| 17:46:41 | cdent | mriedem: is that on both py27 and py35 or or just py35? | |
| 17:47:15 | mriedem | only seeing it on py3 | |
| 17:47:29 | mriedem | py3 jobs are also way chattier about warnings | |
| 17:48:16 | melwitt | gdi looks like logstash.o.o is busted again | |
| 17:48:26 | melwitt | I wanted to check which jobs have "OpenStackApiAuthenticationException: Authentication error" | |
| 17:48:30 | cdent | I get unclosed file errors all over the place in lots of other things besides nova in py35 | |
| 17:48:35 | cdent | s/errors/warnings/ | |
| 17:48:57 | mriedem | so when we get that auth error we aren't providing the original response text | |
| 17:49:10 | mriedem | i had that working in the functional tests but leakypipes removed it | |
| 17:49:24 | mriedem | and then i raged | |
| 17:49:32 | leakypipes | hmm? | |
| 17:49:38 | mriedem | you know what you did | |
| 17:49:42 | mriedem | sec | |
| 17:49:53 | mriedem | https://github.com/openstack/nova/commit/de8096a59d80d10ff1ccf14e0b345be641ba4f07 | |
| 17:50:16 | mriedem | this is actually a bit different | |
| 17:50:23 | mriedem | do we have a bug for this anywhere? | |
| 17:50:45 | cdent | which this? | |
| 17:50:53 | mriedem | the random auth failures in functional tests | |
| 17:51:23 | mriedem | i don't see one | |
| 17:51:34 | melwitt | mriedem: oh, sorry. I think I approved that. I thought bc the jobs were passing it wasn't needed anymore, I didn't know it was a thing to get more info | |
| 17:51:45 | mriedem | yeah it's for debug | |
| 17:52:25 | melwitt | I guess put it back, with a comment that says what it's for | |
| 17:52:36 | mriedem | i assume the "BaseException.message has been deprecated as of Python 2.6" was for something else | |
| 17:52:41 | mriedem | b/c i've seen that before too | |
| 17:52:50 | mriedem | it wouldn't actually help in what we're seeing here | |
| 17:52:54 | mriedem | so i'm going to push something else for that | |
| 17:52:55 | mriedem | it | |
| 17:52:55 | mriedem | push | |
| 17:52:56 | mriedem | real | |
| 17:53:10 | melwitt | I thought that's what was causing those warnings was setting of the message attribute | |
| 17:57:26 | mriedem | i don't remember anymore, i added it here https://github.com/openstack/nova/commit/01dd1a05a213c0cbd0097188418cabe915291c8d | |
| 17:57:33 | mriedem | anywho, not the issue here | |
| 17:58:06 | leakypipes | mriedem: this tempest.api.identity.admin.v3.test_users.UsersV3TestJSON.test_password_history_not_enforced_in_admin_reset failure... grrr. | |
| 17:59:12 | mriedem | i think we have a signature for that one | |
| 17:59:22 | mriedem | http://status.openstack.org/elastic-recheck/#1702211 ? | |
| 17:59:42 | mriedem | cdent: melwitt: https://bugs.launchpad.net/nova/+bug/1705753 | |
| 17:59:43 | openstack | Launchpad bug 1705753 in OpenStack Compute (nova) "Random OpenStackApiAuthenticationException: Authentication error in nova functional tests" [Undecided,New] | |
| 18:01:08 | cdent | thanks, mriedem | |
| 18:01:28 | mriedem | i'll push a debug patch | |
| 18:02:38 | cdent | mriedem: do you want me to try the retry on auth fail thingie? | |
| 18:03:40 | mriedem | sure | |
| 18:03:53 | cdent | it seems like it might of some use but doesn’t really get at whatever the issue is, sadly | |
| 18:03:57 | cdent | but yeah, i’ll make one go | |
| 18:04:00 | melwitt | I wonder if this is relevant https://github.com/openstack/nova/blob/master/nova/tests/fixtures.py#L1436-L1438 | |
| 18:04:32 | mriedem | probably | |
| 18:04:47 | mriedem | the spike in failures started when placement fixture was turned on globally in the IntegratedHelpers mixin | |
| 18:05:35 | melwitt | yeah, that lines up with the fact that "The current placement NoAuthMiddleware returns a 401 in case a token is not provided" | |
| 18:05:40 | melwitt | I just don't know what that means | |
| 18:06:20 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/auth.py#L32 | |
| 18:06:26 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/auth.py#L41 | |
| 18:07:13 | cdent | if that’s playing a part, then it is likely that the problem is when compute requests land on the placement api (which seems to be the core problem here) | |
| 18:07:14 | melwitt | so does that imply that a request is being made that does send the x-auth-token header? is there anything other than GET/PUT/DELETE/POST? | |
| 18:07:22 | melwitt | *does not | |
| 18:07:38 | melwitt | oh. compute requests landing on placement api | |
| 18:07:53 | mriedem | i think there is an eventlet switch that goofs things up | |
| 18:07:55 | mriedem | or that's the theory | |
| 18:08:17 | melwitt | I guess I don't understand that | |
| 18:08:30 | mriedem | this is what compute does https://github.com/openstack/nova/blob/master/nova/api/openstack/auth.py#L32 | |
| 18:09:13 | cdent | placement does what it does to behave like a normal auth middleware and not fake more than it should, it basically stripped that middleware back to the basics | |
| 18:09:24 | cdent | changing it would not fix the real problem here | |
| 18:09:28 | cdent | it would mask it | |
| 18:09:31 | cdent | and we don’t want to do that do we? | |
| 18:09:39 | mriedem | right we do'nt send a fake token for compute requests https://github.com/openstack/nova/blob/master/nova/tests/functional/api/client.py#L142 | |
| 18:10:00 | melwitt | I mean, how does something switch to the wrong api, how does a compute request end up going to the placement api | |
| 18:10:37 | superdan | bad threading | |
| 18:10:39 | cdent | melwitt: the theory is that something is causing eventlet sockets to get confused | |
| 18:10:51 | superdan | yeah | |
| 18:11:23 | melwitt | \:| okay | |
| 18:11:44 | superdan | melwitt: your hair is messed up? | |
| 18:11:59 | melwitt | that's my raised unibrow | |
| 18:12:01 | superdan | eyebrows? | |
| 18:12:02 | superdan | okay | |
| 18:12:03 | superdan | heh | |
| 18:12:09 | cdent | we could run one of the apis (presunably placement) on wsgi intercept instead of a separate server thread, and then it wouldn’t be on threads? | |
| 18:12:31 | cdent | (or rather not in the same way) | |