| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-21 | |||
| 17:01:40 | leakypipes | cdent: not sure I follow you... GET /resource_providers is intended to return a list of the resource providers that can accomodate the requested resources, and in this case, if one of the requested resources is a shared resource, we naturally need to account for sharing providers (we just don't need to return information that indicates a sharing provider would be used) | |
| 17:04:21 | cdent | What I’m saying is that maybe (now that we have /allocation_candidates) listing /resource_providers should just do that: list resource providers that match queries, where resources is just another type of query. the concept of a shared resource isn’t necessarily a concept that is “natural” when request a listing of resource provider for no specific purpose | |
| 17:04:42 | cdent | (i’m not sold on this, just throwing it out there for thinking) | |
| 17:06:42 | leakypipes | cdent: I'm not actually sure what you're advocating :) | |
| 17:07:04 | cdent | i’m advocating leaving /resource_providers as it is right now | |
| 17:08:12 | cdent | leakypipes: or rather not advocating, wondering if it makes sense to do so | |
| 17:08:33 | leakypipes | cdent: I see. I don't suppose it would hurt to leave it as-is for a little while as we decide on that question. | |
| 17:11:04 | cdent | figleaf: you have any thoughts on the last few lines? | |
| 17:27:13 | cdent | have fun stephenfin | |
| 17:35:11 | mriedem | maybe we should put a dumb retry in the functional osapi client fixture if we get a 401 or 403 response for now? | |
| 17:35:20 | mriedem | as a hack until we're to RC1? | |
| 17:38:18 | melwitt | yeah, I forgot about that idea | |
| 17:41:27 | melwitt | I don't understand what the right way to fix this would be | |
| 17:42:14 | mriedem | rm -rf nova/tests/functional | |
| 17:42:25 | melwitt | hah | |
| 17:42:38 | cdent | is the auth error the main offender at this point? | |
| 17:42:54 | mriedem | it's what i've been noticing after the global keepalive=False change | |
| 17:43:14 | cdent | common url, or multiple urls? | |
| 17:43:15 | mriedem | which is odd since we use the noauth middleware i thought in all of the osapi fixture tests | |
| 17:43:27 | melwitt | yeah. I don't get it | |
| 17:43:35 | mriedem | i'll find the last one i just rechecked | |
| 17:44:03 | mriedem | http://logs.openstack.org/11/485011/4/gate/gate-nova-tox-functional-py35-ubuntu-xenial/42d69de/console.html#_2017-07-21_14_55_31_532907 | |
| 17:45:25 | mriedem | i've been seeing these in the functional jobs too | |
| 17:45:26 | mriedem | sys:1: ResourceWarning: unclosed file <_io.FileIO name=1 mode='wb' closefd=True> | |
| 17:46:41 | cdent | mriedem: is that on both py27 and py35 or or just py35? | |
| 17:47:15 | mriedem | only seeing it on py3 | |
| 17:47:29 | mriedem | py3 jobs are also way chattier about warnings | |
| 17:48:16 | melwitt | gdi looks like logstash.o.o is busted again | |
| 17:48:26 | melwitt | I wanted to check which jobs have "OpenStackApiAuthenticationException: Authentication error" | |
| 17:48:30 | cdent | I get unclosed file errors all over the place in lots of other things besides nova in py35 | |
| 17:48:35 | cdent | s/errors/warnings/ | |
| 17:48:57 | mriedem | so when we get that auth error we aren't providing the original response text | |
| 17:49:10 | mriedem | i had that working in the functional tests but leakypipes removed it | |
| 17:49:24 | mriedem | and then i raged | |
| 17:49:32 | leakypipes | hmm? | |
| 17:49:38 | mriedem | you know what you did | |
| 17:49:42 | mriedem | sec | |
| 17:49:53 | mriedem | https://github.com/openstack/nova/commit/de8096a59d80d10ff1ccf14e0b345be641ba4f07 | |
| 17:50:16 | mriedem | this is actually a bit different | |
| 17:50:23 | mriedem | do we have a bug for this anywhere? | |
| 17:50:45 | cdent | which this? | |
| 17:50:53 | mriedem | the random auth failures in functional tests | |
| 17:51:23 | mriedem | i don't see one | |
| 17:51:34 | melwitt | mriedem: oh, sorry. I think I approved that. I thought bc the jobs were passing it wasn't needed anymore, I didn't know it was a thing to get more info | |
| 17:51:45 | mriedem | yeah it's for debug | |
| 17:52:25 | melwitt | I guess put it back, with a comment that says what it's for | |
| 17:52:36 | mriedem | i assume the "BaseException.message has been deprecated as of Python 2.6" was for something else | |
| 17:52:41 | mriedem | b/c i've seen that before too | |
| 17:52:50 | mriedem | it wouldn't actually help in what we're seeing here | |
| 17:52:54 | mriedem | so i'm going to push something else for that | |
| 17:52:55 | mriedem | it | |
| 17:52:55 | mriedem | push | |
| 17:52:56 | mriedem | real | |
| 17:53:10 | melwitt | I thought that's what was causing those warnings was setting of the message attribute | |
| 17:57:26 | mriedem | i don't remember anymore, i added it here https://github.com/openstack/nova/commit/01dd1a05a213c0cbd0097188418cabe915291c8d | |
| 17:57:33 | mriedem | anywho, not the issue here | |
| 17:58:06 | leakypipes | mriedem: this tempest.api.identity.admin.v3.test_users.UsersV3TestJSON.test_password_history_not_enforced_in_admin_reset failure... grrr. | |
| 17:59:12 | mriedem | i think we have a signature for that one | |
| 17:59:22 | mriedem | http://status.openstack.org/elastic-recheck/#1702211 ? | |
| 17:59:42 | mriedem | cdent: melwitt: https://bugs.launchpad.net/nova/+bug/1705753 | |
| 17:59:43 | openstack | Launchpad bug 1705753 in OpenStack Compute (nova) "Random OpenStackApiAuthenticationException: Authentication error in nova functional tests" [Undecided,New] | |
| 18:01:08 | cdent | thanks, mriedem | |
| 18:01:28 | mriedem | i'll push a debug patch | |
| 18:02:38 | cdent | mriedem: do you want me to try the retry on auth fail thingie? | |
| 18:03:40 | mriedem | sure | |
| 18:03:53 | cdent | it seems like it might of some use but doesn’t really get at whatever the issue is, sadly | |
| 18:03:57 | cdent | but yeah, i’ll make one go | |
| 18:04:00 | melwitt | I wonder if this is relevant https://github.com/openstack/nova/blob/master/nova/tests/fixtures.py#L1436-L1438 | |
| 18:04:32 | mriedem | probably | |
| 18:04:47 | mriedem | the spike in failures started when placement fixture was turned on globally in the IntegratedHelpers mixin | |
| 18:05:35 | melwitt | yeah, that lines up with the fact that "The current placement NoAuthMiddleware returns a 401 in case a token is not provided" | |
| 18:05:40 | melwitt | I just don't know what that means | |
| 18:06:20 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/auth.py#L32 | |
| 18:06:26 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/auth.py#L41 | |
| 18:07:13 | cdent | if that’s playing a part, then it is likely that the problem is when compute requests land on the placement api (which seems to be the core problem here) | |
| 18:07:14 | melwitt | so does that imply that a request is being made that does send the x-auth-token header? is there anything other than GET/PUT/DELETE/POST? | |
| 18:07:22 | melwitt | *does not | |
| 18:07:38 | melwitt | oh. compute requests landing on placement api | |
| 18:07:53 | mriedem | i think there is an eventlet switch that goofs things up | |
| 18:07:55 | mriedem | or that's the theory | |
| 18:08:17 | melwitt | I guess I don't understand that | |
| 18:08:30 | mriedem | this is what compute does https://github.com/openstack/nova/blob/master/nova/api/openstack/auth.py#L32 | |
| 18:09:13 | cdent | placement does what it does to behave like a normal auth middleware and not fake more than it should, it basically stripped that middleware back to the basics | |
| 18:09:24 | cdent | changing it would not fix the real problem here | |
| 18:09:28 | cdent | it would mask it | |
| 18:09:31 | cdent | and we don’t want to do that do we? | |
| 18:09:39 | mriedem | right we do'nt send a fake token for compute requests https://github.com/openstack/nova/blob/master/nova/tests/functional/api/client.py#L142 | |
| 18:10:00 | melwitt | I mean, how does something switch to the wrong api, how does a compute request end up going to the placement api | |
| 18:10:37 | superdan | bad threading | |
| 18:10:39 | cdent | melwitt: the theory is that something is causing eventlet sockets to get confused | |
| 18:10:51 | superdan | yeah | |
| 18:11:23 | melwitt | \:| okay | |
| 18:11:44 | superdan | melwitt: your hair is messed up? | |
| 18:11:59 | melwitt | that's my raised unibrow | |
| 18:12:01 | superdan | eyebrows? | |
| 18:12:02 | superdan | okay | |
| 18:12:03 | superdan | heh | |
| 18:12:09 | cdent | we could run one of the apis (presunably placement) on wsgi intercept instead of a separate server thread, and then it wouldn’t be on threads? | |
| 18:12:31 | cdent | (or rather not in the same way) | |
| 18:16:31 | melwitt | so that means we have two wsgi services total now? it seems like this intercept thing would allow us to set up each one separately (with intercept) right? (because of different host/port combos) | |