| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-10 | |||
| 16:43:49 | cdent | efried: that’s why I asked about what’s doing your deployment (devstack, whatever) | |
| 16:44:50 | efried | oh, sorry, I misunderstood the question then. Yeah, I did a devstack a couple weeks ago, been dinking around with the nova code since then. But hadn't yet had occasion to get this far (been just working conf stuff, so stopping early). | |
| 16:45:23 | cdent | a devstack from two weeks ago that is still mod_wsgi surprises me | |
| 16:45:34 | cdent | so it may be that your error is in journalctl | |
| 16:46:06 | cdent | journalctl --unit devstack@placement-api | |
| 16:48:40 | efried | cdent That journalctl unit doesn't exist. | |
| 16:49:11 | cdent | I’ll leave you in sdague’s good hands then. You seem to have some weird. :) | |
| 16:49:16 | efried | cdent The placement log is in /var/log/apache2 | |
| 16:49:27 | efried | cdent But it doesn't have anything in it. Leading me to believe we aren't getting that far. | |
| 16:49:45 | cdent | oh yeah, I forgot that you have that file | |
| 16:50:12 | cdent | look in all the other files in /var/log/apache2 for whatever is the central error file | |
| 16:50:14 | sdague | efried: yeh, it's super weird, it looks like placement isn't running | |
| 16:50:46 | efried | sdague I'm not real savvy here, but I thought placement was (as of yet) part of n-cpu. | |
| 16:51:05 | sdague | efried: it's a wsgi script | |
| 16:51:28 | sdague | so, actually it looks like it's there under apache, but not running for some reason | |
| 16:51:35 | sdague | or running but not connected | |
| 16:51:46 | cdent | or died | |
| 16:52:18 | efried | Should I try restarting httpd? | |
| 16:52:28 | sdague | efried: wait a sec before doing that | |
| 16:52:36 | efried | suresure. | |
| 16:52:38 | sdague | I want to figure out if there is any other postmortem here | |
| 16:52:54 | sdague | http://paste.openstack.org/show/618073/ - apache status | |
| 16:53:17 | efried | FWIW, the stack itself would have been done a couple weeks ago, and the only service I've been mucking with is n-cpu. | |
| 16:55:02 | sdague | hmmm... well all the access for all services is in the horizon access log | |
| 16:55:07 | sdague | which, is surely a bug | |
| 16:55:30 | sdague | however, it's a bread trail | |
| 16:55:50 | efried | sdague If it's unique to mod-wsgi, it might not be a bug that'll get attention. | |
| 16:55:56 | sdague | yeh | |
| 16:56:12 | sdague | we also don't enable horizon in normal testing | |
| 16:56:18 | cdent | that’s not a mod-wsgi bug, that’s an apache misconfiguration, and is part of the driver to switch to uwsgi in devstack | |
| 16:56:59 | cdent | which ever section ends up listening on port 80 (or 443) with a log configuration will end up getting lots of things that listen on that port | |
| 16:57:02 | efried | I can try to restack with uwsgi, but last time I tried, it didn't work. (Which is why we're still using mod-wsgi in our CI.) | |
| 16:57:57 | efried | (esberglu We either need to figure out the uwsgi thing or add /var/log/apache2/* to our CI log dump.) | |
| 16:58:12 | sdague | ah, you know, I think the crux of it is after the uwsgi cut over we started dropping ports | |
| 16:58:33 | cdent | apache makes the assumption that if you are onthe same port, you’re using the same log | |
| 16:58:52 | sdague | yeh | |
| 16:59:11 | sdague | yeh, this is part of why mod-wsgi is such a pain on the dev side | |
| 16:59:11 | cdent | and when horizon is enabled it is consuming port 80 in a weird way | |
| 16:59:34 | sdague | efried: so my starting point was "nova-manage upgrade check" | |
| 16:59:41 | sdague | which went all 500 stack tracy | |
| 17:00:03 | sdague | yes, placement is returning 500s, however I can't find them landing in a log anywhere | |
| 17:01:08 | sdague | http://paste.openstack.org/show/618074/ horizon_access.log | |
| 17:01:38 | efried | yeah, just what I was seeing. | |
| 17:01:54 | efried | Mine were coming from the compute service trying to suss out the resource providers | |
| 17:02:03 | efried | sdague Well, if you're game to help me debug why uwsgi stack is failing, I can try restacking thusly. | |
| 17:02:11 | sdague | efried: yeh, I would do that | |
| 17:02:15 | efried | That would be a big help in general. | |
| 17:02:16 | sdague | make sure you do a ./clean.sh | |
| 17:02:21 | efried | Okay, rippinit. | |
| 17:02:35 | sdague | yeh, I don't think anything else useful can come from this install | |
| 17:03:01 | efried | I did a lot of useful stuff wrt service catalog lookups. | |
| 17:03:24 | efried | But I was breaking right at driver init to muck around in pdb | |
| 17:04:09 | efried | But now I'm working on converting the placement API over (https://review.openstack.org/#/c/492247/) so I kinda wanted to see it working :) | |
| 17:04:44 | efried | ...and you can see from ^^ that our CI isn't having any trouble with it. | |
| 17:05:51 | sdague | :) | |
| 17:05:59 | efried | (just checked the compute logs to be sure - as if we could pass without placement being happy in the first place - and there's no 500s) | |
| 17:06:46 | sdague | yeh, it's in a weird intermediate state I think, there is a reason why mod_wsgi is something i wanted to remove from the dev/test stack. It just hits a bunch of odd apachisms that don't fit well with our other assumptions | |
| 17:07:17 | efried | sdague To be clear, all I need to do to switch is remove WSGI_MODE=mod_wsgi from my local.conf? | |
| 17:07:22 | sdague | yep | |
| 17:07:57 | efried | k. Refreshing other project clones... | |
| 17:10:34 | efried | aaand stacking... | |
| 17:10:40 | melwitt | mriedem_away: I was wondering if we need to document that for multi-cell with nova-network, upcalls from compute are required for quota checks. I was thinking we might not have to because IIRC we're not supporting multi-cell + nova-network | |
| 17:11:30 | melwitt | I think single cell would be okay because there's no isolation from the API DB there | |
| 17:13:12 | efried | Meanwhile, any idea what glorious magic makes a .txt.gz journalctl log show up in the browser with color codes translated? | |
| 17:21:33 | sdague | efried: magic yet to be written | |
| 17:22:07 | efried | sdague But it works | |
| 17:22:15 | efried | At least in my browser | |
| 17:22:19 | sdague | efried: interesting | |
| 17:22:27 | efried | Except that for powervm logs, it doesn't *quite* work. | |
| 17:22:33 | efried | sdague What, it doesn't do that for you? | |
| 17:23:13 | sdague | efried: oh, it's not color codes translated | |
| 17:23:21 | sdague | you mean the coloring of - http://logs.openstack.org/81/488381/7/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/2b8b331/logs/screen-c-api.txt.gz ? | |
| 17:23:40 | efried | sdague Yup. | |
| 17:23:43 | sdague | https://github.com/openstack-infra/os-loganalyze | |
| 17:24:07 | sdague | is a wsgi filter for the logs | |
| 17:24:22 | efried | ...that runs under the auspices of apached? | |
| 17:24:39 | sdague | yeh, log access is proxied through that | |
| 17:24:41 | efried | ...and would need to exist on the log server. | |
| 17:24:48 | sdague | yes | |
| 17:24:56 | sdague | it processes in real time | |
| 17:24:56 | efried | So ours must be outdated or something. Cool. esberglu ^^ ! | |
| 17:25:12 | sdague | there is a regex on the files it will attempt to do it with | |
| 17:25:17 | efried | sdague Cause for us, it shows up HTMLified, with the log filters, but no color translation. | |
| 17:25:28 | efried | http://184.172.12.213/manual/rest_logs/logs/n-cpu.txt.gz | |
| 17:25:40 | efried | (This isn't in production yet, cause we've been naming 'em .log.gz, but we're converting over) | |
| 17:25:42 | sdague | https://github.com/openstack-infra/os-loganalyze/blob/master/os_loganalyze/filter.py#L24-L31 | |
| 17:25:56 | sdague | yeh, you'd need to setup this code on apache | |
| 17:27:13 | sdague | https://github.com/openstack-infra/puppet-openstackci/blob/master/templates/logs.vhost.erb#L106 and friends to see how you'd set it up on apache | |
| 17:27:27 | sdague | it requires a few rewrite rules | |
| 17:27:49 | efried | sdague Is that not the same script that does the filtering capability? | |
| 17:30:34 | efried | sdague Okay that stack failed, on image create. | |
| 17:31:29 | efried | apache error log shows an ENOSPC in /tmp. I've got .8GB in there; is it really trying to copy the image through /tmp?? | |
| 17:40:53 | sdague | efried: it might be | |
| 17:41:03 | efried | sdague Restacking, gonna watch /tmp... | |
| 17:41:04 | sdague | mtreinish: ^^^ | |
| 17:41:14 | sdague | it might be specific on the uwsgi glance | |
| 17:41:25 | sdague | and, if so, that could be an interesting issue | |
| 17:42:30 | sdague | you all are booting a real image, not a cirros, right? | |
| 17:42:44 | sdague | I could totally imagine that issue to have slipped through because cirros | |
| 17:43:33 | cdent | excellent point, sdague, at least one interation of the uwsgiification of glance needed to do some writes to disk, but I don’t know if that version landed, and if it did where on disk | |