Earlier  
Posted Nick Remark
#openstack-nova - 2018-01-25
20:54:29 mriedem i know it's not in the default list of log levels in oslo.log
20:56:05 efried but if it's explicitly in whatever conf (on customer site), we don't want to override it.
20:59:52 mriedem that's why it's default_log_levels yeah?
21:00:30 efried Could be, could be. I don't know. I left comments, but +0 - smarter people can decide whether those things are really an issue.
21:01:40 mriedem look efried, do you want your 20 patches to merge or not
21:01:41 mriedem :)
21:02:11 efried I have NO IDEA what's actually best for my karma at this point. Walkin on eggshells, here.
21:03:48 mriedem yeah, i'm about to just call it a day
21:03:56 mriedem not really much that can be done about anything
21:04:36 efried No, you have to stick around to +2+W stuff so us peons can sit around rechecking them through the night.
21:04:57 mriedem i was rechecking things until about 11:30 last night
21:05:00 mriedem i've done my job
21:06:28 efried BTW, I was tooling around on stackalytics the other day - how did you manage to merge 18KLOC while pushing PTL paperwork 18 hours a day?
21:06:46 efried Impressive. Most impressive.
21:08:11 Roamer` hm, so here's a stupid question about stackalytics: if I put my name and e-mail address in default_data.json, will it pick up the patches that have already been merged, or does it just go forward?
21:09:46 efried Roamer` me no idea. Someone in -infra might know, but they're frying great big fish at the moment.
21:13:12 Roamer` efried, thanks, and yeah, I'm following the discussions there
21:14:18 openstackgerrit Brianna Poulos proposed openstack/nova master: Reduce complexity of _from_db_object https://review.openstack.org/537728
21:14:19 openstackgerrit Brianna Poulos proposed openstack/nova master: Add trusted_certs object https://review.openstack.org/489408
21:14:19 openstackgerrit Brianna Poulos proposed openstack/nova master: Add trusted_certs to instance_extra https://review.openstack.org/537897
21:14:20 openstackgerrit Brianna Poulos proposed openstack/nova master: Add trusted_image_certificates to REST API https://review.openstack.org/486204
21:14:20 openstackgerrit Brianna Poulos proposed openstack/nova master: Implement certificate_utils https://review.openstack.org/479949
21:15:41 openstackgerrit Merged openstack/nova master: Update VMWare vSphere link address https://review.openstack.org/535244
21:17:38 mriedem efried: don't know
21:17:50 mriedem amphetamines
21:18:06 mriedem likely api changes with lots of api samples
21:18:07 efried Tell ya, wait til you're my age. You'll never have the stamina.
21:18:20 mriedem that's why i take that thing that frank thomas is selling
21:18:32 mriedem nugenix
21:18:55 mriedem https://www.ispot.tv/ad/wZA8/nugenix-fan-featuring-frank-thomas
21:19:13 mriedem "stronger and leaner"
21:19:18 efried a powerful, man-boosting innovation
21:19:19 mriedem the guy has been like 300lb for 20 years
21:19:32 mriedem man boosting testafen
21:19:42 mriedem my passion levels will be through the roof!
21:27:01 edleafe It's cute watching you young 'uns talking about being old.
21:29:16 dansmith hush, dad's home!
21:29:34 melwitt hah :D
21:29:54 mriedem i'll grab the hearin' cone
21:30:25 dansmith that said, I'm quite sure edleafe remembers being 36 and feeling like life was over
21:33:07 efried Any time I think of 36... https://youtu.be/q6zhOdroJCE?t=1347
21:33:49 mriedem umm, yikes
21:38:12 melwitt for the native thread logging related gate issues, will there always be a message somewhere in the logs that says "cannot switch to a different thread"?
21:38:26 mriedem yeah
21:39:35 melwitt okay. I was wondering if it could be related to the stable branch cells v1 job 100% fails. libvirt is getting "connection broke" as though something is holding up the tpool and never letting it continue
21:39:55 melwitt (I noticed all the libvirt calls go through eventlet.tpool)
21:40:12 melwitt but the "cannot switch to a different thread" isn't in the logs anywhere
21:41:11 melwitt I've been investigating it but haven't figured out anything yet
21:41:58 dansmith the libvirt driver uses native threads properly
21:42:08 dansmith so it should be fine unless something has regressed (or was never right)
21:42:40 dansmith the connection issue with libvirt almost always means libvirtd crashed,
21:42:51 dansmith which happens a lot more than it should
21:43:59 melwitt yeah. in this case it hasn't crashed. it's as if something held it up so long that the connection dropped from not getting keepalive, or something like that
21:44:19 Roamer` melwitt, what OS is this happening on? Could it be the Ubuntu py27/eventlet problem? http://lists.openstack.org/pipermail/openstack-dev/2018-January/126580.html
21:44:34 melwitt I was wondering if any other library also using eventlet could have made it so it never switches back to the libvirt driver
21:44:46 dansmith melwitt: why do you say it hasn't crashed?
21:45:25 melwitt dansmith: because from syslog and libvirtd I couldn't find any crashing. unless I totally don't know what I'm looking at, which is possible
21:45:45 dansmith melwitt: link?
21:45:46 melwitt Roamer`: it's ubuntu xenial. and that's what I was wondering, but I'm still grasping at straws at this point
21:46:34 melwitt dansmith: here's an example from recently http://logs.openstack.org/04/536904/3/check/legacy-tempest-dsvm-cells/13aacfc/logs/screen-n-cpu.txt.gz#_2018-01-25_19_06_13_271
21:50:24 dansmith hmm, empty journal on that one?
21:50:25 dansmith that's weird
21:51:26 melwitt I feel like I've seen that before
21:52:56 dansmith yeah, libvirtd doesn't look like it was restarted across that connection failure boundary
21:52:58 mriedem interestingly, it's the first guest to be created it looks like http://logs.openstack.org/04/536904/3/check/legacy-tempest-dsvm-cells/13aacfc/logs/libvirt/qemu/instance-0000000a.txt.gz
21:53:17 dansmith this would be the first time I've ever seen that symptom not associated with a crash
21:53:34 dansmith so I still feel a little doubtful, but it's certainly not as glaring as normal
21:53:41 mriedem makes me think of this https://review.openstack.org/#/c/526748/
21:53:50 mriedem where guestfs on the first bring up runs in a tpool
21:53:56 mriedem to get the guestfs capabilities
21:54:14 dansmith so I wonder if it's something like qemu crashes and the libvirt thread we're talking to exits or something
21:54:23 melwitt yeah, that's kinda how I found that patch of yours mriedem
21:54:48 dansmith the journal should have dmesg in it to capture any segv
21:55:07 dansmith but we don't have qemu logs otherwise
21:55:10 melwitt I wonder if the journal collection on the jobs has been busted this whole time
21:55:32 melwitt or is it not always empty?
21:55:51 dansmith it wasn't busted a month or so ago
21:55:52 melwitt I haven't tried to look at it before
21:55:56 melwitt oh okay
21:56:24 dansmith I expect it would be in syslog too, but it's clearer in the journal
21:57:25 melwitt I didn't see anything that looked like crashing in syslog http://logs.openstack.org/04/536904/3/check/legacy-tempest-dsvm-cells/13aacfc/logs/syslog.txt.gz
21:57:38 dansmith 2018-01-25 19:06:13.215+0000: 3732: error : virKeepAliveTimerInternal:143 : internal error: connection closed due to keepalive timeout
21:57:45 dansmith is right before you get it in nova
21:59:13 melwitt where do you see that? libvirtd.log?
21:59:27 dansmith I think this is all in a separate thread in nova so it shouldn't be that nova is hung up and unable to do the keepalive bit, unless it's just system starvation of some sort
21:59:30 dansmith yeah, libvirt
21:59:50 melwitt damn
22:00:06 dansmith and nova is doing stuff milliseconds before that, so it's not hung I don't think
22:00:51 dansmith oh yeah in the nova log:
22:00:52 dansmith 2018-01-25 19:06:13.224 16016 INFO nova.virt.libvirt.driver [req-ea2c4c13-3bbe-4b35-a2b3-84976c11fc4b - -] Connection event '0' reason 'Connection to libvirt lost'
22:01:08 dansmith that's the real thread detecting the connection close
22:01:20 dansmith oooh, but
22:01:49 dansmith 2018-01-25 19:06:09.554 16016 DEBUG oslo_messaging._drivers.amqpdriver [-] received reply msg_id: 9ebc51b2583b4f4da811e951e41068c9 __call__ /usr/local/lib/python2.7/dist-packages/oslo_messaging/_drivers/amqpdriver.py:407
22:01:50 dansmith 2018-01-25 19:06:13.218 16016 DEBUG nova.virt.libvirt.firewall [req-733b7704-8aa3-45a5-8798-67eab246baf8 tempest-InstanceActionsV221TestJSON-176157016 tempest-InstanceActionsV221TestJSON-176157016] Cannot find UUID for filter 'nova-instance-instance-0000000d-fa163e269665': 'Cannot recv data: Connection reset by peer' _get_filter_uuid /opt/stack/new/nova/nova/virt/libvirt/firewall.py:254
22:02:00 dansmith that is five seconds apart with nothing in between
22:02:48 openstackgerrit Jim Rollenhagen proposed openstack/nova master: Add functional tests for virt driver get_traits() method https://review.openstack.org/538047
22:02:58 melwitt hm
22:03:04 dansmith meh, I dunno, looking around that I don't think that's anything related to being stuck though
22:03:14 dansmith because that's the exact instant libvirt logs its message
22:06:48 openstackgerrit Jim Rollenhagen proposed openstack/nova master: Add functional tests for virt driver get_traits() method https://review.openstack.org/538047

Earlier   Later