Earlier  
Posted Nick Remark
#openstack-nova - 2017-12-07
20:35:01 mriedem we should consider ignoring those
20:35:35 jaypipes efried: yes, manually executing SQL statements. of course, the object layer won't allow you to do that, though.
20:36:04 jaypipes efried: since a TraitNotFound or ResourceProviderNotFound would be raised when attempting to do that via ResourceProvider.set_traits()
20:36:25 jaypipes efried: because there's a lookup of trait ID to names supplied in set_traits()
20:36:28 efried jaypipes oh, okay, phew. So the answer from the perspective of a REST API consumer is that we're strict about that stuff.
20:36:36 jaypipes efried: yes.
20:36:40 efried Good.
20:36:53 jaypipes efried: except for aggregates, which are just UUIDs and we have nothing to "check" against.
20:37:10 efried jaypipes Right, but you can't *associate* a random anonymous UUID with an aggregate?
20:37:34 jaypipes efried: no. but you *can* associated a random UUID to a known resource provider.
20:37:43 jaypipes efried: it's the aggregate UUID we have no way of checking.
20:37:50 efried jaypipes Got it, cool, thanks.
20:37:55 jaypipes pas de probleme
20:38:40 openstackgerrit Matt Riedemann proposed openstack/nova master: Update nova-status and docs for nova-compute requiring placement 1.14 https://review.openstack.org/526505
20:38:41 efried jaypipes FYI this is coming from a place where I'm refactoring _ensure_resource_provider: In the _create_resource_provider path I can actually skip refresh_aggregate_map
20:38:44 mriedem efried: jaypipes: ^
20:38:59 jaypipes efried: ack.
20:39:03 jaypipes mriedem: thank you sir.
20:39:08 efried mriedem Thanks
20:40:34 mriedem imacdonn: you can ask but you're going to be hard pressed to find anyone that knows much about libvirt+xen in icehouse in channel right now
20:40:40 mriedem anthonyper is your closest bet
20:42:03 imacdonn yeah, I know ... OK .. so the issue is that nova-compute occassionally gets stuck seemingly in trying to talk to libvirt .. the symptoms are that the resource_tracker no longer reports every minute, and any VM operations that need libvirt fail
20:42:32 imacdonn via guru meditation, I can see that the resource_tracker thread is stuck trying to call libvirt's getLibVersion()
20:42:51 mriedem which is then a call to the hypervisor
20:42:57 mriedem so you're likely deadlocking on something in the hypervisor
20:42:58 imacdonn that call is made through eventlet's thread pooling proxy thingy, which is now holding a lock
20:43:06 mriedem yeah that could also be screwing you
20:43:16 mriedem i'd enable debug logging for libvirt and see if something shows up in there
20:43:31 mriedem or, check to see if things changed around that code since icehouse and see if you need to backport a fix
20:43:52 imacdonn problem is it happens once in a while, and I have like 2k compute nodes ... don't really want to to turn debug on on all of them and wait
20:44:16 imacdonn I've looked around, but not found anything that looks like an obvious related fix
20:45:03 imacdonn I can't tell for sure if it's libvirt hanging on the call, or eventlet getting hung up somehow and not even trying the call
20:48:43 mriedem imacdonn: well, i see this in kilo https://review.openstack.org/#/c/104930/
20:49:16 mriedem https://review.openstack.org/#/c/104930/12/nova/virt/libvirt/host.py@192
20:49:17 mriedem so,
20:49:18 imacdonn yeah, that made it "fun" to try to compare bits of of the code to see what might have changed
20:49:29 mriedem keep in mind that anything running in those threads that logs anything could lock you up
20:49:48 imacdonn yes, saw some stuff about that
20:49:50 mriedem so if you have a GMR when things are locked, i'd look for any libvirt driver/host methods in the thread dump,
20:49:53 mriedem and see if those do loging
20:49:55 mriedem *logging
20:50:22 imacdonn GMR is at https://pastebin.com/1jfgdurJ
20:50:29 mriedem because i'm sure we don't do a good job of auditing stuff like that
20:50:31 imacdonn getLibVersion() call at line 659
20:51:00 imacdonn victim thread waiting to acquire() lock around line 777
20:55:36 imacdonn seems like getLibVersion() isn't asking much of libvirt ... don't think it'd even have to talk to the hypervisor ...
20:56:45 mriedem well, it's making a connection to libvirt i believe
20:56:58 mriedem so it's not like the libvirt-python package version or something
20:57:11 mriedem nova meeting in 3 minutes
20:57:42 imacdonn it'd make a native library call at least .. dunno if it'd have to connect to libvirtd .. OK, I'll shut up for now ;)
20:58:03 mriedem true yeah
20:59:21 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Separate API schemas (usage) https://review.openstack.org/520603
20:59:37 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Separate API schemas (inventory) https://review.openstack.org/520613
21:00:03 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Separate API schemas (resource_class) https://review.openstack.org/520611
21:00:03 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Separate API schemas (aggregate) https://review.openstack.org/520608
21:00:04 mriedem melwitt: might want to check out v
21:00:05 mriedem https://bugs.launchpad.net/nova/+bug/1737011
21:00:06 openstack Launchpad bug 1737011 in OpenStack Compute (nova) "ServerActionsTestJSON.test_reboot_server_hard failed to ssh into instance" [Undecided,New]
21:00:15 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Separate API schemas (trait) https://review.openstack.org/520605
21:00:32 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Add x-openstack-request-id in API ref https://review.openstack.org/523007
21:00:44 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Add 'Location' parameters in API ref https://review.openstack.org/521541
21:01:14 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Fix getting placement request ID https://review.openstack.org/523606
21:01:57 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Add functional tests for resource class API https://review.openstack.org/524506
21:02:09 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Add functional tests for traits API https://review.openstack.org/524094
21:03:03 openstackgerrit Takashi NATSUME proposed openstack/nova master: [cellv2] Improve getting BDMs in multiple cells https://review.openstack.org/521400
21:03:28 openstackgerrit melanie witt proposed openstack/nova master: Add API and nova-manage tests that use the NoopQuotaDriver https://review.openstack.org/526270
21:03:28 openstackgerrit melanie witt proposed openstack/nova master: Follow up on removing old-style quotas code https://review.openstack.org/524234
21:13:44 openstackgerrit Takashi NATSUME proposed openstack/nova master: List/show all server migration types (1/2) https://review.openstack.org/430608
21:14:42 openstackgerrit Takashi NATSUME proposed openstack/nova master: List/show all server migration types (2/2) https://review.openstack.org/459483
21:16:31 melwitt mriedem: is that different than our usual sporadic ssh timeouts?
21:19:33 mriedem melwitt: i don't think ssh timeouts are all that sporadic anymore
21:19:35 mriedem andreaf: ^?
21:20:19 melwitt okay, I saw one recently that I had to recheck so I thought they were still going on. I guess I should look at that one and compare
21:21:38 mriedem i don't see cloud-init run at all
21:24:23 mriedem force_config_drive = True
21:24:32 mriedem so we force a config drive to inject keys and such
21:27:01 mriedem we don't even get to the point of rebooting the serer
21:27:02 mriedem *server
21:27:08 mriedem it's trying to ssh into the guest to check uptime before that
21:27:10 mriedem and that's what fails
21:27:11 mriedem http://logs.openstack.org/83/526183/1/gate/legacy-tempest-dsvm-py35/166f0c9/job-output.txt.gz#_2017-12-07_13_37_58_430147
21:27:32 melwitt good eye
21:28:10 mriedem the console output is all from tempest trying to gather information before the test pukes
21:28:14 jaypipes dansmith: from the API layer, if I want to find which cell a compute node (note: not the service host, but the Ironic baremetal node) was in, how would I do that? do I loop through cells doing a query?
21:28:56 mriedem jaypipes: i think you'd have to
21:29:11 mriedem the host mapping is the compute_node.host, not compute_nodes.hypervisor_hostname which is the node name
21:29:46 jaypipes mriedem: right
21:33:48 melwitt mriedem: this is the change I was thinking of where I had to recheck it about a week ago. looks like the same deal http://logs.openstack.org/22/518022/8/check/legacy-tempest-dsvm-neutron-full/81721fe/job-output.txt.gz#_2017-11-30_21_57_15_827380
21:34:26 mriedem melwitt: http://logs.openstack.org/22/518022/8/check/legacy-tempest-dsvm-neutron-full/81721fe/job-output.txt.gz#_2017-11-30_21_57_15_838147
21:35:02 melwitt oh heh
21:35:41 melwitt so not the same
21:37:46 mriedem melwitt: http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22Kernel%20panic%20-%20not%20syncing%5C%22%20AND%20tags%3A%5C%22console%5C%22&from=7d
21:38:08 mriedem nothing super obvious there, not like a single node provider
21:38:16 mriedem but it's all master branch, so i wonder if we're using a new cirros image in queens
21:39:15 melwitt how do we check that?
21:39:55 mriedem it's in devstack
21:40:10 mriedem https://github.com/openstack-dev/devstack/blob/master/stackrc#L671
21:40:25 mriedem https://github.com/openstack-dev/devstack/commit/9f2dcd333103553626db1924a019e151e3e7252e
21:40:28 melwitt cool thanks
21:40:29 mriedem that's not new so...

Earlier   Later