Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
15:28:28 mriedem sdague: thanks
15:28:34 dansmith mriedem: okay, like I said in the etherpad, I don't think we're doing the fault thing in that patch yet
15:28:47 dansmith and applying the one that does it definitely has an impact
15:29:15 mriedem negative impact?
15:29:37 dansmith yeah
15:30:05 mriedem because we've added a new unconditional join i suppose
15:30:19 dansmith it's not a new join, but it's a new query yeah
15:30:25 mriedem yeah, was just thinking that
15:30:43 dansmith I'm messing with the later patch though to see if I can squash it out
15:30:45 mriedem like you said, we could plumb that in the db api
15:30:53 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif stable/newton: Updated from global requirements https://review.openstack.org/373293
15:31:05 mriedem if faults is in expected_attrs, get the instances in deleted/error and include the faults on those?
15:31:05 dansmith that's what I'm doing yeah
15:32:04 dansmith yes
15:36:47 cdent do dansmith and mriedem have an exciting new etherpad they are willing to share?
15:37:06 dansmith cdent: https://etherpad.openstack.org/p/nova-instance-list
15:37:12 dansmith cdent: but that's just incidental to him finding the problem
15:37:21 dansmith it's not about the thing we were describing
15:37:35 cdent you know I’m a total junkie for all the info
15:37:55 mriedem started as getting a benchmark for testing before and after the big instance list change
15:38:10 mriedem but have also added todos for weird things we've seen, like the concurrent update detected thing with 500 instances
15:42:09 mriedem sdague: can i put [[post-config|$NOVA_CONF]] in stackrc?
15:42:44 cdent mriedem, dansmith, (and sean mooney): If any of you get a chance to look at the discussion on https://review.openstack.org/#/c/504540/ (limiting GET /allocation_candidates ) it’s gotten to the point where we are trying to decide what it is that we are actually optimizing for, so could do with more input
15:43:17 sdague mriedem: in local.conf
15:43:42 mriedem sdague: yeah, locally, but this is for running something through ci
15:44:51 sdague there isn't really anyway to jam it into stackrc, you need to do project-config changes for stuff like that
15:45:15 sdague if this is just for hacktastic stuff, just iniset whatever you want in lib/nova
15:45:16 mriedem ok, just checking, i can hack this other ways
15:45:21 mriedem yeah that's what i'm doing
15:55:01 dansmith mriedem: so with /servers/details on 2.53, we're pegging conductor real hard. With 2.1 there is no conductor interaction at all
15:55:13 dansmith so we must be doing something stupid in the later microversion
15:55:19 dansmith because we shouldn't be making rpc calls at all
15:55:47 dansmith also, defeating all fault loading and all tag loading doesn't make anything faster
15:56:31 mriedem how about services?
15:57:43 dansmith services is included in the 2.1 columns so I didn't try
15:59:04 cdent johnthetubaguy: you still want to hold your -2 on https://review.openstack.org/#/c/270116/ ? It’s got a blueprint now
16:00:41 mriedem oh yeah
16:00:49 melwitt mriedem: I'm working on the regression func test for the reschedule bug, FYI
16:00:51 dansmith however, I can stop compute and conductor and still do the list with no failure (and no difference in speed)
16:00:55 dansmith wtaf
16:01:12 mriedem cdent: the spec on that isn't approved
16:01:53 mriedem cdent: see my comment from may 26 on that patch
16:02:48 cdent mriedem: okay
16:03:13 openstackgerrit Balazs Gibizer proposed openstack/nova master: Fix race in delete allocation in ServerMovingTests https://review.openstack.org/507911
16:03:31 gibi mriedem: I've pushed a fix for https://bugs.launchpad.net/nova/+bug/1719915 ^^
16:03:34 openstack Launchpad bug 1719915 in OpenStack Compute (nova) "test_live_migrate_delete race fail when checking allocations: MismatchError: 2 != 1" [Medium,In progress] - Assigned to Balazs Gibizer (balazs-gibizer)
16:04:36 mriedem thanks
16:04:59 cdent mriedem: I shall continue to remind the keepers of the CI
16:06:09 dansmith mriedem: 2.26 adds ~3s to my runtime
16:07:25 dansmith mriedem: 2.16 adds about 0.250s
16:08:06 cdent sdague: you’ve comment on the bug report related to https://review.openstack.org/#/c/501359/ , can you comment on the fix when you get a chance. _might_ have backport potential. stephenfin you willing to upgrade your +1?
16:08:27 dansmith mriedem: 2.26 was tags, btw
16:08:50 dansmith mriedem: so this is 3s with tags short-circuited at the db layer, so I think the 3s is all api overhead for empty things
16:09:38 stephenfin cdent: Yup, happy to +2 once someone sdague or dansmith has looked at it (I'm no expert in that area)
16:09:53 cdent thanks stephenfin
16:10:11 johnthetubaguy cdent: its normally dropped when the blueprint is approved, I don't remember how spec-less get approved now
16:10:35 cdent johnthetubaguy: s’okay, matt’s cleared things up: until CI is super happy the blueprint won’t get approved
16:10:54 johnthetubaguy cdent: ah, cool, I should read the scrollback better
16:10:54 cdent i hadn’t seen his comment in the middle of the stack
16:11:04 cdent and I should read the comments better :)
16:11:25 johnthetubaguy oh yeah, I see now
16:13:26 mriedem dansmith: cdent: this is my super hack devstack patch to try and recreate the 500 instance burst failure https://review.openstack.org/507918
16:16:18 mriedem dansmith: ok so we still don't know which microversion is making instance list go back through conductor
16:16:44 dansmith mriedem: I have conductor stopped and nothing else is failing
16:16:46 dansmith which I can't explain
16:17:31 mriedem hmm, api going straight to db somewhere?
16:17:48 dansmith api should be going straight to the db everywhere
16:18:03 dansmith I'm not sure why conductor was doing anything during a list in the first place
16:18:11 dansmith I tried turning it off to see what broke and nothing did
16:18:31 mriedem maybe you hit a window where a periodic was hitting conductor at the same time as you were doing the instance list?
16:18:40 dansmith could be, but I did it a few times
16:19:14 dansmith either way, I'm going to go measure the 2.26 impact with just my change (not the short-circuiting i've done) and on master and see what the diff is
16:20:46 mriedem i'm going to go preheat the oven because it's going to be pot pie time in about an hour
16:26:06 openstackgerrit Balazs Gibizer proposed openstack/nova master: Fix race in delete allocation in ServerMovingTests https://review.openstack.org/507911
16:33:57 openstackgerrit Merged openstack/nova master: Fix IoOpsFilter test case class name. https://review.openstack.org/507205
16:34:37 openstackgerrit Merged openstack/nova master: libvirt: bandwidth param should be set in guest migrate https://review.openstack.org/497455
16:35:16 openstackgerrit Merged openstack/nova stable/ocata: Provide hints when nova-manage db sync fails to sync cell0 https://review.openstack.org/501746
16:35:38 openstackgerrit Merged openstack/nova master: Ensure errors_out_migration errors out migration https://review.openstack.org/479802
16:35:49 efried sdague got an opinion on https://review.openstack.org/#/c/488137/21/nova/conf/utils.py@85 ?
16:36:28 johnsom I have an instance booted in nova (master) that nova/neutron shows two plugged ports, but the kernel is not seeing the second network interface. It was hot-plugged with attach. Any pointers for debugging this?
16:37:19 johnsom the qemu process command line (ps -ef) only shows one interface, but I'm not sure if it should show a hot-plugged network interface or not.
16:37:43 johnsom We have been seeing this in our gates off and on during Pike, but I just had it happen local so I can debug, etc.
16:38:20 openstackgerrit Chris Dent proposed openstack/nova master: DNM: Don't monkey patch eventlet in functional https://review.openstack.org/506668
16:38:21 openstackgerrit Chris Dent proposed openstack/nova master: Do not monkey patch eventlet in unit tests https://review.openstack.org/507923
16:39:53 openstackgerrit Merged openstack/python-novaclient stable/pike: Allow boot server with multiple nics https://review.openstack.org/495901
16:46:14 johnthetubaguy johnsom: were there any nova-compute logs about why the attach failed?
16:46:38 johnsom I am looking through and collecting those, n-cpu?
16:47:43 johnthetubaguy yeah, thats the ones I think
16:47:45 johnsom nova show lists it as ACTIVE
16:48:20 johnthetubaguy it wouldn't got to an error state if it failed
16:48:25 openstackgerrit Matt Riedemann proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505673
16:48:52 johnthetubaguy AFAIK
16:48:57 johnsom Yeah, n-cpu doesn't have any ERROR level messages. I can see some vif lines related to the interface, but no ERROR
16:49:06 johnthetubaguy its probably a warn
16:49:28 johnthetubaguy do you see the OVS bits wired up?
16:49:54 johnsom Sep 27 08:48:44 devstackpy27-2 nova-compute[21517]: WARNING nova.compute.manager [None req-48a86b9a-bf96-4f0e-bc60-00682c991e35 service nova] [instance: fb013f87-2e20-42d7-950d-bc9add853f2c] Received unexpected event network-vif-plugged-37ea16ee-b9bc-48c8-b23b-1221bece7c9a for instance with vm_state active and task_state None.
16:50:13 johnthetubaguy how did you do the attach?
16:50:17 johnthetubaguy via the Nova API?
16:50:20 johnsom Yes
16:51:07 johnthetubaguy its probably a case of tracing that through the code following the logs, seeing where it failed, I suspect in n-cpu but it could have been earlier

Earlier   Later