Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-26
18:46:34 melwitt I thought he did too but I'm getting a little confused
18:46:53 mriedem yesterday i didn't have any instances in ERROR state, so they were all in the cell
18:47:04 mriedem i deleted all of those and then archived cell0 and cell1
18:47:21 dansmith deleted them locally?
18:47:21 mriedem today i've been trying to get 500 ERROR during scheduling, and 500 ACTIVE
18:47:24 mriedem based on the flavor i use
18:47:30 mriedem dansmith: deleted via the api
18:47:37 dansmith mriedem: with compute down or no?
18:47:38 mriedem remember me complaining about how long that was taking yesterday?
18:47:39 mriedem no
18:47:48 mriedem took 2+ hours to delete 1000 ACTIVE instances
18:48:03 dansmith I do, but I didn't remember all your details
18:48:20 mriedem yeah i basically trying to get back to clean state before starting today
18:48:33 mriedem so was archiving the db's last night
18:49:19 mriedem btw, bauzas pointed this out before, but we log this way too many times
18:49:19 mriedem Sep 26 18:44:37 devstack nova-compute[30351]: DEBUG nova.compute.resource_tracker [None req-992d494e-d328-4204-bcfe-80d926cf0a65 demo demo] We're on a Pike compute host in a deployment with all Pike compute hosts. Skipping auto-correction of allocations. {{(pid=30351) _update_usage_from_instance /opt/stack/nova/nova/compute/resource_tracker.py:1071}}
18:52:10 dansmith mriedem: unrelated, see this: http://status.openstack.org/openstack-health/#/test/nova.tests.functional.test_servers.ServersTestV219.test_description_errors?duration=P3M
18:52:34 dansmith mriedem: I think this test is occasionally taking up to 240s locally when it should be about 8s
18:53:00 mriedem jesus
18:53:02 dansmith and I think it's because it creates a server that it never cleans up and then abruptly exits where we take down conductor before the compute service finishes waiting on a call or something
18:53:16 dansmith so I have a patch to just make it clean up the server and I _think_ it's working
18:53:32 mriedem the one weird spike in august is, weird
18:53:45 mriedem https://bugs.launchpad.net/nova/+bug/1719714
18:53:46 openstack Launchpad bug 1719714 in OpenStack Compute (nova) "Excessive logging of "We're on a Pike compute host in a deployment with all Pike compute hosts."" [Medium,Confirmed]
18:54:04 dansmith mriedem: it would have just been ordering reasons
18:54:59 dansmith mriedem: note the rising tail at present too
19:18:31 mriedem alright i'm just going to restack
19:18:32 mriedem nuts to this
19:28:40 mriedem dansmith: jaypipes: bauzas: https://review.openstack.org/#/c/498947/6
19:28:45 mriedem that test_servers thing is wrong
19:29:19 openstackgerrit Matthew Treinish proposed openstack/nova master: Add slowest command to tox.ini https://review.openstack.org/507657
19:29:21 mtreinish dansmith: ^^^
19:29:29 mriedem there are 2 tests for failures during evacaute on the dest
19:29:38 mriedem 1. test_evacuate_claim_on_dest_fails - that is testing when the claim fails with ComputeResourcesUnavailable
19:29:57 mriedem 2. test_evacuate_rebuild_on_dest_fails - that is testing when the claim is successful but the driver.rebuild method raises some exception
19:29:57 dansmith mtreinish: sweet
19:30:00 jaypipes mriedem: sorry, I disagree with you.
19:30:19 mriedem i wrote those tests
19:30:29 mriedem so please explain how i'm wrong that they are now made redundant in that change
19:30:36 jaypipes mriedem: that test raising TestingException was not useful. Because TestingException isn't what is ever raised by any code.
19:30:54 mriedem it's simulating the virt driver raising the error during rebuild
19:31:00 mriedem AFTER the successful claim
19:31:10 mriedem it could be ProcessExecutionError
19:31:12 mriedem from driver.spawn()
19:31:14 mriedem if you like
19:31:44 mriedem these 2 tests are testing very specific failures
19:31:45 dansmith mriedem: right but we don't run the claim teardown code in that case
19:32:00 mriedem dansmith: correct, which is why we run the allocation cleanup manually
19:32:03 mriedem and that's what that is testing
19:32:32 mriedem the test you changed isn't meant to test drop_move_claim
19:32:35 mriedem the docstring explains that
19:32:44 jaypipes mriedem: if the point of the test (as is in that docstring) is to ensure allocations are cleaned up after a failed rebuild, then the test should raise the exception that would be raised *after* a claim has been made for the new resources.
19:32:55 dansmith jaypipes: he's saying another one does that
19:33:14 mriedem jaypipes: you realize the virt drivers can raise any kinds of crazy shit right?
19:33:15 dansmith mriedem: so in this case you want the test to validate that the allocations _don't_ get cleaned up is that right?
19:33:31 jaypipes mriedem: Matt, I'm trying to be civil.
19:34:38 mriedem https://review.openstack.org/#/c/499877/
19:35:22 dansmith this is what it's testing: https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2800-L2827
19:35:23 dansmith the except exception case of that
19:35:24 mriedem so ^ is testing that drop_move_claim removes the allocation when the claim was successful but the virt driver raised some exception
19:35:37 jaypipes mriedem: OK, I see that now.
19:36:18 mriedem https://review.openstack.org/#/c/499874/ added the other test
19:36:54 mriedem that was a recreate test for a bug
19:36:55 mriedem fixed in https://review.openstack.org/#/c/499878/
19:37:38 dansmith mriedem: we get it
19:37:51 dansmith mriedem: can you answer my question above about what you want it to do?
19:39:32 mriedem the test should go back to whatever it was testing
19:39:43 mriedem which is the case that the claim passes, but the virt driver raises
19:39:51 mriedem so we'd remove the allocation via drop_move_claim before
19:40:02 dansmith right, but you assert some behavior that happens inside drop_move_claim
19:40:07 dansmith which no longer happens
19:40:35 mriedem then that drop_move_claim behavior has to be replayed elsewhere i guess
19:40:37 jaypipes what if the it's a same-host rebuild? :(
19:40:45 mriedem there is no claim for a same host rebuild
19:40:49 jaypipes k
19:40:54 mriedem so you wouldn't hit ComputeResourcesUnavailable
19:41:06 dansmith exactly, but you could hit other exceptions
19:41:21 jaypipes mriedem: but you *would* hit the TestingException "crazy shit"
19:41:40 jaypipes mriedem: and you're asserting that we'd delete the allocation against the instance in that case, right?
19:41:43 mriedem that doesn't have anything to do with dropping an allocation though
19:41:46 mriedem no
19:42:19 jaypipes mriedem: oh, sorry, you're asserting that the *update_available_resource()* call would clean up allocations for a failed build?
19:42:22 jaypipes rebuild.
19:42:26 mriedem no
19:42:29 dansmith no
19:42:33 jaypipes guh
19:42:40 dansmith the test _is_ asserting that the dest host's allocation was cleaned up by drop_move_claim
19:42:43 mriedem we don't ever want to remove allocations for a *rebuild*
19:42:55 mriedem the tests are specifically for evacuate
19:43:01 mriedem where the scheduler creates allocations on the dest host
19:43:13 mriedem we fail the evacuate on the dest host, so we need to remove those allocatoins created by the scheduler
19:44:04 dansmith there's a specific reason why I made this change,
19:44:11 dansmith and I talked it through with jaypipes which is why I made this
19:44:11 mriedem i'm sorry for being grouchy about this,
19:44:23 mriedem but i've spent the better part of the last 6 weeks fixing these allocation bugs,
19:44:26 dansmith so I'll have to go re-load all my context on this before I can really think about it
19:44:29 mriedem so being told i don't understand the test pisses me off
19:44:36 jaypipes mriedem: understood. and you're saying that you want the drop_move_claim() to remove those resources when ComputeResourcesUnvailable is raised but you want update_available_resource() to delete the allocations when a virt driver exception is raised?
19:44:47 dansmith jaypipes: no
19:45:48 dansmith mriedem: without offending your test sensibilities, you see this right? https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2812

Earlier   Later