| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 18:38:27 | mriedem | yeah | |
| 18:38:30 | mriedem | wasn't a problem yesterday | |
| 18:38:37 | melwitt | oh | |
| 18:38:41 | mriedem | but i had a bit of a cleaner env yesterday | |
| 18:39:46 | melwitt | multi-create will reject you if any one of min_count can't be accommodated. so one at a time would work if you're in that situation, if some/most of them fit | |
| 18:40:22 | mriedem | yesterday i created 100, then like 200, then 500 more or something | |
| 18:40:24 | mriedem | eventually got to 1000 | |
| 18:41:49 | mriedem | i can just restack this env, but it makes me worry that we aren't properly cleaning up allocations somewhere | |
| 18:42:37 | melwitt | yeah | |
| 18:43:15 | melwitt | did you say yesterday you don't have computes, or something like that? I just wonder what happens with FakeDriver, if it somehow doesn't call the healing allocations code | |
| 18:43:38 | mriedem | we don't heal since pike | |
| 18:43:51 | mriedem | if you don't have computes < pike, we don't heal | |
| 18:44:04 | mriedem | this is just a single compute, single node devstack | |
| 18:44:11 | mriedem | with the fake driver and noop quota | |
| 18:44:22 | melwitt | oh wait, sorry I was thinking of local delete | |
| 18:44:53 | melwitt | I was trying to think if with FakeDriver, does the code that deletes allocations when an instance is deleted, run | |
| 18:45:02 | melwitt | or if that even matters | |
| 18:45:17 | mriedem | the compute manager cleans up allocations when an instance is deleted | |
| 18:45:34 | dansmith | mriedem: we heal for deletes | |
| 18:45:44 | dansmith | mriedem: did you archive them after local delete before you started up? | |
| 18:46:01 | mriedem | i've been archiving yeah | |
| 18:46:05 | melwitt | I'm not sure whether he had local deletes | |
| 18:46:06 | dansmith | so that's why | |
| 18:46:14 | dansmith | I thought he did | |
| 18:46:34 | melwitt | I thought he did too but I'm getting a little confused | |
| 18:46:53 | mriedem | yesterday i didn't have any instances in ERROR state, so they were all in the cell | |
| 18:47:04 | mriedem | i deleted all of those and then archived cell0 and cell1 | |
| 18:47:21 | dansmith | deleted them locally? | |
| 18:47:21 | mriedem | today i've been trying to get 500 ERROR during scheduling, and 500 ACTIVE | |
| 18:47:24 | mriedem | based on the flavor i use | |
| 18:47:30 | mriedem | dansmith: deleted via the api | |
| 18:47:37 | dansmith | mriedem: with compute down or no? | |
| 18:47:38 | mriedem | remember me complaining about how long that was taking yesterday? | |
| 18:47:39 | mriedem | no | |
| 18:47:48 | mriedem | took 2+ hours to delete 1000 ACTIVE instances | |
| 18:48:03 | dansmith | I do, but I didn't remember all your details | |
| 18:48:20 | mriedem | yeah i basically trying to get back to clean state before starting today | |
| 18:48:33 | mriedem | so was archiving the db's last night | |
| 18:49:19 | mriedem | btw, bauzas pointed this out before, but we log this way too many times | |
| 18:49:19 | mriedem | Sep 26 18:44:37 devstack nova-compute[30351]: DEBUG nova.compute.resource_tracker [None req-992d494e-d328-4204-bcfe-80d926cf0a65 demo demo] We're on a Pike compute host in a deployment with all Pike compute hosts. Skipping auto-correction of allocations. {{(pid=30351) _update_usage_from_instance /opt/stack/nova/nova/compute/resource_tracker.py:1071}} | |
| 18:52:10 | dansmith | mriedem: unrelated, see this: http://status.openstack.org/openstack-health/#/test/nova.tests.functional.test_servers.ServersTestV219.test_description_errors?duration=P3M | |
| 18:52:34 | dansmith | mriedem: I think this test is occasionally taking up to 240s locally when it should be about 8s | |
| 18:53:00 | mriedem | jesus | |
| 18:53:02 | dansmith | and I think it's because it creates a server that it never cleans up and then abruptly exits where we take down conductor before the compute service finishes waiting on a call or something | |
| 18:53:16 | dansmith | so I have a patch to just make it clean up the server and I _think_ it's working | |
| 18:53:32 | mriedem | the one weird spike in august is, weird | |
| 18:53:45 | mriedem | https://bugs.launchpad.net/nova/+bug/1719714 | |
| 18:53:46 | openstack | Launchpad bug 1719714 in OpenStack Compute (nova) "Excessive logging of "We're on a Pike compute host in a deployment with all Pike compute hosts."" [Medium,Confirmed] | |
| 18:54:04 | dansmith | mriedem: it would have just been ordering reasons | |
| 18:54:59 | dansmith | mriedem: note the rising tail at present too | |
| 19:18:31 | mriedem | alright i'm just going to restack | |
| 19:18:32 | mriedem | nuts to this | |
| 19:28:40 | mriedem | dansmith: jaypipes: bauzas: https://review.openstack.org/#/c/498947/6 | |
| 19:28:45 | mriedem | that test_servers thing is wrong | |
| 19:29:19 | openstackgerrit | Matthew Treinish proposed openstack/nova master: Add slowest command to tox.ini https://review.openstack.org/507657 | |
| 19:29:21 | mtreinish | dansmith: ^^^ | |
| 19:29:29 | mriedem | there are 2 tests for failures during evacaute on the dest | |
| 19:29:38 | mriedem | 1. test_evacuate_claim_on_dest_fails - that is testing when the claim fails with ComputeResourcesUnavailable | |
| 19:29:57 | mriedem | 2. test_evacuate_rebuild_on_dest_fails - that is testing when the claim is successful but the driver.rebuild method raises some exception | |
| 19:29:57 | dansmith | mtreinish: sweet | |
| 19:30:00 | jaypipes | mriedem: sorry, I disagree with you. | |
| 19:30:19 | mriedem | i wrote those tests | |
| 19:30:29 | mriedem | so please explain how i'm wrong that they are now made redundant in that change | |
| 19:30:36 | jaypipes | mriedem: that test raising TestingException was not useful. Because TestingException isn't what is ever raised by any code. | |
| 19:30:54 | mriedem | it's simulating the virt driver raising the error during rebuild | |
| 19:31:00 | mriedem | AFTER the successful claim | |
| 19:31:10 | mriedem | it could be ProcessExecutionError | |
| 19:31:12 | mriedem | from driver.spawn() | |
| 19:31:14 | mriedem | if you like | |
| 19:31:44 | mriedem | these 2 tests are testing very specific failures | |
| 19:31:45 | dansmith | mriedem: right but we don't run the claim teardown code in that case | |
| 19:32:00 | mriedem | dansmith: correct, which is why we run the allocation cleanup manually | |
| 19:32:03 | mriedem | and that's what that is testing | |
| 19:32:32 | mriedem | the test you changed isn't meant to test drop_move_claim | |
| 19:32:35 | mriedem | the docstring explains that | |
| 19:32:44 | jaypipes | mriedem: if the point of the test (as is in that docstring) is to ensure allocations are cleaned up after a failed rebuild, then the test should raise the exception that would be raised *after* a claim has been made for the new resources. | |
| 19:32:55 | dansmith | jaypipes: he's saying another one does that | |
| 19:33:14 | mriedem | jaypipes: you realize the virt drivers can raise any kinds of crazy shit right? | |
| 19:33:15 | dansmith | mriedem: so in this case you want the test to validate that the allocations _don't_ get cleaned up is that right? | |
| 19:33:31 | jaypipes | mriedem: Matt, I'm trying to be civil. | |
| 19:34:38 | mriedem | https://review.openstack.org/#/c/499877/ | |
| 19:35:22 | dansmith | this is what it's testing: https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2800-L2827 | |
| 19:35:23 | dansmith | the except exception case of that | |
| 19:35:24 | mriedem | so ^ is testing that drop_move_claim removes the allocation when the claim was successful but the virt driver raised some exception | |
| 19:35:37 | jaypipes | mriedem: OK, I see that now. | |
| 19:36:18 | mriedem | https://review.openstack.org/#/c/499874/ added the other test | |
| 19:36:54 | mriedem | that was a recreate test for a bug | |
| 19:36:55 | mriedem | fixed in https://review.openstack.org/#/c/499878/ | |
| 19:37:38 | dansmith | mriedem: we get it | |
| 19:37:51 | dansmith | mriedem: can you answer my question above about what you want it to do? | |
| 19:39:32 | mriedem | the test should go back to whatever it was testing | |
| 19:39:43 | mriedem | which is the case that the claim passes, but the virt driver raises | |
| 19:39:51 | mriedem | so we'd remove the allocation via drop_move_claim before | |
| 19:40:02 | dansmith | right, but you assert some behavior that happens inside drop_move_claim | |
| 19:40:07 | dansmith | which no longer happens | |
| 19:40:35 | mriedem | then that drop_move_claim behavior has to be replayed elsewhere i guess | |
| 19:40:37 | jaypipes | what if the it's a same-host rebuild? :( | |
| 19:40:45 | mriedem | there is no claim for a same host rebuild | |
| 19:40:49 | jaypipes | k | |
| 19:40:54 | mriedem | so you wouldn't hit ComputeResourcesUnavailable | |