| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 17:27:00 | openstackgerrit | Moshe Levi proposed openstack/nova master: Don't overwrite binding-profile https://review.openstack.org/505613 | |
| 17:32:29 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404 | |
| 17:38:24 | mriedem | dansmith: aha, i think i'm hitting issues in devstack where placement isn't getting cleaned up for instances that get 'local' deleted in the api | |
| 17:38:45 | mriedem | not totally sure yet, but failing to burst 500 new instances, hitting NoValidHost | |
| 17:38:53 | mriedem | and i assume it's placement b/c it's not the scheduler filters | |
| 17:39:17 | dansmith | mriedem: and why do you have locally-deleted instances for this test? | |
| 17:39:17 | melwitt | for local deletes, allocations aren't cleaned up till the compute host heals it | |
| 17:39:24 | dansmith | right, what melwitt said | |
| 17:39:27 | mriedem | mysql> select count(id) from consumers; | |
| 17:39:27 | mriedem | | count(id) | | |
| 17:39:27 | mriedem | | 2002 | | |
| 17:39:27 | mriedem | +-----------+ | |
| 17:39:28 | mriedem | 1 row in set (0.01 sec) | |
| 17:40:05 | mriedem | melwitt: there is no compute for these | |
| 17:40:07 | mriedem | they failed during scheduling | |
| 17:40:22 | mriedem | although yeah why would placement have allocations for these... | |
| 17:40:23 | mriedem | wtf | |
| 17:40:37 | mriedem | stack@devstack:~$ nova list | grep -c ERROR | |
| 17:40:37 | mriedem | 1000 | |
| 17:40:39 | melwitt | oh, hm | |
| 17:40:48 | mriedem | so i've got 1000 instances in ERROR state, and 2002 consumers in the api db | |
| 17:41:20 | melwitt | allocations are written at claim time? | |
| 17:41:27 | mriedem | from the scheduler yeah | |
| 17:42:10 | melwitt | so that would explain the ones you do have. but I guess your point is why are there more allocation consumers than non error instances | |
| 17:42:49 | mriedem | that's because i've deleted 1000 over time | |
| 17:43:17 | mriedem | i was hitting messaging timeouts between conductor and the scheduler earlier today, so had 500 in error which i needed to be active, so deleted all of those, restarted conductor and scheduler, and was able to create a single instance | |
| 17:43:22 | mriedem | so tried with 500 more again | |
| 17:43:26 | mriedem | and hit novalidhost on all of those | |
| 17:44:18 | openstackgerrit | Merged openstack/nova master: cleanup test-requirements https://review.openstack.org/507063 | |
| 17:55:32 | openstackgerrit | Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638 | |
| 17:55:45 | dansmith | jaypipes: cdent: ^ quick stab at the live migrate version of this | |
| 17:56:02 | dansmith | it's probably rough at this point, but worth a look I think | |
| 18:32:22 | cdent | dansmith: haven’t had a chance to give it a proper look, but saw a weird when skimming the live migrate thing | |
| 18:32:48 | dansmith | lol | |
| 18:33:04 | dansmith | it's returning True-ish which is what I wanted for the functional tests | |
| 18:33:09 | dansmith | so.. working as designed? :) | |
| 18:33:26 | dansmith | s/returning/being/ | |
| 18:34:36 | openstackgerrit | Merged openstack/nova master: Set the Pike release version for scheduler RPC https://review.openstack.org/507245 | |
| 18:34:51 | cdent | go python! | |
| 18:38:00 | mriedem | wtf, so i can't create multiple instances, i get novalidhost, but i can create one at a time | |
| 18:38:21 | melwitt | are you using multi-create? | |
| 18:38:27 | mriedem | yeah | |
| 18:38:30 | mriedem | wasn't a problem yesterday | |
| 18:38:37 | melwitt | oh | |
| 18:38:41 | mriedem | but i had a bit of a cleaner env yesterday | |
| 18:39:46 | melwitt | multi-create will reject you if any one of min_count can't be accommodated. so one at a time would work if you're in that situation, if some/most of them fit | |
| 18:40:22 | mriedem | yesterday i created 100, then like 200, then 500 more or something | |
| 18:40:24 | mriedem | eventually got to 1000 | |
| 18:41:49 | mriedem | i can just restack this env, but it makes me worry that we aren't properly cleaning up allocations somewhere | |
| 18:42:37 | melwitt | yeah | |
| 18:43:15 | melwitt | did you say yesterday you don't have computes, or something like that? I just wonder what happens with FakeDriver, if it somehow doesn't call the healing allocations code | |
| 18:43:38 | mriedem | we don't heal since pike | |
| 18:43:51 | mriedem | if you don't have computes < pike, we don't heal | |
| 18:44:04 | mriedem | this is just a single compute, single node devstack | |
| 18:44:11 | mriedem | with the fake driver and noop quota | |
| 18:44:22 | melwitt | oh wait, sorry I was thinking of local delete | |
| 18:44:53 | melwitt | I was trying to think if with FakeDriver, does the code that deletes allocations when an instance is deleted, run | |
| 18:45:02 | melwitt | or if that even matters | |
| 18:45:17 | mriedem | the compute manager cleans up allocations when an instance is deleted | |
| 18:45:34 | dansmith | mriedem: we heal for deletes | |
| 18:45:44 | dansmith | mriedem: did you archive them after local delete before you started up? | |
| 18:46:01 | mriedem | i've been archiving yeah | |
| 18:46:05 | melwitt | I'm not sure whether he had local deletes | |
| 18:46:06 | dansmith | so that's why | |
| 18:46:14 | dansmith | I thought he did | |
| 18:46:34 | melwitt | I thought he did too but I'm getting a little confused | |
| 18:46:53 | mriedem | yesterday i didn't have any instances in ERROR state, so they were all in the cell | |
| 18:47:04 | mriedem | i deleted all of those and then archived cell0 and cell1 | |
| 18:47:21 | dansmith | deleted them locally? | |
| 18:47:21 | mriedem | today i've been trying to get 500 ERROR during scheduling, and 500 ACTIVE | |
| 18:47:24 | mriedem | based on the flavor i use | |
| 18:47:30 | mriedem | dansmith: deleted via the api | |
| 18:47:37 | dansmith | mriedem: with compute down or no? | |
| 18:47:38 | mriedem | remember me complaining about how long that was taking yesterday? | |
| 18:47:39 | mriedem | no | |
| 18:47:48 | mriedem | took 2+ hours to delete 1000 ACTIVE instances | |
| 18:48:03 | dansmith | I do, but I didn't remember all your details | |
| 18:48:20 | mriedem | yeah i basically trying to get back to clean state before starting today | |
| 18:48:33 | mriedem | so was archiving the db's last night | |
| 18:49:19 | mriedem | btw, bauzas pointed this out before, but we log this way too many times | |
| 18:49:19 | mriedem | Sep 26 18:44:37 devstack nova-compute[30351]: DEBUG nova.compute.resource_tracker [None req-992d494e-d328-4204-bcfe-80d926cf0a65 demo demo] We're on a Pike compute host in a deployment with all Pike compute hosts. Skipping auto-correction of allocations. {{(pid=30351) _update_usage_from_instance /opt/stack/nova/nova/compute/resource_tracker.py:1071}} | |
| 18:52:10 | dansmith | mriedem: unrelated, see this: http://status.openstack.org/openstack-health/#/test/nova.tests.functional.test_servers.ServersTestV219.test_description_errors?duration=P3M | |
| 18:52:34 | dansmith | mriedem: I think this test is occasionally taking up to 240s locally when it should be about 8s | |
| 18:53:00 | mriedem | jesus | |
| 18:53:02 | dansmith | and I think it's because it creates a server that it never cleans up and then abruptly exits where we take down conductor before the compute service finishes waiting on a call or something | |
| 18:53:16 | dansmith | so I have a patch to just make it clean up the server and I _think_ it's working | |
| 18:53:32 | mriedem | the one weird spike in august is, weird | |
| 18:53:45 | mriedem | https://bugs.launchpad.net/nova/+bug/1719714 | |
| 18:53:46 | openstack | Launchpad bug 1719714 in OpenStack Compute (nova) "Excessive logging of "We're on a Pike compute host in a deployment with all Pike compute hosts."" [Medium,Confirmed] | |
| 18:54:04 | dansmith | mriedem: it would have just been ordering reasons | |
| 18:54:59 | dansmith | mriedem: note the rising tail at present too | |
| 19:18:31 | mriedem | alright i'm just going to restack | |
| 19:18:32 | mriedem | nuts to this | |
| 19:28:40 | mriedem | dansmith: jaypipes: bauzas: https://review.openstack.org/#/c/498947/6 | |
| 19:28:45 | mriedem | that test_servers thing is wrong | |
| 19:29:19 | openstackgerrit | Matthew Treinish proposed openstack/nova master: Add slowest command to tox.ini https://review.openstack.org/507657 | |
| 19:29:21 | mtreinish | dansmith: ^^^ | |
| 19:29:29 | mriedem | there are 2 tests for failures during evacaute on the dest | |
| 19:29:38 | mriedem | 1. test_evacuate_claim_on_dest_fails - that is testing when the claim fails with ComputeResourcesUnavailable | |
| 19:29:57 | mriedem | 2. test_evacuate_rebuild_on_dest_fails - that is testing when the claim is successful but the driver.rebuild method raises some exception | |