Earlier  
Posted Nick Remark
#openstack-nova - 2017-10-06
19:04:20 ronlund {"server": {"name": "test-alloc-conflict", "imageRef": "947823a9-8c15-420a-95d8-f35cd2f024b9", "flavorRef": "1", "max_count": 1000, "min_count": 1000, "networks": "none"}}
19:04:50 ronlund i see a successful allocation for the failed instance 2 seconds after the final 409 retry failure
19:04:51 ronlund http://logs.openstack.org/18/507918/6/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/d4f175d/logs/screen-placement-api.txt.gz#_Oct_04_18_34_13_854776
19:05:40 ronlund wth, why would we have 3 greenthreads?
19:08:16 openstackgerrit Jay Pipes proposed openstack/nova master: rp: de-ORM ResourceProvider.get_by_uuid() https://review.openstack.org/509025
19:08:17 openstackgerrit Jay Pipes proposed openstack/nova master: rp: Remove RP.get_traits() method https://review.openstack.org/509027
19:08:17 openstackgerrit Jay Pipes proposed openstack/nova master: rp: Move RP._get|set_aggregates() to module scope https://review.openstack.org/509026
19:08:18 openstackgerrit Jay Pipes proposed openstack/nova master: rp: remove CRUD operations on Inventory class https://review.openstack.org/509029
19:08:18 openstackgerrit Jay Pipes proposed openstack/nova master: rp: move RP._set_traits() to module scope https://review.openstack.org/509028
19:08:19 openstackgerrit Jay Pipes proposed openstack/nova master: rp: remove dead code in Allocation._create_in_db() https://review.openstack.org/509031
19:08:19 openstackgerrit Jay Pipes proposed openstack/nova master: rp: streamline InventoryList.get_all_by_rp_uuid() https://review.openstack.org/509030
19:08:20 openstackgerrit Jay Pipes proposed openstack/nova master: rp: remove ability to delete 1 allocation record https://review.openstack.org/509032
19:08:21 openstackgerrit Jay Pipes proposed openstack/nova master: rp: remove _HasAResourceProvider mixin https://review.openstack.org/509036
19:08:21 openstackgerrit Jay Pipes proposed openstack/nova master: rp: rework AllocList.get_all_by_consumer_id() https://review.openstack.org/509035
19:08:21 openstackgerrit Jay Pipes proposed openstack/nova master: rp: fix up AllocList.get_by_resource_provider_uuid https://review.openstack.org/509033
19:08:22 openstackgerrit Jay Pipes proposed openstack/nova master: rp: break functions out of _set_traits() https://review.openstack.org/509908
19:09:02 ronlund http://logs.openstack.org/18/507918/6/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/d4f175d/logs/screen-n-super-cond.txt.gz#_Oct_04_18_32_07_052398
19:09:02 ronlund aha
19:09:21 ronlund select_destinations took too long, so oslo.messaging retried
19:09:53 ronlund well that explains that i think
19:09:58 ronlund i'm going to bump up the rpc timeout
19:10:18 cdent that doesn’t sound safe
19:10:47 cdent (dishwasher not done yet)
19:11:14 ronlund huh, given we don't rate limit min_count at all,
19:11:23 ronlund you could totally bork some stuff up with making a huge request
19:11:35 ronlund if the server's rpc timeout isn't configured high enough to handle it
19:11:45 cfriesen ronlund: I think we check quota fairly early, no?
19:11:50 ronlund this isn't quota
19:12:07 ronlund although, sure you wouldn't hit this for real if you had default quota of 10
19:12:11 sean-k-mooney jgwentworth: https://github.com/openstack-infra/project-config/blob/dbdef981de7cb56e9cd44514a41102270bfc9bac/jenkins/scripts/run-tox.sh#L32-L48 this is why its failing we only grab the last result dir
19:12:12 ronlund i've disabled quota
19:12:24 ronlund if you did allow a tenant 1000 instances to burst at once,
19:12:33 ronlund then you are going to have to deal with big ass rpc timeouts
19:13:03 leakypipes ronlund: there's a crapload of timeouts on the MQ in that log file...
19:14:05 ronlund yeah i know
19:14:14 ronlund and by default oslo.messaging retries twice
19:14:51 superdan so it just adds to the load?
19:15:03 superdan because it's starting extra scheduling runs for the same set of stuff?
19:15:14 superdan seems like we'd be breaking hard on that already anyway, pre-placement
19:17:02 ronlund yeah this is definitely user error on my part :)
19:17:21 ronlund maybe it's a decent simulator of sorts, but probably not
19:17:42 ronlund "how to dos your devstack"
19:18:00 superdan user error why?
19:18:08 superdan just because all the limits are removed?
19:18:12 ronlund yeah
19:18:35 ronlund we could definitely hammer scheduler/placement with concurrent requests,
19:18:38 cdent in a perfect world it would at least fail gracefully rather than confusedly
19:18:45 ronlund but within a single tenant, default quota is 10 so that'd be your max
19:18:56 superdan well, we probably really should never retry a call to the scheduler like this after a timeout, I'm thinking
19:19:00 ronlund well, the good news is it totally does the allocation cleanup properly
19:19:07 ronlund and everything is put into error state and shoved in cell0
19:19:24 superdan because you could hit that timeout for other reasons
19:19:48 superdan well, I guess it has already created the instance records, so maybe not a huge deal I guess,
19:20:16 superdan I expect there is a case where you could do a boot, fail to hear from scheduler, never send boots to compute,
19:20:21 superdan but the scheduler made allocations
19:20:48 superdan I guess we just live with that and assume they're cleaned up at delete, but technically it's holding space for those dead ones
19:21:06 superdan actually, maybe we wouldn't clean up allocations on delete in that case since instance.host=None?
19:23:19 ronlund superdan: well i think if instance.host == None we assume the allocations are already gone
19:23:24 ronlund either it failed to schedule,
19:23:27 ronlund or it was shelved offloaded
19:23:36 ronlund and we remove allocations when shelve offloading
19:23:36 superdan right, my point
19:23:41 superdan we call to scheduler, timeout,
19:23:44 superdan scheduler has made allocations
19:23:55 superdan we just delete from db because it never scheduled
19:24:02 ronlund instance goes to error
19:24:04 ronlund but has allocations
19:24:13 superdan yes
19:24:17 ronlund heh yeah
19:26:05 ronlund checking for something like that on every delete kind of sucks if it's a super edge case
19:26:26 superdan but no healing, so.. leaking capacity will anger people and rightly so :)
19:26:38 ronlund right
19:26:54 ronlund plus a delete request for allocations that never existing should be fast
19:26:58 ronlund *existed
19:27:02 superdan yes
19:27:47 ronlund i know huawei customers love nfv, i need to see what their instance quota limit is quick... :)
19:29:48 ronlund ha
19:29:49 ronlund @utils.retry_select_destinations
19:29:53 ronlund that's what's causing the retry
19:29:54 ronlund it's by design
19:29:59 ronlund it's not oslo.messaging, it's nova
19:31:02 ronlund https://github.com/openstack/nova/blob/353db2d1932965b6502e002b8be510440ff529c0/nova/scheduler/utils.py#L599
19:31:38 superdan yeah
19:31:49 ronlund so yeah, now that we're doing claims in the scheduler, that seems like a bad idea...
19:32:14 ronlund it does it up to max_attempts-1, so by default 2 retries
19:32:27 superdan that doesn't fix the allocation leak, mind you,
19:32:43 superdan but yeah, seems like if you fail talking to it, you're just going to hurt things by adding to the load with a retry
19:33:11 ronlund i wonder if we double up the 2nd allocation request for the same consumer
19:33:25 ronlund maybe not if there is only 1 rp uuid in the request
19:33:40 ronlund note the 2nd time through the scheduler on the retry, we could likely target a completely different host :)
19:33:49 ronlund thus totally fucking up things for everything
19:33:52 superdan yeah
19:34:08 ronlund huh, well this is fun
19:50:57 leakypipes fried_rice: putting: "blueprint: XXXX" does the same thing.
19:51:27 fried_rice leakypipes Coolio. Is there a Source Of Truth for these taggy things?
19:54:31 leakypipes fried_rice: meh, https://wiki.openstack.org/wiki/GitCommitMessages
19:54:45 leakypipes fried_rice: but it only mentions using Implements: blueprint XXX
19:54:50 fried_rice mm
19:55:23 leakypipes fried_rice: that's not necessary though. the word "blueprint" followed by a tag-like thing is all that's needed to link the patch with the blueprint on Launchpad.
19:55:42 fried_rice leakypipes Including having whatever bot add the URL to the whiteboard on the bp?

Earlier   Later