Earlier  
Posted Nick Remark
#openstack-nova - 2017-10-06
19:23:19 ronlund superdan: well i think if instance.host == None we assume the allocations are already gone
19:23:24 ronlund either it failed to schedule,
19:23:27 ronlund or it was shelved offloaded
19:23:36 ronlund and we remove allocations when shelve offloading
19:23:36 superdan right, my point
19:23:41 superdan we call to scheduler, timeout,
19:23:44 superdan scheduler has made allocations
19:23:55 superdan we just delete from db because it never scheduled
19:24:02 ronlund instance goes to error
19:24:04 ronlund but has allocations
19:24:13 superdan yes
19:24:17 ronlund heh yeah
19:26:05 ronlund checking for something like that on every delete kind of sucks if it's a super edge case
19:26:26 superdan but no healing, so.. leaking capacity will anger people and rightly so :)
19:26:38 ronlund right
19:26:54 ronlund plus a delete request for allocations that never existing should be fast
19:26:58 ronlund *existed
19:27:02 superdan yes
19:27:47 ronlund i know huawei customers love nfv, i need to see what their instance quota limit is quick... :)
19:29:48 ronlund ha
19:29:49 ronlund @utils.retry_select_destinations
19:29:53 ronlund that's what's causing the retry
19:29:54 ronlund it's by design
19:29:59 ronlund it's not oslo.messaging, it's nova
19:31:02 ronlund https://github.com/openstack/nova/blob/353db2d1932965b6502e002b8be510440ff529c0/nova/scheduler/utils.py#L599
19:31:38 superdan yeah
19:31:49 ronlund so yeah, now that we're doing claims in the scheduler, that seems like a bad idea...
19:32:14 ronlund it does it up to max_attempts-1, so by default 2 retries
19:32:27 superdan that doesn't fix the allocation leak, mind you,
19:32:43 superdan but yeah, seems like if you fail talking to it, you're just going to hurt things by adding to the load with a retry
19:33:11 ronlund i wonder if we double up the 2nd allocation request for the same consumer
19:33:25 ronlund maybe not if there is only 1 rp uuid in the request
19:33:40 ronlund note the 2nd time through the scheduler on the retry, we could likely target a completely different host :)
19:33:49 ronlund thus totally fucking up things for everything
19:33:52 superdan yeah
19:34:08 ronlund huh, well this is fun
19:50:57 leakypipes fried_rice: putting: "blueprint: XXXX" does the same thing.
19:51:27 fried_rice leakypipes Coolio. Is there a Source Of Truth for these taggy things?
19:54:31 leakypipes fried_rice: meh, https://wiki.openstack.org/wiki/GitCommitMessages
19:54:45 leakypipes fried_rice: but it only mentions using Implements: blueprint XXX
19:54:50 fried_rice mm
19:55:23 leakypipes fried_rice: that's not necessary though. the word "blueprint" followed by a tag-like thing is all that's needed to link the patch with the blueprint on Launchpad.
19:55:42 fried_rice leakypipes Including having whatever bot add the URL to the whiteboard on the bp?
19:55:58 fried_rice Cause that seems to be a thing.
19:56:06 leakypipes fried_rice: correct, that's what I mean.
19:56:22 fried_rice k, thought you were just talking about gerrit turning it into a nice hyperlink to the LP page.
19:56:41 fried_rice Anyway, I dig it.
19:58:07 ronlund bp also works i think
19:58:12 ronlund maybe not
19:59:56 mtreinish there is a lot of backscroll, but I couldn't see a link to what I should be looking at
19:59:56 mtreinish ronlund, jgwentworth: do you have a link to the thing you're seeing?
20:00:11 jgwentworth sec
20:00:43 jgwentworth mtreinish: this is happening on stable/ocata and stable/pike only http://logs.openstack.org/39/509439/1/check/gate-nova-python27-ubuntu-xenial/e456c8f/testr_results.html.gz
20:01:05 jgwentworth I think it's just a display issue, showing the os profiler result instead of the unit tests result. in the console you can see that both ran
20:01:23 mtreinish jgwentworth: ok, yeah that's because you have 2 test runs in the tox command
20:01:43 mtreinish for the post processing to generate that we run testr last --subunit and pipe that into subunit2html to generate that page
20:01:46 jgwentworth it shows the right thing on master and stable/newton for some reason even though we have 2 runs
20:01:59 mtreinish but testr doesn't let you combine the results
20:02:10 mtreinish so it's just showing the results from the second one
20:02:12 jgwentworth yeah, that's what sean-k-mooney was saying
20:02:37 jgwentworth well, I think testr is showing the first one. os profiler always runs last
20:02:39 mtreinish it works on master because stestr has a --combine flag that treats the 2 commands as a single run
20:03:23 ronlund gah 2017-10-06 04:30:20.654 | /opt/stack/new/devstack/inc/meta-config: line 209: /opt/stack/new/devstack/.localrc.auto: Permission denied
20:03:41 jgwentworth mtreinish: oh, on stable/newton we're not running the os profiler thing
20:04:00 jgwentworth that's why that one shows up correctly
20:04:28 mtreinish yep
20:04:53 mtreinish this was the same thing I was seeing on openstack-health, which is why I pushed https://review.openstack.org/#/c/501842/ before doing the stestr migration
20:04:59 jgwentworth weird. it seems like this would always have been happening before stestr but I could have sworn I had seen full lists of the unit tests on a pass run prior to stestr
20:05:32 jgwentworth maybe I dreamed it
20:05:37 mtreinish jgwentworth: this was a longstanding issue, I just don't think anyone noticed it
20:05:45 mtreinish before the osprofiler tests were added it wasn't an issue
20:05:51 jgwentworth right
20:06:15 jgwentworth oh well, sorry for the noise. I thought something had changed
20:06:54 mtreinish no worries, I'm just glad other people are looking at this stuff :)
20:08:04 mtreinish jgwentworth: fwiw, if you want to fix that output you could probably backport 501842
20:08:36 mtreinish or just rip the osprofiler test out on the stable branches, tbh I'm not sure what value it adds
20:10:03 jgwentworth yeah, I like running profiler last locally so tests I'm working on run and fail first without having to wait for profiler
20:10:08 openstackgerrit Matt Riedemann proposed openstack/nova stable/pike: Set regex flag on ostestr command for osprofiler tests https://review.openstack.org/510226
20:10:11 ronlund let's see
20:29:30 openstackgerrit Ildiko Vancsa proposed openstack/nova master: update live migration to use v3 cinder api https://review.openstack.org/463987
21:07:39 superdan ronlund: I think those two patches are good to go
21:08:06 superdan ronlund: cool if you want to wait for andreykurilin to run a test, but he confirmed on one of the earlier iterations, I have a repro test in there, etc
21:09:00 andreykurilin superdan: Does the new revision ready for testing? I can recheck the patch now
21:09:21 superdan andreykurilin: oh yeah, I poked you earlier but no response :)
21:09:32 andreykurilin oh.
21:09:38 superdan andreykurilin: would be awesome if you could, and maybe point me at the review so I can watch?
21:11:48 andreykurilin superdan: heh, sure. https://review.openstack.org/#/c/510144 gate-rally-dsvm-neutron-existing-users-rally job. grep tag-to-search to see debug messages which actually shows that everything works or not
21:13:23 superdan andreykurilin: doesn't it just fail if it fails?
21:15:27 superdan andreykurilin: I don't see that job in the first run of that patch...
21:16:14 ronlund ok, writing up a draft spec for discussion on this max_count rate limiting thing
21:16:21 andreykurilin superdan: no :( I tried to make this change as soon as possible, so i have not removed a hack to limit the loop. I'll do it now. the first results were not published, it looks like you posted a new revision before the whole jobs were finished
21:16:37 superdan ah okay
21:19:33 andreykurilin superdan: so originally the listing doesn't fail. it was just a inf loop. To add debug messages which will not flood the log file, I added a limit in the loop(10). That is why it stopped failing. Now I resubmitted a patch which fails in case of 10 iterations of the loop. so you do not need to check the logs
21:20:11 superdan andreykurilin: if len(log) == 10g -> fail? :)
21:20:12 superdan andreykurilin: awesome thanks
21:20:21 andreykurilin ha
21:29:00 openstackgerrit Matt Riedemann proposed openstack/nova-specs master: Limit instance create max_count (spec) https://review.openstack.org/510235
21:29:44 ronlund superdan: cfriesen: leakypipes: i know it's late so probably just ignore until next week, but ^ has some thoughts on the concurrent scheduling issue
21:30:10 superdan ack
21:31:11 mriedem now lemme take a looksie at these crazy patches

Earlier   Later