| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-06 | |||
| 19:29:48 | ronlund | ha | |
| 19:29:49 | ronlund | @utils.retry_select_destinations | |
| 19:29:53 | ronlund | that's what's causing the retry | |
| 19:29:54 | ronlund | it's by design | |
| 19:29:59 | ronlund | it's not oslo.messaging, it's nova | |
| 19:31:02 | ronlund | https://github.com/openstack/nova/blob/353db2d1932965b6502e002b8be510440ff529c0/nova/scheduler/utils.py#L599 | |
| 19:31:38 | superdan | yeah | |
| 19:31:49 | ronlund | so yeah, now that we're doing claims in the scheduler, that seems like a bad idea... | |
| 19:32:14 | ronlund | it does it up to max_attempts-1, so by default 2 retries | |
| 19:32:27 | superdan | that doesn't fix the allocation leak, mind you, | |
| 19:32:43 | superdan | but yeah, seems like if you fail talking to it, you're just going to hurt things by adding to the load with a retry | |
| 19:33:11 | ronlund | i wonder if we double up the 2nd allocation request for the same consumer | |
| 19:33:25 | ronlund | maybe not if there is only 1 rp uuid in the request | |
| 19:33:40 | ronlund | note the 2nd time through the scheduler on the retry, we could likely target a completely different host :) | |
| 19:33:49 | ronlund | thus totally fucking up things for everything | |
| 19:33:52 | superdan | yeah | |
| 19:34:08 | ronlund | huh, well this is fun | |
| 19:50:57 | leakypipes | fried_rice: putting: "blueprint: XXXX" does the same thing. | |
| 19:51:27 | fried_rice | leakypipes Coolio. Is there a Source Of Truth for these taggy things? | |
| 19:54:31 | leakypipes | fried_rice: meh, https://wiki.openstack.org/wiki/GitCommitMessages | |
| 19:54:45 | leakypipes | fried_rice: but it only mentions using Implements: blueprint XXX | |
| 19:54:50 | fried_rice | mm | |
| 19:55:23 | leakypipes | fried_rice: that's not necessary though. the word "blueprint" followed by a tag-like thing is all that's needed to link the patch with the blueprint on Launchpad. | |
| 19:55:42 | fried_rice | leakypipes Including having whatever bot add the URL to the whiteboard on the bp? | |
| 19:55:58 | fried_rice | Cause that seems to be a thing. | |
| 19:56:06 | leakypipes | fried_rice: correct, that's what I mean. | |
| 19:56:22 | fried_rice | k, thought you were just talking about gerrit turning it into a nice hyperlink to the LP page. | |
| 19:56:41 | fried_rice | Anyway, I dig it. | |
| 19:58:07 | ronlund | bp also works i think | |
| 19:58:12 | ronlund | maybe not | |
| 19:59:56 | mtreinish | ronlund, jgwentworth: do you have a link to the thing you're seeing? | |
| 19:59:56 | mtreinish | there is a lot of backscroll, but I couldn't see a link to what I should be looking at | |
| 20:00:11 | jgwentworth | sec | |
| 20:00:43 | jgwentworth | mtreinish: this is happening on stable/ocata and stable/pike only http://logs.openstack.org/39/509439/1/check/gate-nova-python27-ubuntu-xenial/e456c8f/testr_results.html.gz | |
| 20:01:05 | jgwentworth | I think it's just a display issue, showing the os profiler result instead of the unit tests result. in the console you can see that both ran | |
| 20:01:23 | mtreinish | jgwentworth: ok, yeah that's because you have 2 test runs in the tox command | |
| 20:01:43 | mtreinish | for the post processing to generate that we run testr last --subunit and pipe that into subunit2html to generate that page | |
| 20:01:46 | jgwentworth | it shows the right thing on master and stable/newton for some reason even though we have 2 runs | |
| 20:01:59 | mtreinish | but testr doesn't let you combine the results | |
| 20:02:10 | mtreinish | so it's just showing the results from the second one | |
| 20:02:12 | jgwentworth | yeah, that's what sean-k-mooney was saying | |
| 20:02:37 | jgwentworth | well, I think testr is showing the first one. os profiler always runs last | |
| 20:02:39 | mtreinish | it works on master because stestr has a --combine flag that treats the 2 commands as a single run | |
| 20:03:23 | ronlund | gah 2017-10-06 04:30:20.654 | /opt/stack/new/devstack/inc/meta-config: line 209: /opt/stack/new/devstack/.localrc.auto: Permission denied | |
| 20:03:41 | jgwentworth | mtreinish: oh, on stable/newton we're not running the os profiler thing | |
| 20:04:00 | jgwentworth | that's why that one shows up correctly | |
| 20:04:28 | mtreinish | yep | |
| 20:04:53 | mtreinish | this was the same thing I was seeing on openstack-health, which is why I pushed https://review.openstack.org/#/c/501842/ before doing the stestr migration | |
| 20:04:59 | jgwentworth | weird. it seems like this would always have been happening before stestr but I could have sworn I had seen full lists of the unit tests on a pass run prior to stestr | |
| 20:05:32 | jgwentworth | maybe I dreamed it | |
| 20:05:37 | mtreinish | jgwentworth: this was a longstanding issue, I just don't think anyone noticed it | |
| 20:05:45 | mtreinish | before the osprofiler tests were added it wasn't an issue | |
| 20:05:51 | jgwentworth | right | |
| 20:06:15 | jgwentworth | oh well, sorry for the noise. I thought something had changed | |
| 20:06:54 | mtreinish | no worries, I'm just glad other people are looking at this stuff :) | |
| 20:08:04 | mtreinish | jgwentworth: fwiw, if you want to fix that output you could probably backport 501842 | |
| 20:08:36 | mtreinish | or just rip the osprofiler test out on the stable branches, tbh I'm not sure what value it adds | |
| 20:10:03 | jgwentworth | yeah, I like running profiler last locally so tests I'm working on run and fail first without having to wait for profiler | |
| 20:10:08 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Set regex flag on ostestr command for osprofiler tests https://review.openstack.org/510226 | |
| 20:10:11 | ronlund | let's see | |
| 20:29:30 | openstackgerrit | Ildiko Vancsa proposed openstack/nova master: update live migration to use v3 cinder api https://review.openstack.org/463987 | |
| 21:07:39 | superdan | ronlund: I think those two patches are good to go | |
| 21:08:06 | superdan | ronlund: cool if you want to wait for andreykurilin to run a test, but he confirmed on one of the earlier iterations, I have a repro test in there, etc | |
| 21:09:00 | andreykurilin | superdan: Does the new revision ready for testing? I can recheck the patch now | |
| 21:09:21 | superdan | andreykurilin: oh yeah, I poked you earlier but no response :) | |
| 21:09:32 | andreykurilin | oh. | |
| 21:09:38 | superdan | andreykurilin: would be awesome if you could, and maybe point me at the review so I can watch? | |
| 21:11:48 | andreykurilin | superdan: heh, sure. https://review.openstack.org/#/c/510144 gate-rally-dsvm-neutron-existing-users-rally job. grep tag-to-search to see debug messages which actually shows that everything works or not | |
| 21:13:23 | superdan | andreykurilin: doesn't it just fail if it fails? | |
| 21:15:27 | superdan | andreykurilin: I don't see that job in the first run of that patch... | |
| 21:16:14 | ronlund | ok, writing up a draft spec for discussion on this max_count rate limiting thing | |
| 21:16:21 | andreykurilin | superdan: no :( I tried to make this change as soon as possible, so i have not removed a hack to limit the loop. I'll do it now. the first results were not published, it looks like you posted a new revision before the whole jobs were finished | |
| 21:16:37 | superdan | ah okay | |
| 21:19:33 | andreykurilin | superdan: so originally the listing doesn't fail. it was just a inf loop. To add debug messages which will not flood the log file, I added a limit in the loop(10). That is why it stopped failing. Now I resubmitted a patch which fails in case of 10 iterations of the loop. so you do not need to check the logs | |
| 21:20:11 | superdan | andreykurilin: if len(log) == 10g -> fail? :) | |
| 21:20:12 | superdan | andreykurilin: awesome thanks | |
| 21:20:21 | andreykurilin | ha | |
| 21:29:00 | openstackgerrit | Matt Riedemann proposed openstack/nova-specs master: Limit instance create max_count (spec) https://review.openstack.org/510235 | |
| 21:29:44 | ronlund | superdan: cfriesen: leakypipes: i know it's late so probably just ignore until next week, but ^ has some thoughts on the concurrent scheduling issue | |
| 21:30:10 | superdan | ack | |
| 21:31:11 | mriedem | now lemme take a looksie at these crazy patches | |
| 21:44:07 | mriedem | wow rally runs a lot of jobs | |
| 21:44:29 | mriedem | that's got to be a challenge to both tempest and trove | |
| 21:45:40 | mtreinish | mriedem: I think htat's more than tempest | |
| 21:56:10 | cfriesen | mriedem: did we ever fix the issues with multi-boot where scheduling more than min_count but less than max_count would result in instances in the "error" state? | |
| 21:58:40 | mriedem | cfriesen: i'd need more details than that | |
| 21:59:04 | mriedem | why would that put them in error state? | |
| 21:59:31 | mriedem | if i request min=5 max=10 but there are only 7 ports available in my port quota, the api changes max=7 | |
| 21:59:50 | mriedem | you wrote that patch | |
| 22:02:44 | jgwentworth | zzzeek: are you around? I'm looking into something with oslo.db and was wondering, given a TransactionContextManager, how does it switch from reader to writer mode and vice versa between separate transactions? | |
| 22:07:37 | mriedem | superdan: done https://review.openstack.org/#/c/510203/ | |
| 22:07:55 | mriedem | superdan: 2 issues, (1) the tests aren't using the args passed in and (2) apparently we have to handle boolean sort keys differently | |
| 22:08:14 | mriedem | superdan: presumably because not all db backends model booleans the same | |
| 22:08:17 | mriedem | some are bools, some are ints | |
| 22:08:31 | cfriesen | mriedem: I'm thinking about the case where we request min=5 and max=10 and the scheduler only finds hosts for 7 of them. | |
| 22:08:36 | superdan | mriedem: ah I didn't think we had any booleans, but I guess we do | |
| 22:09:14 | superdan | mriedem: hah, sorry, last minute cleanup on the tests and I didn't remember to plumb those through | |
| 22:09:38 | melwitt | zzzeek: it looks like it's controlled by the "independent" attribute, but I don't see that we use that | |
| 22:10:06 | cfriesen | mriedem: I think the ones without hosts might get left in the BUILD state | |
| 22:13:20 | mriedem | cfriesen: if you get fewer hosts than requested instances, it's NoValidHost https://github.com/openstack/nova/blob/master/nova/scheduler/filter_scheduler.py#L86 | |