| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-21 | |||
| 15:01:56 | openstackgerrit | Dan Smith proposed openstack/nova master: Add tests to validate instance_list handles faults correctly https://review.openstack.org/505392 | |
| 15:01:57 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix a pagination logic bug in test_bug_1689692 https://review.openstack.org/505661 | |
| 15:01:57 | openstackgerrit | Dan Smith proposed openstack/nova master: Copy some tests to a cellsv1 mixin https://review.openstack.org/505442 | |
| 15:01:58 | openstackgerrit | Dan Smith proposed openstack/nova master: Remove legacy fault-loading routines https://review.openstack.org/505456 | |
| 15:01:58 | openstackgerrit | Dan Smith proposed openstack/nova master: Use improved instance_list module in compute API https://review.openstack.org/505418 | |
| 15:02:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: api-ref: fix default sort key when listing servers https://review.openstack.org/506227 | |
| 15:02:09 | andreykurilin | mriedem: hi! basically, I like to do great things, so rally is my choice :) but yeah, scale and performance testing are somehow related :) | |
| 15:03:22 | gibi | dansmith: I don't need proofs I have faith :) | |
| 15:03:46 | dansmith | gibi: I need proof for my own sanity :) | |
| 15:03:55 | gibi | :) | |
| 15:04:25 | gibi | yeah, sanity is important | |
| 15:05:42 | mriedem | andreykurilin: ok some of us got talking at the PTG about how we need some scale testing done on nova, especially with multiple cells, and thought you might be interested in doing that, | |
| 15:06:01 | mriedem | especially since godaddy is using cells v1 and presumably has to investigate migrating to cells v2 at some point | |
| 15:06:40 | mriedem | andreykurilin: more specifically, dansmith is working on a series of changes to make listing instances across cells more efficient, | |
| 15:06:48 | mriedem | it would be nice if we could get benchmarks before/after | |
| 15:07:39 | mriedem | i'd also be interested in general in how long it takes to create instances in ocata vs pike, since pike is now doing quotas differently and doing resource claims in the scheduler rather than the computes | |
| 15:08:28 | mriedem | i could do this a bit sloppy with devstack and using the fake virt driver, but that doesn't give me multiple cells | |
| 15:10:06 | andreykurilin | mriedem: so basically, it is quite easy to add scenarios for cells v1/v2 (and generate the load), but the problem is in where to launch it. I'll talk to some folks to try find the place for such research | |
| 15:15:51 | mriedem | bauzas: a few comments inline https://review.openstack.org/#/c/506092/ | |
| 15:16:06 | mriedem | andreykurilin: great, thanks | |
| 15:16:26 | mriedem | we've been trying to get feedback from large cells v1 users for awhile now and so far we don't get much feedback | |
| 15:18:26 | efried | sdague You want me to do anything about https://review.openstack.org/#/c/488137/17/nova/utils.py@1309 ? (There or in a fup?) | |
| 15:22:19 | bauzas | mriedem: k, cool, reading | |
| 15:24:50 | bauzas | mriedem: just a question for just one comment, I'm not sure why we need to disable the source host | |
| 15:25:05 | bauzas | mriedem: because you'd like to make sure you end up in some specific host? | |
| 15:30:26 | dansmith | edleafe: I dunno why, but gerrit is being confusing about the stack of actual code patches for your selection/alternates bit | |
| 15:30:36 | dansmith | edleafe: is this the bottom? https://review.openstack.org/#/c/486215/7 | |
| 15:33:47 | mriedem | bauzas: that's what was reported in the bug for the recreate, | |
| 15:33:54 | mriedem | but i guess that's not really necessary to reproduce the issue | |
| 15:34:07 | mriedem | because the scheduler will compare 1 host to 2 instances and fail | |
| 15:34:14 | mriedem | bauzas: yeah so nevermind that part | |
| 15:34:33 | bauzas | okay | |
| 15:35:14 | edleafe | dansmith: yep | |
| 15:35:47 | edleafe | dansmith: it seems to be adding specs and code together | |
| 15:37:13 | dansmith | edleafe: well, it is also jumping around with the related-changes view of the code | |
| 15:37:51 | dansmith | I guess it's because the add-allocations patch has been orphaned for a while | |
| 15:39:44 | mriedem | sdague: https://review.openstack.org/#/c/500190/ is the "re-proposal" of the ksa adapters spec if you want to go over that. it's not so much a re-proposal as nearly a re-write though. | |
| 15:40:33 | mriedem | edleafe: do i need to review those alternative hosts specs in any order? there is the one about the object model and then the other one. | |
| 15:40:49 | mriedem | or should those be merged? | |
| 15:41:15 | edleafe | mriedem: alternate hosts was a concept we agreed to in Pike but never wrote up in a spec | |
| 15:41:31 | edleafe | mriedem: the Selection object is independent | |
| 15:44:02 | dansmith | edleafe: is the host.updated=True thing the flag we use to avoid choosing that same host for another instance of a multi-create? | |
| 15:44:12 | sdague | mriedem: will look | |
| 15:44:28 | sdague | mriedem: I'm trying to sort out a final set of unit tests fails on the QEMU_IMG version thing | |
| 15:45:27 | edleafe | dansmith: I think so. That host state stuff is still black magic to me | |
| 15:45:39 | dansmith | edleafe: oh actually it's host.updated=None.. how. strange. | |
| 15:45:55 | dansmith | edleafe: okay I will comment based on that assumption and hedge a bit | |
| 15:47:00 | drwahl | does placement-api (in newton) require keystone v3? | |
| 15:47:42 | bauzas | dansmith: that's because updated is the flag for saying "please refresh my state" | |
| 15:48:19 | dansmith | bauzas: okay I'm confused about why we need to do that in this context then | |
| 15:48:39 | bauzas | dansmith: what's the exact concern you have in mind ? | |
| 15:48:48 | dansmith | oohh, | |
| 15:48:51 | bauzas | I just jumped on your above question | |
| 15:49:02 | dansmith | maybe this is "I'm going to consume this host, so make sure to refresh it soon" ? | |
| 15:49:08 | dansmith | https://review.openstack.org/#/c/486215/7/nova/scheduler/filter_scheduler.py | |
| 15:49:09 | dansmith | L92 | |
| 15:50:05 | bauzas | dansmith: no | |
| 15:50:14 | bauzas | dansmith: that's because we consume_from_request() before | |
| 15:50:35 | bauzas | dansmith: so the in-memory HostState is modified | |
| 15:50:42 | mriedem | stvnoyes: the scsi volume live migration tempest test failed for the new style attach flows in a grenade job, meaning one of the computes in the live migration is pike and one is queens http://logs.openstack.org/90/481290/6/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/d8962cc/console.html#_2017-09-17_04_21_05_125937 | |
| 15:51:01 | bauzas | dansmith: in case the calculations were wrong, we need to unset those modifications and force a fresh new DB call | |
| 15:51:26 | dansmith | bauzas: meaning if we fail to do the consume, we want to trigger a refresh? | |
| 15:51:26 | bauzas | we generally keep in memory the HostStates once per request | |
| 15:51:44 | bauzas | but for a multi-schedule call, then we could iterate wrongly | |
| 15:51:46 | openstackgerrit | Merged openstack/os-vif master: Rehome OVO unit tests to tests.unit.test_object.py https://review.openstack.org/489922 | |
| 15:51:53 | bauzas | dansmith: yup that sort of | |
| 15:51:58 | dansmith | hrm | |
| 15:52:09 | dansmith | not sure I fully get it, but it's not the thing I thought, so that's fine | |
| 15:52:33 | mriedem | stvnoyes: looks like it fails here https://review.openstack.org/#/c/487884/7/tempest/api/compute/admin/test_live_migration.py@185 | |
| 15:52:38 | dansmith | edleafe: so, what follows is some rambling for improvements we can make later, but listen and nod (or not): | |
| 15:52:57 | bauzas | dansmith: see HostManager.consume_from_request() and you'll understand | |
| 15:53:21 | dansmith | edleafe: I think what you're doing here is going through the list of hosts like we do today, and picking the target and alternates for each one, | |
| 15:54:02 | dansmith | edleafe: which may mean your first alternate for instance 2 may be the the target for instance 3 and so on, such that all the alternates for early instances are almost definitely going to be burned by the later instances for any num_instances>max_retries | |
| 15:54:03 | dansmith | right? | |
| 15:54:29 | dansmith | so, hmm, yeah if that's right, that seems like a problem | |
| 15:54:31 | edleafe | dansmith: yes. Selecting an alternate doesn't consume anything | |
| 15:54:52 | dansmith | what we probably need to do is go through and pick primaries, and then pick alternates for each after that | |
| 15:55:03 | dansmith | the alternates can overlap, but they shouldn't overlap with any primaries | |
| 15:55:21 | edleafe | dansmith: ok, that could be done | |
| 15:55:48 | dansmith | edleafe: I'm just thinking that we're basically never going to be able to reschedule those early instances if we do it in the order you have, which is probably bad :) | |
| 15:56:04 | dansmith | edleafe: I shall add comments about this | |
| 15:56:06 | edleafe | I think in an earlier version of this we didn't think this would be a significant issue | |
| 15:56:54 | dansmith | yeah, I think we discussed it before, indeed | |
| 15:56:55 | dansmith | I think i was focusing on the overlapping of the alternates, | |
| 15:56:57 | edleafe | since a) alternates should rarely be used and b) hosts might still fit an additional instance | |
| 15:56:59 | dansmith | and not considering the overlapping of primaries | |
| 15:57:01 | dansmith | yeah | |
| 15:57:26 | stvnoyes | mriedem: ok, Ill take a look | |
| 15:58:41 | gibi | dansmith: I think I found why the target_cell context manager eats our MarkerNotFoundException, see my comment in https://review.openstack.org/#/c/504986/6/nova/compute/instance_list.py@85 | |
| 15:59:21 | dansmith | gibi: ahh, that makes more sense.. I forgot we were using the figure's target_cell here | |
| 15:59:40 | dansmith | gibi: I wrote that fixture, so I'll fix that up and make this change at the end of this series | |
| 16:00:00 | dansmith | gibi: now, please, go find bugs in someone else's code :) | |
| 16:00:34 | efried | sdague mordred Design point about bp/use-ksa-adapter-for-endpoints: When setting up for a service where we have the ability to get auth from context, what should we do about auth in conf? Options: a) Always use the context auth, don't even register auth options in the conf; b) Register auth conf options and allow them to override the context auth. | |
| 16:00:45 | efried | cdent ^ you may also have an opinion | |
| 16:01:03 | openstackgerrit | Merged openstack/nova-specs master: Spec: Use keystoneauth1 Adapter for endpoints https://review.openstack.org/500190 | |
| 16:01:36 | sdague | efried: it depends on what the subcall is doing | |
| 16:01:48 | sdague | if it is acting on behalf of the user, it should use the context auth | |
| 16:02:02 | gibi | dansmith: OK, I will review that follow up too. But instead of looking at others code I will just stop looking at any code for today | |
| 16:02:04 | sdague | optionally wrapped in a service token so that it doesn't expire | |
| 16:02:23 | dansmith | gibi: as long as it's not finding bugs in _my_ code I'm fine with whatever :P | |