Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-21
15:06:48 mriedem it would be nice if we could get benchmarks before/after
15:07:39 mriedem i'd also be interested in general in how long it takes to create instances in ocata vs pike, since pike is now doing quotas differently and doing resource claims in the scheduler rather than the computes
15:08:28 mriedem i could do this a bit sloppy with devstack and using the fake virt driver, but that doesn't give me multiple cells
15:10:06 andreykurilin mriedem: so basically, it is quite easy to add scenarios for cells v1/v2 (and generate the load), but the problem is in where to launch it. I'll talk to some folks to try find the place for such research
15:15:51 mriedem bauzas: a few comments inline https://review.openstack.org/#/c/506092/
15:16:06 mriedem andreykurilin: great, thanks
15:16:26 mriedem we've been trying to get feedback from large cells v1 users for awhile now and so far we don't get much feedback
15:18:26 efried sdague You want me to do anything about https://review.openstack.org/#/c/488137/17/nova/utils.py@1309 ? (There or in a fup?)
15:22:19 bauzas mriedem: k, cool, reading
15:24:50 bauzas mriedem: just a question for just one comment, I'm not sure why we need to disable the source host
15:25:05 bauzas mriedem: because you'd like to make sure you end up in some specific host?
15:30:26 dansmith edleafe: I dunno why, but gerrit is being confusing about the stack of actual code patches for your selection/alternates bit
15:30:36 dansmith edleafe: is this the bottom? https://review.openstack.org/#/c/486215/7
15:33:47 mriedem bauzas: that's what was reported in the bug for the recreate,
15:33:54 mriedem but i guess that's not really necessary to reproduce the issue
15:34:07 mriedem because the scheduler will compare 1 host to 2 instances and fail
15:34:14 mriedem bauzas: yeah so nevermind that part
15:34:33 bauzas okay
15:35:14 edleafe dansmith: yep
15:35:47 edleafe dansmith: it seems to be adding specs and code together
15:37:13 dansmith edleafe: well, it is also jumping around with the related-changes view of the code
15:37:51 dansmith I guess it's because the add-allocations patch has been orphaned for a while
15:39:44 mriedem sdague: https://review.openstack.org/#/c/500190/ is the "re-proposal" of the ksa adapters spec if you want to go over that. it's not so much a re-proposal as nearly a re-write though.
15:40:33 mriedem edleafe: do i need to review those alternative hosts specs in any order? there is the one about the object model and then the other one.
15:40:49 mriedem or should those be merged?
15:41:15 edleafe mriedem: alternate hosts was a concept we agreed to in Pike but never wrote up in a spec
15:41:31 edleafe mriedem: the Selection object is independent
15:44:02 dansmith edleafe: is the host.updated=True thing the flag we use to avoid choosing that same host for another instance of a multi-create?
15:44:12 sdague mriedem: will look
15:44:28 sdague mriedem: I'm trying to sort out a final set of unit tests fails on the QEMU_IMG version thing
15:45:27 edleafe dansmith: I think so. That host state stuff is still black magic to me
15:45:39 dansmith edleafe: oh actually it's host.updated=None.. how. strange.
15:45:55 dansmith edleafe: okay I will comment based on that assumption and hedge a bit
15:47:00 drwahl does placement-api (in newton) require keystone v3?
15:47:42 bauzas dansmith: that's because updated is the flag for saying "please refresh my state"
15:48:19 dansmith bauzas: okay I'm confused about why we need to do that in this context then
15:48:39 bauzas dansmith: what's the exact concern you have in mind ?
15:48:48 dansmith oohh,
15:48:51 bauzas I just jumped on your above question
15:49:02 dansmith maybe this is "I'm going to consume this host, so make sure to refresh it soon" ?
15:49:08 dansmith https://review.openstack.org/#/c/486215/7/nova/scheduler/filter_scheduler.py
15:49:09 dansmith L92
15:50:05 bauzas dansmith: no
15:50:14 bauzas dansmith: that's because we consume_from_request() before
15:50:35 bauzas dansmith: so the in-memory HostState is modified
15:50:42 mriedem stvnoyes: the scsi volume live migration tempest test failed for the new style attach flows in a grenade job, meaning one of the computes in the live migration is pike and one is queens http://logs.openstack.org/90/481290/6/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/d8962cc/console.html#_2017-09-17_04_21_05_125937
15:51:01 bauzas dansmith: in case the calculations were wrong, we need to unset those modifications and force a fresh new DB call
15:51:26 dansmith bauzas: meaning if we fail to do the consume, we want to trigger a refresh?
15:51:26 bauzas we generally keep in memory the HostStates once per request
15:51:44 bauzas but for a multi-schedule call, then we could iterate wrongly
15:51:46 openstackgerrit Merged openstack/os-vif master: Rehome OVO unit tests to tests.unit.test_object.py https://review.openstack.org/489922
15:51:53 bauzas dansmith: yup that sort of
15:51:58 dansmith hrm
15:52:09 dansmith not sure I fully get it, but it's not the thing I thought, so that's fine
15:52:33 mriedem stvnoyes: looks like it fails here https://review.openstack.org/#/c/487884/7/tempest/api/compute/admin/test_live_migration.py@185
15:52:38 dansmith edleafe: so, what follows is some rambling for improvements we can make later, but listen and nod (or not):
15:52:57 bauzas dansmith: see HostManager.consume_from_request() and you'll understand
15:53:21 dansmith edleafe: I think what you're doing here is going through the list of hosts like we do today, and picking the target and alternates for each one,
15:54:02 dansmith edleafe: which may mean your first alternate for instance 2 may be the the target for instance 3 and so on, such that all the alternates for early instances are almost definitely going to be burned by the later instances for any num_instances>max_retries
15:54:03 dansmith right?
15:54:29 dansmith so, hmm, yeah if that's right, that seems like a problem
15:54:31 edleafe dansmith: yes. Selecting an alternate doesn't consume anything
15:54:52 dansmith what we probably need to do is go through and pick primaries, and then pick alternates for each after that
15:55:03 dansmith the alternates can overlap, but they shouldn't overlap with any primaries
15:55:21 edleafe dansmith: ok, that could be done
15:55:48 dansmith edleafe: I'm just thinking that we're basically never going to be able to reschedule those early instances if we do it in the order you have, which is probably bad :)
15:56:04 dansmith edleafe: I shall add comments about this
15:56:06 edleafe I think in an earlier version of this we didn't think this would be a significant issue
15:56:54 dansmith yeah, I think we discussed it before, indeed
15:56:55 dansmith I think i was focusing on the overlapping of the alternates,
15:56:57 edleafe since a) alternates should rarely be used and b) hosts might still fit an additional instance
15:56:59 dansmith and not considering the overlapping of primaries
15:57:01 dansmith yeah
15:57:26 stvnoyes mriedem: ok, Ill take a look
15:58:41 gibi dansmith: I think I found why the target_cell context manager eats our MarkerNotFoundException, see my comment in https://review.openstack.org/#/c/504986/6/nova/compute/instance_list.py@85
15:59:21 dansmith gibi: ahh, that makes more sense.. I forgot we were using the figure's target_cell here
15:59:40 dansmith gibi: I wrote that fixture, so I'll fix that up and make this change at the end of this series
16:00:00 dansmith gibi: now, please, go find bugs in someone else's code :)
16:00:34 efried sdague mordred Design point about bp/use-ksa-adapter-for-endpoints: When setting up for a service where we have the ability to get auth from context, what should we do about auth in conf? Options: a) Always use the context auth, don't even register auth options in the conf; b) Register auth conf options and allow them to override the context auth.
16:00:45 efried cdent ^ you may also have an opinion
16:01:03 openstackgerrit Merged openstack/nova-specs master: Spec: Use keystoneauth1 Adapter for endpoints https://review.openstack.org/500190
16:01:36 sdague efried: it depends on what the subcall is doing
16:01:48 sdague if it is acting on behalf of the user, it should use the context auth
16:02:02 gibi dansmith: OK, I will review that follow up too. But instead of looking at others code I will just stop looking at any code for today
16:02:04 sdague optionally wrapped in a service token so that it doesn't expire
16:02:23 dansmith gibi: as long as it's not finding bugs in _my_ code I'm fine with whatever :P
16:03:15 gibi dansmith: :)
16:03:30 efried sdague I guess there's also c) Don't use the context auth; require the conf auth.
16:03:38 sdague efried: yeh
16:03:39 efried sdague I hear you saying it's gonna be case by case.
16:03:41 dansmith gibi: I'll add you to the review of the fixture cleanup when I have it
16:03:45 sdague it's going to be case by case
16:03:59 efried sdague I'll likely need some help identifying which is which.
16:04:12 sdague efried: so, I think the answer ends up being roughly
16:04:23 sdague glance, cinder, barbican, keystone always act as user
16:04:41 sdague ironic always from conf, because nova is the ironic multi tenancy solution
16:04:51 sdague and neutron, it depends on the operation
16:05:31 efried sdague Okay. For that first set: You already pushed for glance to be able to get the auth from context, so not c). Should we b) allow conf override or just a) always use context?
16:05:59 sdague efried: for glance, I don't think so
16:06:13 sdague I can't think of a case where we should "sudo" on image things

Earlier   Later