Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-21
15:18:26 efried sdague You want me to do anything about https://review.openstack.org/#/c/488137/17/nova/utils.py@1309 ? (There or in a fup?)
15:22:19 bauzas mriedem: k, cool, reading
15:24:50 bauzas mriedem: just a question for just one comment, I'm not sure why we need to disable the source host
15:25:05 bauzas mriedem: because you'd like to make sure you end up in some specific host?
15:30:26 dansmith edleafe: I dunno why, but gerrit is being confusing about the stack of actual code patches for your selection/alternates bit
15:30:36 dansmith edleafe: is this the bottom? https://review.openstack.org/#/c/486215/7
15:33:47 mriedem bauzas: that's what was reported in the bug for the recreate,
15:33:54 mriedem but i guess that's not really necessary to reproduce the issue
15:34:07 mriedem because the scheduler will compare 1 host to 2 instances and fail
15:34:14 mriedem bauzas: yeah so nevermind that part
15:34:33 bauzas okay
15:35:14 edleafe dansmith: yep
15:35:47 edleafe dansmith: it seems to be adding specs and code together
15:37:13 dansmith edleafe: well, it is also jumping around with the related-changes view of the code
15:37:51 dansmith I guess it's because the add-allocations patch has been orphaned for a while
15:39:44 mriedem sdague: https://review.openstack.org/#/c/500190/ is the "re-proposal" of the ksa adapters spec if you want to go over that. it's not so much a re-proposal as nearly a re-write though.
15:40:33 mriedem edleafe: do i need to review those alternative hosts specs in any order? there is the one about the object model and then the other one.
15:40:49 mriedem or should those be merged?
15:41:15 edleafe mriedem: alternate hosts was a concept we agreed to in Pike but never wrote up in a spec
15:41:31 edleafe mriedem: the Selection object is independent
15:44:02 dansmith edleafe: is the host.updated=True thing the flag we use to avoid choosing that same host for another instance of a multi-create?
15:44:12 sdague mriedem: will look
15:44:28 sdague mriedem: I'm trying to sort out a final set of unit tests fails on the QEMU_IMG version thing
15:45:27 edleafe dansmith: I think so. That host state stuff is still black magic to me
15:45:39 dansmith edleafe: oh actually it's host.updated=None.. how. strange.
15:45:55 dansmith edleafe: okay I will comment based on that assumption and hedge a bit
15:47:00 drwahl does placement-api (in newton) require keystone v3?
15:47:42 bauzas dansmith: that's because updated is the flag for saying "please refresh my state"
15:48:19 dansmith bauzas: okay I'm confused about why we need to do that in this context then
15:48:39 bauzas dansmith: what's the exact concern you have in mind ?
15:48:48 dansmith oohh,
15:48:51 bauzas I just jumped on your above question
15:49:02 dansmith maybe this is "I'm going to consume this host, so make sure to refresh it soon" ?
15:49:08 dansmith https://review.openstack.org/#/c/486215/7/nova/scheduler/filter_scheduler.py
15:49:09 dansmith L92
15:50:05 bauzas dansmith: no
15:50:14 bauzas dansmith: that's because we consume_from_request() before
15:50:35 bauzas dansmith: so the in-memory HostState is modified
15:50:42 mriedem stvnoyes: the scsi volume live migration tempest test failed for the new style attach flows in a grenade job, meaning one of the computes in the live migration is pike and one is queens http://logs.openstack.org/90/481290/6/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/d8962cc/console.html#_2017-09-17_04_21_05_125937
15:51:01 bauzas dansmith: in case the calculations were wrong, we need to unset those modifications and force a fresh new DB call
15:51:26 dansmith bauzas: meaning if we fail to do the consume, we want to trigger a refresh?
15:51:26 bauzas we generally keep in memory the HostStates once per request
15:51:44 bauzas but for a multi-schedule call, then we could iterate wrongly
15:51:46 openstackgerrit Merged openstack/os-vif master: Rehome OVO unit tests to tests.unit.test_object.py https://review.openstack.org/489922
15:51:53 bauzas dansmith: yup that sort of
15:51:58 dansmith hrm
15:52:09 dansmith not sure I fully get it, but it's not the thing I thought, so that's fine
15:52:33 mriedem stvnoyes: looks like it fails here https://review.openstack.org/#/c/487884/7/tempest/api/compute/admin/test_live_migration.py@185
15:52:38 dansmith edleafe: so, what follows is some rambling for improvements we can make later, but listen and nod (or not):
15:52:57 bauzas dansmith: see HostManager.consume_from_request() and you'll understand
15:53:21 dansmith edleafe: I think what you're doing here is going through the list of hosts like we do today, and picking the target and alternates for each one,
15:54:02 dansmith edleafe: which may mean your first alternate for instance 2 may be the the target for instance 3 and so on, such that all the alternates for early instances are almost definitely going to be burned by the later instances for any num_instances>max_retries
15:54:03 dansmith right?
15:54:29 dansmith so, hmm, yeah if that's right, that seems like a problem
15:54:31 edleafe dansmith: yes. Selecting an alternate doesn't consume anything
15:54:52 dansmith what we probably need to do is go through and pick primaries, and then pick alternates for each after that
15:55:03 dansmith the alternates can overlap, but they shouldn't overlap with any primaries
15:55:21 edleafe dansmith: ok, that could be done
15:55:48 dansmith edleafe: I'm just thinking that we're basically never going to be able to reschedule those early instances if we do it in the order you have, which is probably bad :)
15:56:04 dansmith edleafe: I shall add comments about this
15:56:06 edleafe I think in an earlier version of this we didn't think this would be a significant issue
15:56:54 dansmith yeah, I think we discussed it before, indeed
15:56:55 dansmith I think i was focusing on the overlapping of the alternates,
15:56:57 edleafe since a) alternates should rarely be used and b) hosts might still fit an additional instance
15:56:59 dansmith and not considering the overlapping of primaries
15:57:01 dansmith yeah
15:57:26 stvnoyes mriedem: ok, Ill take a look
15:58:41 gibi dansmith: I think I found why the target_cell context manager eats our MarkerNotFoundException, see my comment in https://review.openstack.org/#/c/504986/6/nova/compute/instance_list.py@85
15:59:21 dansmith gibi: ahh, that makes more sense.. I forgot we were using the figure's target_cell here
15:59:40 dansmith gibi: I wrote that fixture, so I'll fix that up and make this change at the end of this series
16:00:00 dansmith gibi: now, please, go find bugs in someone else's code :)
16:00:34 efried sdague mordred Design point about bp/use-ksa-adapter-for-endpoints: When setting up for a service where we have the ability to get auth from context, what should we do about auth in conf? Options: a) Always use the context auth, don't even register auth options in the conf; b) Register auth conf options and allow them to override the context auth.
16:00:45 efried cdent ^ you may also have an opinion
16:01:03 openstackgerrit Merged openstack/nova-specs master: Spec: Use keystoneauth1 Adapter for endpoints https://review.openstack.org/500190
16:01:36 sdague efried: it depends on what the subcall is doing
16:01:48 sdague if it is acting on behalf of the user, it should use the context auth
16:02:02 gibi dansmith: OK, I will review that follow up too. But instead of looking at others code I will just stop looking at any code for today
16:02:04 sdague optionally wrapped in a service token so that it doesn't expire
16:02:23 dansmith gibi: as long as it's not finding bugs in _my_ code I'm fine with whatever :P
16:03:15 gibi dansmith: :)
16:03:30 efried sdague I guess there's also c) Don't use the context auth; require the conf auth.
16:03:38 sdague efried: yeh
16:03:39 efried sdague I hear you saying it's gonna be case by case.
16:03:41 dansmith gibi: I'll add you to the review of the fixture cleanup when I have it
16:03:45 sdague it's going to be case by case
16:03:59 efried sdague I'll likely need some help identifying which is which.
16:04:12 sdague efried: so, I think the answer ends up being roughly
16:04:23 sdague glance, cinder, barbican, keystone always act as user
16:04:41 sdague ironic always from conf, because nova is the ironic multi tenancy solution
16:04:51 sdague and neutron, it depends on the operation
16:05:31 efried sdague Okay. For that first set: You already pushed for glance to be able to get the auth from context, so not c). Should we b) allow conf override or just a) always use context?
16:05:59 sdague efried: for glance, I don't think so
16:06:13 sdague I can't think of a case where we should "sudo" on image things
16:06:27 sdague or that image things have to happen outside of a user context
16:07:08 sdague we should, whenever possible, really operate in the user context of the user token we got, because it ensures we don't have an unintended priv escalation
16:07:26 sdague ironic is a special case, ironic doesn't have regular users
16:07:31 dansmith edleafe: sanity check my tome of words on that review?
16:07:38 efried (It's worth noting that part of the mission statement of this bp was consistency. Sigh.)
16:08:09 mordred efried, sdague: "optionally wrapped in a service token" doesn't need nova-specific creds does it?
16:08:21 edleafe dansmith: on a call and an IRC meeting, so it'll be later

Earlier   Later