Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-28
14:52:40 edleafe mriedem: if you think it's wise, I can expand it to cover every use of select_destinations()
14:52:44 mriedem edleafe: ok, but we also got into a lot of hot water late in pike because the claims in the scheduler stuff didn't take into account move operations
14:53:03 mriedem so the main issue we're trying to solve is that the cell conductor can't reach the scheduler
14:53:08 mriedem there are only 2 times that happens,
14:53:09 mriedem build and resize
14:53:28 mriedem live migrate reschedules happen in super condcutor which can reach the scheduler, so that's fine - you could make a note of it as something we know about and don't need to change
14:53:30 edleafe mriedem: ok, I'll dig into resize, and add some stuff about that
14:53:39 johnthetubaguy evacuate or shelve? not sure if they retry at all?
14:53:44 mriedem johnthetubaguy: they don't
14:53:52 johnthetubaguy OK
14:54:05 mriedem prep_resize on the compute will call back to resize_instance in the conductor, which calls migrate_server, which calls the scheduler for a new destination
14:54:15 mriedem so that is the flow, besides build, that has to also be fixed
14:54:47 johnthetubaguy OK, so if its only two, we should do them together I think
14:54:51 mriedem after 5 years working in nova, i think i finally have these conductor flows memorized
14:55:04 johnthetubaguy they changed after I last did that
14:55:12 johnthetubaguy and I slept since then
14:55:21 cdent you sleep?
14:55:33 johnthetubaguy yeah, I know, old school
14:56:02 edleafe mriedem: I don't know conductor flows that clearly, but that sounds like a very different problem than retries
14:56:20 johnthetubaguy its the same compute -> wrong conductor right?
14:56:25 edleafe mriedem: that sounds like the entire flow needs to change, and alternate hosts won't address that
14:57:02 johnthetubaguy you replace a call to select_destinations to a claim the next candidate host right?
14:57:14 johnthetubaguy (wibble passing the data through)
14:57:38 edleafe johnthetubaguy: the call up happens before select_destinations, at least in mriedem's flow
14:57:42 mriedem edleafe: it's essentially the same as the build flow
14:57:49 mriedem just different methods
14:58:04 mriedem compute:build_and_run_instances calls up to conductor:build_and_run_instances
14:58:13 mriedem compute:prep_resize calls up to conductor:resize_instance
14:58:33 mriedem both of those methods in the cell conductor eventually ask the scheduler for a new dest
14:58:34 edleafe oh, you're talking about API-level compute, not cell compute
14:58:44 dansmith what is api-level compute?
14:58:46 mriedem edleafe: there is no such thing
14:59:08 edleafe mriedem: I didn't think so, but it sounded like you were
14:59:23 mriedem https://docs.openstack.org/nova/pike/user/cellsv2_layout.html#multiple-cells
14:59:34 edleafe the call up from compute isn't after select_destinations; it's before in a resize
14:59:38 mriedem in ^ the compute only has access to the cell conductor
14:59:54 mriedem i think we're talking about different things
15:00:07 edleafe yes, we are - that's what I've been trying to say
15:00:26 edleafe alternate hosts just removes the need for a retry to have to call up from the cell
15:00:37 mriedem resize flow is, summarized: api -> superconductor -> scheduler -> superconductor -> compute (reschedule) -> cell conductor -> compute (with alternate hosts)
15:00:50 edleafe in your resize flow, the problem is that the call from the cell already is happening, and needs to change
15:00:50 mriedem yes, in ^ we can't upcall from the cell conductor to the scheduler
15:00:55 mriedem hence the need to pass the alternate hosts through
15:00:58 mriedem for both build and resize
15:01:46 johnthetubaguy https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L3847
15:01:49 openstackgerrit Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948
15:01:49 openstackgerrit Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
15:01:50 openstackgerrit Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949
15:01:50 openstackgerrit Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950
15:01:51 openstackgerrit Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420
15:01:51 openstackgerrit Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419
15:01:52 openstackgerrit Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638
15:02:19 johnthetubaguy vs https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L1880
15:02:22 johnthetubaguy seems the same flow
15:02:38 johnthetubaguy i.e. +1 mriedem
15:03:43 johnthetubaguy I guess its the second visit here that should not call the scheduler again: https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/tasks/migrate.py#L67
15:03:53 mriedem correct
15:04:33 mriedem just like this one https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/manager.py#L552
15:04:40 johnthetubaguy ++
15:04:42 mriedem ^ is the build reschedule within the cell conductor
15:06:30 edleafe well, you've succeeded in completely confusing me as to what you want in the spec
15:08:04 edleafe Will it be enough to add that resize can also make use of the alternates? Or are there some other logical flows that need to be changed?
15:08:26 mriedem i think that is enough
15:08:37 mriedem and point out that live migration does reschedules, but within superconductor so we don't need to worry about those
15:09:46 johnthetubaguy edleafe: given the bits I just found out that doesn't do retries, I am +1 what mriedem just said
15:11:22 edleafe ok, I'll push another revision soon
15:39:19 efried jaypipes Any plan to include a "friendly name" or "description" field on resource provider?
15:39:34 jaypipes efried: there already is.
15:39:37 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce framework for RFB authentication https://review.openstack.org/345397
15:39:37 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce basic framework for security proxying https://review.openstack.org/345396
15:39:38 openstackgerrit Stephen Finucane proposed openstack/nova master: console: provide an RFB security proxy implementation https://review.openstack.org/345399
15:39:38 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce the VeNCrypt RFB authentication scheme https://review.openstack.org/345398
15:39:39 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544
15:40:19 jaypipes efried: https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/resource_provider.py#L33
15:40:20 efried woot.
15:40:32 efried I see it in the API doc now. Not sure how I missed it.
15:40:33 efried Thanks.
15:40:47 jaypipes efried: that's why the ProviderTree allows finding a provider by name or UUID...
15:40:59 efried oh yeah. I should really get back to reviewing that series.
15:41:14 cdent jaypipes, efried: we currently uniq on that name field, if I remember right, and that’s caused problems for some situations, should we change it?
15:41:18 jaypipes efried: hold off. pushing a new series of revisions after fixing up comments from gibi
15:41:30 efried jaypipes rgr wilco
15:41:37 cdent https://bugs.launchpad.net/nova/+bug/1714248
15:41:38 openstack Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed]
15:41:41 jaypipes cdent: you mean we're *not* unique?
15:42:47 efried but no skin in the game. Reading bugh...
15:42:57 cdent jaypipes: we enforce uniqueness on the name column
15:43:15 jaypipes cdent: ah, great. that's good then.
15:43:25 cdent jaypipes: which people are saying is bad
15:43:41 cdent because they want to blue green a compute node or something
15:43:55 cdent and names get in the way
15:45:09 efried cdent Skimming this bug, it seems to me like the problem is that the compute service that's taking over should *not* be attempting to create a new RP.
15:45:19 efried It should be looking up and using the old one.
15:45:30 openstackgerrit Matt Riedemann proposed openstack/nova master: Ensure instance can migrate when launched concurrently https://review.openstack.org/506093
15:45:35 mriedem bauzas: ^ is all cleaned up now
15:45:45 cdent efried: that would be one way to solve that particular problem
15:45:47 efried Is the RP somehow associated with the compute process rather than the node(s) it's managing?
15:46:00 cfriesen efried: do we ensure hypervisor host name uniqueness?
15:46:34 cfriesen efried: actually, I guess this is for ironic, so it's not even hypervisor, just host
15:46:59 cfriesen or is there some other way to uniquely identify a host in ironic (I know nothing)

Earlier   Later