Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-28
14:58:44 dansmith what is api-level compute?
14:58:46 mriedem edleafe: there is no such thing
14:59:08 edleafe mriedem: I didn't think so, but it sounded like you were
14:59:23 mriedem https://docs.openstack.org/nova/pike/user/cellsv2_layout.html#multiple-cells
14:59:34 edleafe the call up from compute isn't after select_destinations; it's before in a resize
14:59:38 mriedem in ^ the compute only has access to the cell conductor
14:59:54 mriedem i think we're talking about different things
15:00:07 edleafe yes, we are - that's what I've been trying to say
15:00:26 edleafe alternate hosts just removes the need for a retry to have to call up from the cell
15:00:37 mriedem resize flow is, summarized: api -> superconductor -> scheduler -> superconductor -> compute (reschedule) -> cell conductor -> compute (with alternate hosts)
15:00:50 mriedem yes, in ^ we can't upcall from the cell conductor to the scheduler
15:00:50 edleafe in your resize flow, the problem is that the call from the cell already is happening, and needs to change
15:00:55 mriedem hence the need to pass the alternate hosts through
15:00:58 mriedem for both build and resize
15:01:46 johnthetubaguy https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L3847
15:01:49 openstackgerrit Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
15:01:49 openstackgerrit Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948
15:01:50 openstackgerrit Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950
15:01:50 openstackgerrit Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949
15:01:51 openstackgerrit Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419
15:01:51 openstackgerrit Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420
15:01:52 openstackgerrit Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638
15:02:19 johnthetubaguy vs https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L1880
15:02:22 johnthetubaguy seems the same flow
15:02:38 johnthetubaguy i.e. +1 mriedem
15:03:43 johnthetubaguy I guess its the second visit here that should not call the scheduler again: https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/tasks/migrate.py#L67
15:03:53 mriedem correct
15:04:33 mriedem just like this one https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/manager.py#L552
15:04:40 johnthetubaguy ++
15:04:42 mriedem ^ is the build reschedule within the cell conductor
15:06:30 edleafe well, you've succeeded in completely confusing me as to what you want in the spec
15:08:04 edleafe Will it be enough to add that resize can also make use of the alternates? Or are there some other logical flows that need to be changed?
15:08:26 mriedem i think that is enough
15:08:37 mriedem and point out that live migration does reschedules, but within superconductor so we don't need to worry about those
15:09:46 johnthetubaguy edleafe: given the bits I just found out that doesn't do retries, I am +1 what mriedem just said
15:11:22 edleafe ok, I'll push another revision soon
15:39:19 efried jaypipes Any plan to include a "friendly name" or "description" field on resource provider?
15:39:34 jaypipes efried: there already is.
15:39:37 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce basic framework for security proxying https://review.openstack.org/345396
15:39:37 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce framework for RFB authentication https://review.openstack.org/345397
15:39:38 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce the VeNCrypt RFB authentication scheme https://review.openstack.org/345398
15:39:38 openstackgerrit Stephen Finucane proposed openstack/nova master: console: provide an RFB security proxy implementation https://review.openstack.org/345399
15:39:39 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544
15:40:19 jaypipes efried: https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/resource_provider.py#L33
15:40:20 efried woot.
15:40:32 efried I see it in the API doc now. Not sure how I missed it.
15:40:33 efried Thanks.
15:40:47 jaypipes efried: that's why the ProviderTree allows finding a provider by name or UUID...
15:40:59 efried oh yeah. I should really get back to reviewing that series.
15:41:14 cdent jaypipes, efried: we currently uniq on that name field, if I remember right, and that’s caused problems for some situations, should we change it?
15:41:18 jaypipes efried: hold off. pushing a new series of revisions after fixing up comments from gibi
15:41:30 efried jaypipes rgr wilco
15:41:37 cdent https://bugs.launchpad.net/nova/+bug/1714248
15:41:38 openstack Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed]
15:41:41 jaypipes cdent: you mean we're *not* unique?
15:42:47 efried but no skin in the game. Reading bugh...
15:42:57 cdent jaypipes: we enforce uniqueness on the name column
15:43:15 jaypipes cdent: ah, great. that's good then.
15:43:25 cdent jaypipes: which people are saying is bad
15:43:41 cdent because they want to blue green a compute node or something
15:43:55 cdent and names get in the way
15:45:09 efried cdent Skimming this bug, it seems to me like the problem is that the compute service that's taking over should *not* be attempting to create a new RP.
15:45:19 efried It should be looking up and using the old one.
15:45:30 openstackgerrit Matt Riedemann proposed openstack/nova master: Ensure instance can migrate when launched concurrently https://review.openstack.org/506093
15:45:35 mriedem bauzas: ^ is all cleaned up now
15:45:45 cdent efried: that would be one way to solve that particular problem
15:45:47 efried Is the RP somehow associated with the compute process rather than the node(s) it's managing?
15:46:00 cfriesen efried: do we ensure hypervisor host name uniqueness?
15:46:34 cfriesen efried: actually, I guess this is for ironic, so it's not even hypervisor, just host
15:46:59 cfriesen or is there some other way to uniquely identify a host in ironic (I know nothing)
15:47:43 efried Yeah, I don't know from ironic. Is it like, compute runs on the hypervisor and that represents a single possible instance?
15:47:58 efried (johnthetubaguy this conversation may interest you)
15:48:05 cfriesen efried: I believe nova-compute represents a bunch of baremetal machines
15:49:01 efried And that nova-compute could run on any one of several possible "hypervisors" that all "manage" that bunch of baremetal machines?
15:49:02 cfriesen efried: so then the node where nova-compute is running crashes, and they want a different nova-compute process to take over managing those baremetal machines.
15:49:43 efried Cool. So my question (cdent, johnthetubaguy): Why is there a RP associated with the nova-compute at all?? Seems like there should just be a RP per baremetal machine.
15:50:34 cdent efried: you’d think so, but no
15:50:44 cdent a node has inventory of classes of baremetal
15:51:01 cdent (last I recall, it’s hard to remember/keep straight)
15:51:03 cfriesen efried: I think the idea is that each baremetal machine is a resource
15:51:27 cfriesen or is several types of resource
15:51:45 cfriesen since you always claim a whole machine at a time
15:52:40 efried But but but... that would mean that every baremetal node in that nova-compute's purview is "identical".
15:53:21 mriedem johnthetubaguy: while you're awake, this is holding up the live migration new style volume attach change, and is simple https://review.openstack.org/#/c/506805/
15:56:56 cdent efried: ironic’s virtdriver’s get_inventory may explain things a bit: https://github.com/openstack/nova/blob/ae4b5d0147cb3e345bf57034221e9c8fedf3cad2/nova/virt/ironic/driver.py#L751
15:57:14 efried cdent Was just looking at that.
15:57:43 efried And was looking for the change jaypipes was going TODO.
15:57:48 cdent i’m not sure, but I think you can different classes of baremental node, resprsented by multiple custom resource classes
15:57:58 efried And was going to look for where the RP is set up, but not sure where to start looking for that.
15:58:01 openstackgerrit Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275
15:58:17 edleafe mriedem: johnthetubaguy: ^^
15:58:43 efried cdent Yeah, sounds like that's where they're going, but haven't yet.
15:59:07 cdent efried: edleafe did some further work on that chunk of code
15:59:48 johnthetubaguy I am keen to help in the ironic scheduling side of things btw, trying to write up all the traits discussions
16:00:29 cfriesen cdent: why wouldn't you follow efried's suggestion and have an RP per baremetal machine that reports resources for memory/disk/cpu? Then when you want to allocate a baremetal node you look for the smallest machine that can provide what you're looking for.
16:02:31 cfriesen I suspect this has been discussed already, so if anyone has links to the discussion....
16:02:47 efried johnthetubaguy Yeah, I was just poring over your https://review.openstack.org/#/c/507052/2/specs/queens/approved/ironic-traits.rst
16:03:13 efried which is actually what led me down the rabbit hole and ultimately prompted this discussion :)
16:03:37 efried Though the name conflict bug thing is a tangent-to-a-tangent...
16:03:52 johnthetubaguy so for ironic, resource class matching, and requesting no VCPU,MEM, etc, is the current way forward I thought?

Earlier   Later