| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-28 | |||
| 14:56:02 | edleafe | mriedem: I don't know conductor flows that clearly, but that sounds like a very different problem than retries | |
| 14:56:20 | johnthetubaguy | its the same compute -> wrong conductor right? | |
| 14:56:25 | edleafe | mriedem: that sounds like the entire flow needs to change, and alternate hosts won't address that | |
| 14:57:02 | johnthetubaguy | you replace a call to select_destinations to a claim the next candidate host right? | |
| 14:57:14 | johnthetubaguy | (wibble passing the data through) | |
| 14:57:38 | edleafe | johnthetubaguy: the call up happens before select_destinations, at least in mriedem's flow | |
| 14:57:42 | mriedem | edleafe: it's essentially the same as the build flow | |
| 14:57:49 | mriedem | just different methods | |
| 14:58:04 | mriedem | compute:build_and_run_instances calls up to conductor:build_and_run_instances | |
| 14:58:13 | mriedem | compute:prep_resize calls up to conductor:resize_instance | |
| 14:58:33 | mriedem | both of those methods in the cell conductor eventually ask the scheduler for a new dest | |
| 14:58:34 | edleafe | oh, you're talking about API-level compute, not cell compute | |
| 14:58:44 | dansmith | what is api-level compute? | |
| 14:58:46 | mriedem | edleafe: there is no such thing | |
| 14:59:08 | edleafe | mriedem: I didn't think so, but it sounded like you were | |
| 14:59:23 | mriedem | https://docs.openstack.org/nova/pike/user/cellsv2_layout.html#multiple-cells | |
| 14:59:34 | edleafe | the call up from compute isn't after select_destinations; it's before in a resize | |
| 14:59:38 | mriedem | in ^ the compute only has access to the cell conductor | |
| 14:59:54 | mriedem | i think we're talking about different things | |
| 15:00:07 | edleafe | yes, we are - that's what I've been trying to say | |
| 15:00:26 | edleafe | alternate hosts just removes the need for a retry to have to call up from the cell | |
| 15:00:37 | mriedem | resize flow is, summarized: api -> superconductor -> scheduler -> superconductor -> compute (reschedule) -> cell conductor -> compute (with alternate hosts) | |
| 15:00:50 | edleafe | in your resize flow, the problem is that the call from the cell already is happening, and needs to change | |
| 15:00:50 | mriedem | yes, in ^ we can't upcall from the cell conductor to the scheduler | |
| 15:00:55 | mriedem | hence the need to pass the alternate hosts through | |
| 15:00:58 | mriedem | for both build and resize | |
| 15:01:46 | johnthetubaguy | https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L3847 | |
| 15:01:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948 | |
| 15:01:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947 | |
| 15:01:50 | openstackgerrit | Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949 | |
| 15:01:50 | openstackgerrit | Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950 | |
| 15:01:51 | openstackgerrit | Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420 | |
| 15:01:51 | openstackgerrit | Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419 | |
| 15:01:52 | openstackgerrit | Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638 | |
| 15:02:19 | johnthetubaguy | vs https://github.com/openstack/nova/blob/8a386b055c82df67092a1abc683e7225ef80671e/nova/compute/manager.py#L1880 | |
| 15:02:22 | johnthetubaguy | seems the same flow | |
| 15:02:38 | johnthetubaguy | i.e. +1 mriedem | |
| 15:03:43 | johnthetubaguy | I guess its the second visit here that should not call the scheduler again: https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/tasks/migrate.py#L67 | |
| 15:03:53 | mriedem | correct | |
| 15:04:33 | mriedem | just like this one https://github.com/openstack/nova/blob/7cd9e3b8bb7fc0601786847f19cdf3f706ec079f/nova/conductor/manager.py#L552 | |
| 15:04:40 | johnthetubaguy | ++ | |
| 15:04:42 | mriedem | ^ is the build reschedule within the cell conductor | |
| 15:06:30 | edleafe | well, you've succeeded in completely confusing me as to what you want in the spec | |
| 15:08:04 | edleafe | Will it be enough to add that resize can also make use of the alternates? Or are there some other logical flows that need to be changed? | |
| 15:08:26 | mriedem | i think that is enough | |
| 15:08:37 | mriedem | and point out that live migration does reschedules, but within superconductor so we don't need to worry about those | |
| 15:09:46 | johnthetubaguy | edleafe: given the bits I just found out that doesn't do retries, I am +1 what mriedem just said | |
| 15:11:22 | edleafe | ok, I'll push another revision soon | |
| 15:39:19 | efried | jaypipes Any plan to include a "friendly name" or "description" field on resource provider? | |
| 15:39:34 | jaypipes | efried: there already is. | |
| 15:39:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: console: introduce framework for RFB authentication https://review.openstack.org/345397 | |
| 15:39:37 | openstackgerrit | Stephen Finucane proposed openstack/nova master: console: introduce basic framework for security proxying https://review.openstack.org/345396 | |
| 15:39:38 | openstackgerrit | Stephen Finucane proposed openstack/nova master: console: provide an RFB security proxy implementation https://review.openstack.org/345399 | |
| 15:39:38 | openstackgerrit | Stephen Finucane proposed openstack/nova master: console: introduce the VeNCrypt RFB authentication scheme https://review.openstack.org/345398 | |
| 15:39:39 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Document TLS security setup for noVNC proxy https://review.openstack.org/500544 | |
| 15:40:19 | jaypipes | efried: https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/resource_provider.py#L33 | |
| 15:40:20 | efried | woot. | |
| 15:40:32 | efried | I see it in the API doc now. Not sure how I missed it. | |
| 15:40:33 | efried | Thanks. | |
| 15:40:47 | jaypipes | efried: that's why the ProviderTree allows finding a provider by name or UUID... | |
| 15:40:59 | efried | oh yeah. I should really get back to reviewing that series. | |
| 15:41:14 | cdent | jaypipes, efried: we currently uniq on that name field, if I remember right, and that’s caused problems for some situations, should we change it? | |
| 15:41:18 | jaypipes | efried: hold off. pushing a new series of revisions after fixing up comments from gibi | |
| 15:41:30 | efried | jaypipes rgr wilco | |
| 15:41:37 | cdent | https://bugs.launchpad.net/nova/+bug/1714248 | |
| 15:41:38 | openstack | Launchpad bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] | |
| 15:41:41 | jaypipes | cdent: you mean we're *not* unique? | |
| 15:42:47 | efried | but no skin in the game. Reading bugh... | |
| 15:42:57 | cdent | jaypipes: we enforce uniqueness on the name column | |
| 15:43:15 | jaypipes | cdent: ah, great. that's good then. | |
| 15:43:25 | cdent | jaypipes: which people are saying is bad | |
| 15:43:41 | cdent | because they want to blue green a compute node or something | |
| 15:43:55 | cdent | and names get in the way | |
| 15:45:09 | efried | cdent Skimming this bug, it seems to me like the problem is that the compute service that's taking over should *not* be attempting to create a new RP. | |
| 15:45:19 | efried | It should be looking up and using the old one. | |
| 15:45:30 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Ensure instance can migrate when launched concurrently https://review.openstack.org/506093 | |
| 15:45:35 | mriedem | bauzas: ^ is all cleaned up now | |
| 15:45:45 | cdent | efried: that would be one way to solve that particular problem | |
| 15:45:47 | efried | Is the RP somehow associated with the compute process rather than the node(s) it's managing? | |
| 15:46:00 | cfriesen | efried: do we ensure hypervisor host name uniqueness? | |
| 15:46:34 | cfriesen | efried: actually, I guess this is for ironic, so it's not even hypervisor, just host | |
| 15:46:59 | cfriesen | or is there some other way to uniquely identify a host in ironic (I know nothing) | |
| 15:47:43 | efried | Yeah, I don't know from ironic. Is it like, compute runs on the hypervisor and that represents a single possible instance? | |
| 15:47:58 | efried | (johnthetubaguy this conversation may interest you) | |
| 15:48:05 | cfriesen | efried: I believe nova-compute represents a bunch of baremetal machines | |
| 15:49:01 | efried | And that nova-compute could run on any one of several possible "hypervisors" that all "manage" that bunch of baremetal machines? | |
| 15:49:02 | cfriesen | efried: so then the node where nova-compute is running crashes, and they want a different nova-compute process to take over managing those baremetal machines. | |
| 15:49:43 | efried | Cool. So my question (cdent, johnthetubaguy): Why is there a RP associated with the nova-compute at all?? Seems like there should just be a RP per baremetal machine. | |
| 15:50:34 | cdent | efried: you’d think so, but no | |
| 15:50:44 | cdent | a node has inventory of classes of baremetal | |
| 15:51:01 | cdent | (last I recall, it’s hard to remember/keep straight) | |
| 15:51:03 | cfriesen | efried: I think the idea is that each baremetal machine is a resource | |
| 15:51:27 | cfriesen | or is several types of resource | |
| 15:51:45 | cfriesen | since you always claim a whole machine at a time | |
| 15:52:40 | efried | But but but... that would mean that every baremetal node in that nova-compute's purview is "identical". | |
| 15:53:21 | mriedem | johnthetubaguy: while you're awake, this is holding up the live migration new style volume attach change, and is simple https://review.openstack.org/#/c/506805/ | |
| 15:56:56 | cdent | efried: ironic’s virtdriver’s get_inventory may explain things a bit: https://github.com/openstack/nova/blob/ae4b5d0147cb3e345bf57034221e9c8fedf3cad2/nova/virt/ironic/driver.py#L751 | |
| 15:57:14 | efried | cdent Was just looking at that. | |
| 15:57:43 | efried | And was looking for the change jaypipes was going TODO. | |
| 15:57:48 | cdent | i’m not sure, but I think you can different classes of baremental node, resprsented by multiple custom resource classes | |