| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-21 | |||
| 16:42:03 | mriedem | dansmith: i think i'm just going to slap that into our official docs | |
| 16:42:17 | mriedem | "Technical reference deep dive > live migration > THIS" | |
| 16:42:27 | dansmith | heh | |
| 16:43:05 | openstackgerrit | Elod Illes proposed openstack/nova master: Add instance.interface_detach notification https://review.openstack.org/506284 | |
| 16:43:49 | mriedem | if only i were on the twitters | |
| 16:44:17 | sdague | mriedem: that is a solvable problem | |
| 16:45:18 | efried | mriedem What does "cast" mean? | |
| 16:45:44 | efried | async invocation? | |
| 16:46:28 | mriedem | rpc cast | |
| 16:46:29 | mriedem | vs call | |
| 16:46:42 | mriedem | which is important to understand the hot potato between the computes during live migration | |
| 16:46:51 | mriedem | especially if you love rpc timeouts | |
| 16:46:59 | mriedem | because your instance has 20 ports and 20 volumes attached to it | |
| 16:47:09 | mriedem | and token timeouts | |
| 16:47:33 | dansmith | mmmm, rpc timeout | |
| 16:49:07 | sdague | mriedem: service tokens fix that | |
| 16:49:22 | sdague | well, some of it | |
| 16:50:54 | mriedem | yeah i know | |
| 16:50:59 | mriedem | wonder if anyone is using those yet | |
| 17:04:25 | bauzas | dansmith: mriedem: question, when I was testing https://review.openstack.org/#/c/506093/2 I discovered a weird stack http://paste.openstack.org/show/621641/ | |
| 17:05:35 | bauzas | dansmith: mriedem: when looking at https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1227 it lookups the in-memory dict of all compute nodes for finding the right node | |
| 17:06:59 | bauzas | dansmith: mriedem: but AFAICS, we're passing the source node as attribute to the destination RT https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5735 | |
| 17:07:22 | bauzas | dansmith: mriedem: so since the destination RT doesn't know the source node, it fails with a KeyError | |
| 17:07:26 | bauzas | amirite? | |
| 17:07:44 | bauzas | unless post_live_mig runs on the source node | |
| 17:07:52 | bauzas | that's confusing | |
| 17:07:54 | mriedem | post live migrate runs on the source onde | |
| 17:07:56 | mriedem | *node | |
| 17:08:11 | mriedem | see my awesome call flow diagram above | |
| 17:08:21 | mriedem | gd this is already paying for itself | |
| 17:08:27 | mriedem | https://photos.app.goo.gl/Q8JdpjM0PZhAzsv32 | |
| 17:08:28 | bauzas | grmblbl | |
| 17:08:36 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix CellDatabases fixture swallowing exceptions https://review.openstack.org/506312 | |
| 17:09:08 | bauzas | mriedem: any reason you would see why I'm getting that KeyError ? I guess it's because I don't run the periodics ? | |
| 17:09:36 | openstackgerrit | Merged openstack/nova master: Updated from global requirements https://review.openstack.org/505839 | |
| 17:10:19 | openstackgerrit | Merged openstack/nova master: Fix hyperlinks in document https://review.openstack.org/506101 | |
| 17:12:25 | mriedem | bauzas: that should get set in memory when the compute service starts up | |
| 17:12:29 | mriedem | when it calls pre_start_hook | |
| 17:12:46 | bauzas | mmm | |
| 17:12:50 | mriedem | that will call update_available_resource_for_node | |
| 17:13:01 | bauzas | looking at the func tests, jay needed to explicitly start them | |
| 17:13:09 | mriedem | start what? | |
| 17:13:10 | mriedem | the computes? | |
| 17:13:14 | bauzas | the RT update things | |
| 17:13:14 | jaypipes | hmm? | |
| 17:13:14 | mriedem | the fixture starts those | |
| 17:13:43 | bauzas | here, I'm getting a trace because self.compute_nodes[my_node] isn't set yet | |
| 17:13:44 | mriedem | when the compute service starts, it will get the available nodes from the driver, and use those to call rt.update_available_resource | |
| 17:13:54 | mriedem | passing in the node name which gets populated in the rt.compute_nodes dict | |
| 17:13:55 | bauzas | which is populated AFAIK by running the RT update call | |
| 17:15:23 | bauzas | mmm, I'm introspecting | |
| 17:15:32 | mriedem | what you have in the test setup looks fine to me | |
| 17:15:36 | mriedem | and it's what we're doing in other tests | |
| 17:16:57 | bauzas | I'll upload my latest rev and see what the job insults me | |
| 17:17:12 | edleafe | mriedem: IMO, we need lots more flow diagrams like that | |
| 17:18:17 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: Add a regression test for bug 1718455 https://review.openstack.org/506092 | |
| 17:18:19 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: Fix definitely move single instance when created concurrently https://review.openstack.org/506093 | |
| 17:18:19 | openstack | bug 1718455 in OpenStack Compute (nova) "[pike] Nova host disable and Live Migrate all instances fail." [Medium,In progress] https://launchpad.net/bugs/1718455 - Assigned to Sylvain Bauza (sylvain-bauza) | |
| 17:18:49 | openstackgerrit | Merged openstack/nova master: Remove compatibility code for flavors https://review.openstack.org/460377 | |
| 17:19:28 | openstackgerrit | Merged openstack/nova master: Test InstanceNotFound handling in 'nova usage' https://review.openstack.org/468514 | |
| 17:26:14 | openstackgerrit | Merged openstack/nova master: VMware: Factor out relocate_vm() https://review.openstack.org/270115 | |
| 17:26:47 | mriedem | edleafe: agree, just needs to be put into docs | |
| 17:26:49 | openstackgerrit | Merged openstack/nova master: neutron: handle binding:profile=None during migration https://review.openstack.org/504260 | |
| 17:33:49 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: neutron: handle binding:profile=None during migration https://review.openstack.org/506319 | |
| 17:37:46 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/ocata: neutron: handle binding:profile=None during migration https://review.openstack.org/506320 | |
| 17:49:41 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/newton: neutron: handle binding:profile=None during migration https://review.openstack.org/506323 | |
| 17:54:38 | efried | mriedem Who are we looking to for the second +2 & +W for bp/use-ksa-adapter-for-endpoints work? (E.g. https://review.openstack.org/#/c/488137/) | |
| 17:55:41 | mriedem | me | |
| 17:55:45 | mriedem | i guess | |
| 17:55:55 | mriedem | i'm doing about 6 things at once right now though, so it's going to have to wait | |
| 17:56:55 | dansmith | efried: and three of those six things are my patches, which are very important | |
| 17:57:01 | efried | mriedem Sure, no worries. From this morning's meeting, you indicated it should all be done in the next month, and there's actually a nontrivial amount of code left to write (cinder, barbican, keystone) | |
| 17:57:06 | efried | dansmith No doubt. | |
| 17:57:34 | openstackgerrit | Merged openstack/nova master: Update docs to include standardization of VM diagnostics https://review.openstack.org/500408 | |
| 18:05:06 | sdague | mtreinish: http://logs.openstack.org/74/501874/2/gate/gate-nova-python35/0dcede7/console.html#_2017-09-21_18_02_47_155101 - is that an stestr issue? | |
| 18:05:21 | sdague | worker 7 just hung, and eventually that was a fail | |
| 18:11:51 | mriedem | stvnoyes: the bdm attachment ids stuff in the migrate_data object is going to get weird when we have multiattach | |
| 18:12:04 | mriedem | since in your change it's a 1:1 between volume id and attachment id | |
| 18:12:06 | mriedem | although, | |
| 18:12:23 | mriedem | i guess you wouldn't have a volume attached multiple times to the same instance, unless we're migrating it | |
| 18:12:27 | mriedem | so maybe i'm overthinking things | |
| 18:15:54 | openstackgerrit | Ildiko Vancsa proposed openstack/nova-specs master: Add multiattach support to Nova https://review.openstack.org/499777 | |
| 18:29:05 | mriedem | and just when i was about to -1 this change | |
| 18:38:14 | mriedem | dansmith: so for your 2nd change in the series to honor the global limit, | |
| 18:38:19 | mriedem | i was posting this comment, but can't | |
| 18:38:33 | mriedem | "So today we iterate the cells in order and just pull $limit instances and if we keep going to another cell we subtract how many we pulled out of the previous cell for a new limit on the next cell. Now we're going to be pulling $limit instances out of each cell concurrently, merge sorting them all and then enforcing the limit below. That is more traffic, right? Maybe it's negligible.", | |
| 18:38:34 | mriedem | "I was wondering if we wanted to divide the limit by the number of cells and then pull that many instances from each. So if we had 2 cells and the limit was 1000 (which is the default in the API), we'd just pull 500 instances from each cell and merge sort them." | |
| 18:44:03 | dansmith | mriedem: we can't do that partitioning of the limit until/unless we support going back to the cells for more instances | |
| 18:44:24 | dansmith | mriedem: otherwise we'd hit the end of one limit batch, assume there are no more in that cell that sort after anything in other cells and then get out of sync | |
| 18:44:39 | dansmith | mriedem: we basically have to have at least $limit results from each cell to prevent that | |
| 18:44:55 | dansmith | or emulate it by refilling our batch when we run out before we continue the merge | |
| 18:45:29 | dansmith | that distinction is the thing that made us think we couldn't do this easily initially and why we went down the path of punting on the problem | |
| 18:45:54 | dansmith | also, yes it may be more traffic at times, but something else I was discussing with people was: | |
| 18:46:33 | dansmith | if each tenant has N instances now, and then later those N instances are spread across a bunch of cells, this list operation is still pulling the same number of rows from the DB in total, but smaller pieces | |
| 18:46:51 | dansmith | the db driver doesn't even fetch partial results today or with this, it fetches everything in the result | |
| 18:47:27 | dansmith | does that make sense at all? | |
| 18:50:14 | mriedem | yeah, hadn't considered the case that there are like 200 instances in cell1 and 600 insteances in cell2, if we split the limit to 500 each, we'd end up with 700 instead of 800 | |
| 18:51:10 | dansmith | well, there's that, but we'd also be lossy in that we'd ignore things from cell1 that sorted before some in cell2, but just didn't make the limit cut | |
| 18:51:34 | dansmith | and if those were ever after max_limit/N cells, you'd never be able to limit them ever | |
| 18:51:56 | dansmith | er, never be able to list them | |