| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-01-04 | |||
| 22:30:34 | mriedem | yikes, super rbd specific code in there | |
| 22:31:37 | mriedem | hemna_: at the ptg in denver we said we'd just always refresh the connection_info when we needed it http://lists.openstack.org/pipermail/openstack-dev/2017-September/122170.html | |
| 22:32:04 | hemna_ | yah, that was just a customer's hacked patch to get it to work | |
| 22:32:22 | lyarwood | that's imagebackend rbd btw | |
| 22:32:29 | lyarwood | not connection_info volume rbd | |
| 22:32:30 | mriedem | sure. the forced refresh_connection_info=True thing would arguably be much simpler, if it actually works | |
| 22:32:53 | hemna_ | my customer tried to live migrate and it didn't get the updated info, hence his hack | |
| 22:33:18 | lyarwood | hemna_: vms/3b97914e-3f9b-410a-b3d9-6c1a83244136_disk isn't a volume however right? | |
| 22:33:47 | jaypipes | efried: and instance group is an abomination otherwise called a server group. never mind, though... I'm gonna sleep on it and tackle manana | |
| 22:33:58 | mriedem | hemnahttps://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5957 | |
| 22:34:06 | hemna_ | I'm not sure, this is a dump direct from the customer's env. | |
| 22:34:09 | mriedem | hemna_: there is a refresh_conn_info kwarg to _get_instance_block_device_info | |
| 22:34:11 | efried | jaypipes Okay. I'll play with it a bit if I get some time here. | |
| 22:34:40 | mriedem | so the idea from the ptg was just always pass refresh_conn_info=True there when we were going to do something like this | |
| 22:34:47 | hemna_ | mriedem, is that an option at nova cmdln time ? | |
| 22:34:52 | mriedem | no | |
| 22:35:04 | mriedem | it was a way to do this w/o any api changes | |
| 22:35:17 | lyarwood | mriedem: the issue isn't with the rbd volume, but the imagebackend rbd images | |
| 22:35:34 | mriedem | i don't know what that means | |
| 22:35:51 | mriedem | i thought the issue was stale rbd information in the connection_info, which nova gets from cinder | |
| 22:35:53 | lyarwood | mriedem: volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54 is the volume | |
| 22:35:56 | mriedem | and uses to populate the disk config | |
| 22:36:22 | lyarwood | mriedem: not for ephemerial rbd images, it's all hard coded from the local nova.conf iirc | |
| 22:36:37 | mriedem | lyarwood: yeah i'm not talking about ephemeral, | |
| 22:36:38 | mriedem | just volumes | |
| 22:37:13 | mriedem | this https://github.com/openstack/nova/blob/master/nova/virt/libvirt/volume/net.py#L56 | |
| 22:37:22 | lyarwood | mriedem: yeah that appears to be updated in the LM flow | |
| 22:37:32 | lyarwood | mriedem: <source protocol='rbd' name='volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54'> <-- this one is changed, new ips | |
| 22:37:44 | mriedem | right, because the live migration flow calls _get_volume_config | |
| 22:37:50 | mriedem | using the bdm.connection_info | |
| 22:37:53 | mriedem | which is stale | |
| 22:38:00 | hemna_ | yup | |
| 22:38:01 | lyarwood | yup | |
| 22:38:06 | mriedem | what we said in denver, | |
| 22:38:18 | mriedem | was when we start live migration, and get the bdms, we call cinder to refresh the connection info | |
| 22:38:37 | mriedem | by pass refresh_conn_info=True to https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L1683 | |
| 22:38:52 | mriedem | that will make a new os-initialize_connection call to cinder and return the latest connection_info | |
| 22:39:07 | mriedem | then we save that on the bdm.connection_info | |
| 22:39:18 | mriedem | for reasons, 'ive just never written that patch, | |
| 22:39:26 | mriedem | but i also have concerns about the new attach flow making that no longer work | |
| 22:39:33 | hemna_ | :( | |
| 22:39:46 | mriedem | with the new flow, nova calls cinder to get the attachment record which has the connection_info stored in the cinder db, | |
| 22:39:58 | mriedem | so i don't think it actually creates a new connection (export?) so we wouldn't refresh | |
| 22:40:05 | hemna_ | which could be stale... | |
| 22:40:12 | mriedem | this https://github.com/openstack/nova/blob/master/nova/virt/block_device.py#L557 | |
| 22:40:20 | hemna_ | in this case we just moved the stale info from nova to cinder, and cinder is stale? | |
| 22:40:39 | mriedem | hemna_: right, if cinder doesn't update that, and just returns whta's in the db, it would be stale - and we'd have the same problem as using the stale nova db info | |
| 22:40:43 | mriedem | yes | |
| 22:40:45 | mriedem | exactly | |
| 22:40:51 | hemna_ | bleh | |
| 22:42:22 | mriedem | yeah idk, nova could still call os-initialize_connection if we wanted to, | |
| 22:42:31 | mriedem | i don't know if that would update the attachment record in cinder or not | |
| 22:43:09 | hemna_ | so I thought the purpose of storing the conn info in cinder's db is for force delete time when nova doesn't know anything | |
| 22:43:12 | mriedem | otherwise we'd likely need like a refresh=True query parameter to GET /attachments/{id} | |
| 22:43:37 | mriedem | hemna_: that's part of it yeah | |
| 22:43:38 | hemna_ | in which case, can't cinder always refresh that if nova asks for it? | |
| 22:43:45 | hemna_ | I dunno | |
| 22:43:55 | mriedem | but how does cinder know that we're asking for a refresh? | |
| 22:43:58 | mriedem | or just a simple read-only get | |
| 22:44:13 | mriedem | this is where we got talking about API changes to force a refresh | |
| 22:44:13 | hemna_ | true | |
| 22:44:52 | hemna_ | I guess this could affect other cinder backends too | |
| 22:44:58 | hemna_ | not just ceph | |
| 22:46:05 | hemna_ | multipath IPs can change in between attaching the same volume, if the cinder backend gets changed | |
| 22:46:50 | hemna_ | ok so for now, I'll just tell my customer to bounce their VMs. :( | |
| 22:46:55 | mriedem | like a retype? | |
| 22:47:14 | hemna_ | well, more like a storage array loses one of it's interfaces | |
| 22:47:16 | hemna_ | and the IP changes | |
| 22:47:16 | mriedem | or you mean the backend backend | |
| 22:47:19 | mriedem | ok | |
| 22:47:20 | hemna_ | or new interfaces are added | |
| 22:47:39 | hemna_ | it's kinda the same thing as a new ceph monitor IP | |
| 22:48:11 | lyarwood | hemna_: FWIW restarting the instance is the only way to update the mon IPs for the ephemerial rbd images. | |
| 22:48:21 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: [placement] Add sending global request ID in put (1) https://review.openstack.org/531258 | |
| 22:48:29 | lyarwood | ephemeral* | |
| 22:48:41 | hemna_ | lyarwood, yah. they were trying to avoid that, as there could be tons of VMs to bounce | |
| 22:49:03 | mriedem | hmm, actually it looks like for live migration we do refresh the connection_info https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5882 | |
| 22:49:10 | mriedem | in pre-live migration on the dest host | |
| 22:49:58 | lyarwood | right, the issue hemna_ reported in the bug isn't with rbd volumes :) | |
| 22:50:08 | mriedem | oh | |
| 22:50:30 | mriedem | now i get it | |
| 22:51:22 | mriedem | well, i'm reminded once again that i'd like to just give up and buy a farm in austria and just live out my days there | |
| 22:51:37 | mriedem | because this whole software business isn't worth it | |
| 22:51:41 | hemna_ | lolz | |
| 22:52:01 | hemna_ | you aren't the only one.....I've been in a youtube black whole, watching dudes build log cabins... | |
| 22:52:21 | mriedem | that's the manliest thing i've heard all day | |
| 22:58:49 | hemna_ | lyarwood, should I file a separate bug then to cover the rbd image update during LM ? | |
| 23:00:01 | lyarwood | hemna_: yeah I think so, it's a different codepath etc | |
| 23:00:09 | hemna_ | ok will do | |
| 23:00:25 | lyarwood | hemna_: cool thanks, I'll take a look tomorrow | |
| 23:06:47 | openstackgerrit | Eric Fried proposed openstack/nova master: Raise on API errors getting aggregates/traits https://review.openstack.org/526540 | |
| 23:06:47 | openstackgerrit | Eric Fried proposed openstack/nova master: Track associated sharing RPs in report client https://review.openstack.org/526539 | |
| 23:06:48 | openstackgerrit | Eric Fried proposed openstack/nova master: Track tree-associated providers in report client https://review.openstack.org/526541 | |
| 23:06:48 | openstackgerrit | Eric Fried proposed openstack/nova master: ProviderTree.populate_from_iterable https://review.openstack.org/520756 | |
| 23:06:49 | openstackgerrit | Eric Fried proposed openstack/nova master: WIP: ComputeDriver.update_provider_tree() https://review.openstack.org/521187 | |
| 23:06:49 | openstackgerrit | Eric Fried proposed openstack/nova master: WIP: Scheduler[Report]Client.get_provider_tree https://review.openstack.org/521098 | |
| 23:06:50 | openstackgerrit | Eric Fried proposed openstack/nova master: Fix nits in update_provider_tree series https://review.openstack.org/531260 | |
| 23:06:50 | openstackgerrit | Eric Fried proposed openstack/nova master: WIP: Use update_provider_tree from resource tracker https://review.openstack.org/520246 | |
| 23:07:15 | efried | jaypipes ^ -- and response to review comments on the way... | |
| 23:07:25 | jaypipes | coolio. | |