| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-01-04 | |||
| 22:07:19 | mriedem | just like a tag | |
| 22:07:20 | jaypipes | efried: in any case, meh, will hit it later and tomorrow but if you have any time, could use your eyeballs. | |
| 22:07:46 | jaypipes | efried: it will fail the assertion here: https://review.openstack.org/#/c/531243/1/nova/tests/functional/db/test_instance_group.py on line 351 | |
| 22:07:56 | mriedem | mdbooth: and that always tells us, the volume representing this bdm was attached and supported multiattach at that time, so treat it like that until it's detached and the bdm is deleted | |
| 22:22:08 | efried | jaypipes Sorry, I'm back now. Catching up... | |
| 22:23:46 | efried | jaypipes wtf is an instance group? | |
| 22:25:07 | efried | jaypipes And are these host aggregates (as opposed to placement aggregates)? | |
| 22:26:24 | hemna_ | mriedem, hey man, I added an update to bug https://bugs.launchpad.net/nova/+bug/1452641 | |
| 22:26:25 | openstack | Launchpad bug 1452641 in OpenStack Compute (nova) "Static Ceph mon IP addresses in connection_info can prevent VM startup" [Medium,Confirmed] | |
| 22:26:39 | hemna_ | mriedem, thought you might want to take a look and see what you thought. | |
| 22:30:34 | mriedem | yikes, super rbd specific code in there | |
| 22:31:37 | mriedem | hemna_: at the ptg in denver we said we'd just always refresh the connection_info when we needed it http://lists.openstack.org/pipermail/openstack-dev/2017-September/122170.html | |
| 22:32:04 | hemna_ | yah, that was just a customer's hacked patch to get it to work | |
| 22:32:22 | lyarwood | that's imagebackend rbd btw | |
| 22:32:29 | lyarwood | not connection_info volume rbd | |
| 22:32:30 | mriedem | sure. the forced refresh_connection_info=True thing would arguably be much simpler, if it actually works | |
| 22:32:53 | hemna_ | my customer tried to live migrate and it didn't get the updated info, hence his hack | |
| 22:33:18 | lyarwood | hemna_: vms/3b97914e-3f9b-410a-b3d9-6c1a83244136_disk isn't a volume however right? | |
| 22:33:47 | jaypipes | efried: and instance group is an abomination otherwise called a server group. never mind, though... I'm gonna sleep on it and tackle manana | |
| 22:33:58 | mriedem | hemnahttps://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5957 | |
| 22:34:06 | hemna_ | I'm not sure, this is a dump direct from the customer's env. | |
| 22:34:09 | mriedem | hemna_: there is a refresh_conn_info kwarg to _get_instance_block_device_info | |
| 22:34:11 | efried | jaypipes Okay. I'll play with it a bit if I get some time here. | |
| 22:34:40 | mriedem | so the idea from the ptg was just always pass refresh_conn_info=True there when we were going to do something like this | |
| 22:34:47 | hemna_ | mriedem, is that an option at nova cmdln time ? | |
| 22:34:52 | mriedem | no | |
| 22:35:04 | mriedem | it was a way to do this w/o any api changes | |
| 22:35:17 | lyarwood | mriedem: the issue isn't with the rbd volume, but the imagebackend rbd images | |
| 22:35:34 | mriedem | i don't know what that means | |
| 22:35:51 | mriedem | i thought the issue was stale rbd information in the connection_info, which nova gets from cinder | |
| 22:35:53 | lyarwood | mriedem: volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54 is the volume | |
| 22:35:56 | mriedem | and uses to populate the disk config | |
| 22:36:22 | lyarwood | mriedem: not for ephemerial rbd images, it's all hard coded from the local nova.conf iirc | |
| 22:36:37 | mriedem | lyarwood: yeah i'm not talking about ephemeral, | |
| 22:36:38 | mriedem | just volumes | |
| 22:37:13 | mriedem | this https://github.com/openstack/nova/blob/master/nova/virt/libvirt/volume/net.py#L56 | |
| 22:37:22 | lyarwood | mriedem: yeah that appears to be updated in the LM flow | |
| 22:37:32 | lyarwood | mriedem: <source protocol='rbd' name='volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54'> <-- this one is changed, new ips | |
| 22:37:44 | mriedem | right, because the live migration flow calls _get_volume_config | |
| 22:37:50 | mriedem | using the bdm.connection_info | |
| 22:37:53 | mriedem | which is stale | |
| 22:38:00 | hemna_ | yup | |
| 22:38:01 | lyarwood | yup | |
| 22:38:06 | mriedem | what we said in denver, | |
| 22:38:18 | mriedem | was when we start live migration, and get the bdms, we call cinder to refresh the connection info | |
| 22:38:37 | mriedem | by pass refresh_conn_info=True to https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L1683 | |
| 22:38:52 | mriedem | that will make a new os-initialize_connection call to cinder and return the latest connection_info | |
| 22:39:07 | mriedem | then we save that on the bdm.connection_info | |
| 22:39:18 | mriedem | for reasons, 'ive just never written that patch, | |
| 22:39:26 | mriedem | but i also have concerns about the new attach flow making that no longer work | |
| 22:39:33 | hemna_ | :( | |
| 22:39:46 | mriedem | with the new flow, nova calls cinder to get the attachment record which has the connection_info stored in the cinder db, | |
| 22:39:58 | mriedem | so i don't think it actually creates a new connection (export?) so we wouldn't refresh | |
| 22:40:05 | hemna_ | which could be stale... | |
| 22:40:12 | mriedem | this https://github.com/openstack/nova/blob/master/nova/virt/block_device.py#L557 | |
| 22:40:20 | hemna_ | in this case we just moved the stale info from nova to cinder, and cinder is stale? | |
| 22:40:39 | mriedem | hemna_: right, if cinder doesn't update that, and just returns whta's in the db, it would be stale - and we'd have the same problem as using the stale nova db info | |
| 22:40:43 | mriedem | yes | |
| 22:40:45 | mriedem | exactly | |
| 22:40:51 | hemna_ | bleh | |
| 22:42:22 | mriedem | yeah idk, nova could still call os-initialize_connection if we wanted to, | |
| 22:42:31 | mriedem | i don't know if that would update the attachment record in cinder or not | |
| 22:43:09 | hemna_ | so I thought the purpose of storing the conn info in cinder's db is for force delete time when nova doesn't know anything | |
| 22:43:12 | mriedem | otherwise we'd likely need like a refresh=True query parameter to GET /attachments/{id} | |
| 22:43:37 | mriedem | hemna_: that's part of it yeah | |
| 22:43:38 | hemna_ | in which case, can't cinder always refresh that if nova asks for it? | |
| 22:43:45 | hemna_ | I dunno | |
| 22:43:55 | mriedem | but how does cinder know that we're asking for a refresh? | |
| 22:43:58 | mriedem | or just a simple read-only get | |
| 22:44:13 | mriedem | this is where we got talking about API changes to force a refresh | |
| 22:44:13 | hemna_ | true | |
| 22:44:52 | hemna_ | I guess this could affect other cinder backends too | |
| 22:44:58 | hemna_ | not just ceph | |
| 22:46:05 | hemna_ | multipath IPs can change in between attaching the same volume, if the cinder backend gets changed | |
| 22:46:50 | hemna_ | ok so for now, I'll just tell my customer to bounce their VMs. :( | |
| 22:46:55 | mriedem | like a retype? | |
| 22:47:14 | hemna_ | well, more like a storage array loses one of it's interfaces | |
| 22:47:16 | hemna_ | and the IP changes | |
| 22:47:16 | mriedem | or you mean the backend backend | |
| 22:47:19 | mriedem | ok | |
| 22:47:20 | hemna_ | or new interfaces are added | |
| 22:47:39 | hemna_ | it's kinda the same thing as a new ceph monitor IP | |
| 22:48:11 | lyarwood | hemna_: FWIW restarting the instance is the only way to update the mon IPs for the ephemerial rbd images. | |
| 22:48:21 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: [placement] Add sending global request ID in put (1) https://review.openstack.org/531258 | |
| 22:48:29 | lyarwood | ephemeral* | |
| 22:48:41 | hemna_ | lyarwood, yah. they were trying to avoid that, as there could be tons of VMs to bounce | |
| 22:49:03 | mriedem | hmm, actually it looks like for live migration we do refresh the connection_info https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5882 | |
| 22:49:10 | mriedem | in pre-live migration on the dest host | |
| 22:49:58 | lyarwood | right, the issue hemna_ reported in the bug isn't with rbd volumes :) | |
| 22:50:08 | mriedem | oh | |
| 22:50:30 | mriedem | now i get it | |
| 22:51:22 | mriedem | well, i'm reminded once again that i'd like to just give up and buy a farm in austria and just live out my days there | |
| 22:51:37 | mriedem | because this whole software business isn't worth it | |
| 22:51:41 | hemna_ | lolz | |
| 22:52:01 | hemna_ | you aren't the only one.....I've been in a youtube black whole, watching dudes build log cabins... | |
| 22:52:21 | mriedem | that's the manliest thing i've heard all day | |
| 22:58:49 | hemna_ | lyarwood, should I file a separate bug then to cover the rbd image update during LM ? | |
| 23:00:01 | lyarwood | hemna_: yeah I think so, it's a different codepath etc | |
| 23:00:09 | hemna_ | ok will do | |
| 23:00:25 | lyarwood | hemna_: cool thanks, I'll take a look tomorrow | |