Earlier  
Posted Nick Remark
#openstack-nova - 2018-01-04
22:33:18 lyarwood hemna_: vms/3b97914e-3f9b-410a-b3d9-6c1a83244136_disk isn't a volume however right?
22:33:47 jaypipes efried: and instance group is an abomination otherwise called a server group. never mind, though... I'm gonna sleep on it and tackle manana
22:33:58 mriedem hemnahttps://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5957
22:34:06 hemna_ I'm not sure, this is a dump direct from the customer's env.
22:34:09 mriedem hemna_: there is a refresh_conn_info kwarg to _get_instance_block_device_info
22:34:11 efried jaypipes Okay. I'll play with it a bit if I get some time here.
22:34:40 mriedem so the idea from the ptg was just always pass refresh_conn_info=True there when we were going to do something like this
22:34:47 hemna_ mriedem, is that an option at nova cmdln time ?
22:34:52 mriedem no
22:35:04 mriedem it was a way to do this w/o any api changes
22:35:17 lyarwood mriedem: the issue isn't with the rbd volume, but the imagebackend rbd images
22:35:34 mriedem i don't know what that means
22:35:51 mriedem i thought the issue was stale rbd information in the connection_info, which nova gets from cinder
22:35:53 lyarwood mriedem: volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54 is the volume
22:35:56 mriedem and uses to populate the disk config
22:36:22 lyarwood mriedem: not for ephemerial rbd images, it's all hard coded from the local nova.conf iirc
22:36:37 mriedem lyarwood: yeah i'm not talking about ephemeral,
22:36:38 mriedem just volumes
22:37:13 mriedem this https://github.com/openstack/nova/blob/master/nova/virt/libvirt/volume/net.py#L56
22:37:22 lyarwood mriedem: yeah that appears to be updated in the LM flow
22:37:32 lyarwood mriedem: <source protocol='rbd' name='volumes/volume-6d04520d-0029-499c-af81-516a7ba37a54'> <-- this one is changed, new ips
22:37:44 mriedem right, because the live migration flow calls _get_volume_config
22:37:50 mriedem using the bdm.connection_info
22:37:53 mriedem which is stale
22:38:00 hemna_ yup
22:38:01 lyarwood yup
22:38:06 mriedem what we said in denver,
22:38:18 mriedem was when we start live migration, and get the bdms, we call cinder to refresh the connection info
22:38:37 mriedem by pass refresh_conn_info=True to https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L1683
22:38:52 mriedem that will make a new os-initialize_connection call to cinder and return the latest connection_info
22:39:07 mriedem then we save that on the bdm.connection_info
22:39:18 mriedem for reasons, 'ive just never written that patch,
22:39:26 mriedem but i also have concerns about the new attach flow making that no longer work
22:39:33 hemna_ :(
22:39:46 mriedem with the new flow, nova calls cinder to get the attachment record which has the connection_info stored in the cinder db,
22:39:58 mriedem so i don't think it actually creates a new connection (export?) so we wouldn't refresh
22:40:05 hemna_ which could be stale...
22:40:12 mriedem this https://github.com/openstack/nova/blob/master/nova/virt/block_device.py#L557
22:40:20 hemna_ in this case we just moved the stale info from nova to cinder, and cinder is stale?
22:40:39 mriedem hemna_: right, if cinder doesn't update that, and just returns whta's in the db, it would be stale - and we'd have the same problem as using the stale nova db info
22:40:43 mriedem yes
22:40:45 mriedem exactly
22:40:51 hemna_ bleh
22:42:22 mriedem yeah idk, nova could still call os-initialize_connection if we wanted to,
22:42:31 mriedem i don't know if that would update the attachment record in cinder or not
22:43:09 hemna_ so I thought the purpose of storing the conn info in cinder's db is for force delete time when nova doesn't know anything
22:43:12 mriedem otherwise we'd likely need like a refresh=True query parameter to GET /attachments/{id}
22:43:37 mriedem hemna_: that's part of it yeah
22:43:38 hemna_ in which case, can't cinder always refresh that if nova asks for it?
22:43:45 hemna_ I dunno
22:43:55 mriedem but how does cinder know that we're asking for a refresh?
22:43:58 mriedem or just a simple read-only get
22:44:13 mriedem this is where we got talking about API changes to force a refresh
22:44:13 hemna_ true
22:44:52 hemna_ I guess this could affect other cinder backends too
22:44:58 hemna_ not just ceph
22:46:05 hemna_ multipath IPs can change in between attaching the same volume, if the cinder backend gets changed
22:46:50 hemna_ ok so for now, I'll just tell my customer to bounce their VMs. :(
22:46:55 mriedem like a retype?
22:47:14 hemna_ well, more like a storage array loses one of it's interfaces
22:47:16 hemna_ and the IP changes
22:47:16 mriedem or you mean the backend backend
22:47:19 mriedem ok
22:47:20 hemna_ or new interfaces are added
22:47:39 hemna_ it's kinda the same thing as a new ceph monitor IP
22:48:11 lyarwood hemna_: FWIW restarting the instance is the only way to update the mon IPs for the ephemerial rbd images.
22:48:21 openstackgerrit Takashi NATSUME proposed openstack/nova master: [placement] Add sending global request ID in put (1) https://review.openstack.org/531258
22:48:29 lyarwood ephemeral*
22:48:41 hemna_ lyarwood, yah. they were trying to avoid that, as there could be tons of VMs to bounce
22:49:03 mriedem hmm, actually it looks like for live migration we do refresh the connection_info https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5882
22:49:10 mriedem in pre-live migration on the dest host
22:49:58 lyarwood right, the issue hemna_ reported in the bug isn't with rbd volumes :)
22:50:08 mriedem oh
22:50:30 mriedem now i get it
22:51:22 mriedem well, i'm reminded once again that i'd like to just give up and buy a farm in austria and just live out my days there
22:51:37 mriedem because this whole software business isn't worth it
22:51:41 hemna_ lolz
22:52:01 hemna_ you aren't the only one.....I've been in a youtube black whole, watching dudes build log cabins...
22:52:21 mriedem that's the manliest thing i've heard all day
22:58:49 hemna_ lyarwood, should I file a separate bug then to cover the rbd image update during LM ?
23:00:01 lyarwood hemna_: yeah I think so, it's a different codepath etc
23:00:09 hemna_ ok will do
23:00:25 lyarwood hemna_: cool thanks, I'll take a look tomorrow
23:06:47 openstackgerrit Eric Fried proposed openstack/nova master: Raise on API errors getting aggregates/traits https://review.openstack.org/526540
23:06:47 openstackgerrit Eric Fried proposed openstack/nova master: Track associated sharing RPs in report client https://review.openstack.org/526539
23:06:48 openstackgerrit Eric Fried proposed openstack/nova master: Track tree-associated providers in report client https://review.openstack.org/526541
23:06:48 openstackgerrit Eric Fried proposed openstack/nova master: ProviderTree.populate_from_iterable https://review.openstack.org/520756
23:06:49 openstackgerrit Eric Fried proposed openstack/nova master: WIP: ComputeDriver.update_provider_tree() https://review.openstack.org/521187
23:06:49 openstackgerrit Eric Fried proposed openstack/nova master: WIP: Scheduler[Report]Client.get_provider_tree https://review.openstack.org/521098
23:06:50 openstackgerrit Eric Fried proposed openstack/nova master: Fix nits in update_provider_tree series https://review.openstack.org/531260
23:06:50 openstackgerrit Eric Fried proposed openstack/nova master: WIP: Use update_provider_tree from resource tracker https://review.openstack.org/520246
23:07:15 efried jaypipes ^ -- and response to review comments on the way...
23:07:25 jaypipes coolio.
23:08:01 hemna_ lyarwood, https://bugs.launchpad.net/nova/+bug/1741364
23:08:02 openstack Launchpad bug 1741364 in OpenStack Compute (nova) "ceph ephemeral info not updated during live migrate" [Undecided,New]
23:11:54 efried jaypipes See response on https://review.openstack.org/#/c/526539/ -- not sure if I've answered the right question there.
23:26:09 mnaser does nova not set the instance status to ERROR if cinder volume detach fails?
23:26:39 mnaser I have some cinder volumes which show attached servers, but the server is deleted
23:31:10 mriedem mnaser: correct
23:31:40 mriedem normal volume detach? or volume detach during instance delete?

Earlier   Later