Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-21
18:59:21 sean-k-mooney (source host)
18:59:51 mnaser now if you’re using rbd for images_type and you have 2 clusters with each az using different cluster
18:59:51 mnaser now if you’re using rbd for images_type and you have 2 clusters with each az using different cluster
19:00:23 mnaser And you do a live migrate and end up with vm running on the other side and but using it’s original storage
19:00:23 mnaser And you do a live migrate and end up with vm running on the other side and but using it’s original storage
19:00:27 sean-k-mooney then you need to configure your filters to ensure that you target the vsm to spcific cluster using a flavor or simialr
19:00:45 mnaser And then on resize ops it blows up horribly because it’s trying to use the destination cluster id
19:00:45 mnaser And then on resize ops it blows up horribly because it’s trying to use the destination cluster id
19:01:11 sean-k-mooney yep that operator error if they did not configure things properly to prevent this
19:01:37 sean-k-mooney adressing theses usecase is somethign that could be done but it would be a feature not a bug
19:01:37 mnaser How? So if you have 3 azs you create 3 flavors?
19:01:43 sean-k-mooney yep
19:01:53 mnaser Do you think that’s user friendly at all
19:01:53 mnaser Do you think that’s user friendly at all
19:02:13 sean-k-mooney nope but its how its currently desigined
19:02:23 sean-k-mooney and fixing it would not eb a bug fix
19:02:41 mnaser So really what you’re saying is nova will do live migrations that will break your vm
19:02:41 mnaser So really what you’re saying is nova will do live migrations that will break your vm
19:02:45 mnaser And that’s not a bug
19:02:45 mnaser And that’s not a bug
19:02:58 sean-k-mooney nope
19:03:19 sean-k-mooney nova check if it can attach the volcumes to the select host before it live migrates
19:03:27 sean-k-mooney so it will pass the schduler but fail in pre live migrate
19:03:33 mnaser ok, lets put that aside and talk about the users who use images_type=rbd
19:03:36 mnaser with different az's
19:03:44 mnaser it will break thoes vms
19:03:45 sean-k-mooney also live migrate is an admin only api and we allow you as an admin to select the host
19:04:06 mnaser ok when we're deploying openstack for customers they don't expect to sit and decide which host they are going to move things into at scale
19:04:15 sean-k-mooney mnaser: not if you use cross_az_atch=false
19:04:27 mnaser if i tell them 'sorry, openstack is kinda silly, it picks the wrong hosts, you just pick the right host yourself instead'
19:04:53 mnaser non-bfv, images_type=rbd, 2 az's with ceph cluster each will result in broken live migrations
19:04:54 sean-k-mooney if you want to propsoe a new feature for this im open to review that
19:05:16 sean-k-mooney what i do not think woudl be corerct it considerign this a bug when we previously declared it out of scope and backproting this
19:05:53 sean-k-mooney mnaser: it would break if the ceph cluster was inaccable yes
19:06:09 sean-k-mooney although i belvie
19:06:20 sean-k-mooney the vm would stay runnign on the souce host in active
19:06:24 sean-k-mooney with the migration in error
19:06:30 mnaser and any reasonable operator would make a sane assumption that the cloud would not live migrate across az's
19:06:42 sean-k-mooney libvirt will detect teh qemu instance was not able to connect
19:06:44 mnaser the vm does migrate if the cluster is accessible, and then all further operations like resize/migrate/etc are broken
19:06:47 sean-k-mooney and it shoudl abort the migration
19:06:57 mnaser so it goes into a user-facing broken state
19:07:05 sean-k-mooney az are not fault domain
19:07:12 sean-k-mooney or isolated segments
19:07:36 sean-k-mooney mnaser: i do not belive you will get into a user facing broken state for live migration
19:07:58 mnaser you will.. if both ceph clusters are accessible, then the further operations will try to use the fsid of the target vm
19:08:12 mnaser i can ask to get tracebacks and logs from teh customer
19:08:39 mnaser but it makes sense since now its trying to use the _new_ cluster fsid, but doesnt find the volume, since its attached from the old cluster fsid
19:08:43 sean-k-mooney if both are accsabel and you only have ceph cred for one of them on the compute host then qemu will not be able to conenct
19:09:24 sean-k-mooney mnaser: that sound like they are trying to use the same user/keyring between both clusters
19:09:29 mnaser ok, assume one cluster with different pools when you're using ceph then
19:09:31 sean-k-mooney which is incorect
19:09:33 mnaser i havent dug that deep into their stuff
19:10:03 mnaser now when nova tries to do things it'll do it on the new pool but cant find that _disk image
19:10:44 sean-k-mooney which will fail when we try to create the qemu instance on the dest
19:10:49 sean-k-mooney but the migraiton shoudl abort then
19:11:03 mnaser isnt the old xml get transferred
19:11:05 sean-k-mooney and the vm shoudl stay runing on the souce node in actie
19:11:06 mnaser so it successfully completes?
19:11:16 mnaser s/isnt/doesnt/
19:11:23 sean-k-mooney no the vm get created really really early on the dest
19:11:34 mnaser i dont think we rebuild xml from scratch on target but rather rely on shipping the xml from the old libvirt to the new one?
19:11:36 sean-k-mooney we have to create the vm on the dest so that the ram can be copied
19:11:54 sean-k-mooney mnaser: we generate a new xml on the souce for the dest
19:12:03 mnaser ok something is not adding up then
19:12:08 sean-k-mooney so my expectation is that it shoudl use the old cluster
19:12:19 sean-k-mooney so you woudl have cross az traffic
19:12:22 mnaser oh ok right yes, it would add up nevermind
19:12:29 mnaser if we generate xml on source for the dest it'll have the old
19:12:30 sean-k-mooney what might break is a hard reboot after that
19:12:38 mnaser yes exactly, or resize, etc
19:13:00 sean-k-mooney right but thats a complete differnt issue
19:13:13 sean-k-mooney we do not supprot move operations across diffent stroagge backends at all
19:13:30 sean-k-mooney and preventing that is left to the operator today and it has alyas been that way in nova
19:13:52 mnaser so as someone whos trying to get people to use openstack, giving them a big gun to shoot themselves in the foot
19:14:12 mnaser and then when they do that because it doesnt seem very trivial and obvious that what they did is wrong
19:14:18 sean-k-mooney mnaser: the simpelr approch si to use cells
19:14:20 mnaser when they went ahead, created az, aggregates, etc
19:14:36 sean-k-mooney we do not allwo cross cell live migration
19:14:37 mnaser that's a really good point
19:14:57 mnaser so ensure same storage backend inside a cell
19:15:03 mnaser seems like pretty sane advice
19:15:04 sean-k-mooney yes
19:15:16 sean-k-mooney with all that said we coudl work on a feature to adress this
19:15:45 sean-k-mooney but it would be a new feature and it would have to still allow usecase wehre cross az move operations make sense
19:16:14 sean-k-mooney mnaser: for example we recently added a similar feature for neutron routed networks
19:16:47 sean-k-mooney https://specs.openstack.org/openstack/nova-specs/specs/wallaby/implemented/routed-networks-scheduling.html
19:16:51 mnaser sometimes i really feel letting users create az's was a massive mistake lol
19:17:02 sean-k-mooney well users cant
19:17:04 mnaser it was always so loose and there's so many people who get shot in the foot with it
19:17:08 mnaser nah i mean from an operator perspective
19:17:11 sean-k-mooney its admin only unless you change the policy
19:17:26 mnaser people build out something and then it almost never gives them what they want
19:17:29 sean-k-mooney oh well the issue is peopel consufe nova az with aws
19:17:35 sean-k-mooney and they are nothign like each other
19:17:55 mnaser yeah
19:18:04 sean-k-mooney so before wallaybe tehre was no schduler supprot for route l3 networks
19:18:07 mnaser aws has a strong presence so its natural to think of it that way
19:18:22 sean-k-mooney i.e. there was nothign preventing you form cold/live migrating to a host where that ip coudl not be routed
19:18:36 sean-k-mooney https://specs.openstack.org/openstack/nova-specs/specs/wallaby/implemented/routed-networks-scheduling.html added support for this

Earlier   Later