| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-02-21 | |||
| 19:06:09 | sean-k-mooney | although i belvie | |
| 19:06:20 | sean-k-mooney | the vm would stay runnign on the souce host in active | |
| 19:06:24 | sean-k-mooney | with the migration in error | |
| 19:06:30 | mnaser | and any reasonable operator would make a sane assumption that the cloud would not live migrate across az's | |
| 19:06:42 | sean-k-mooney | libvirt will detect teh qemu instance was not able to connect | |
| 19:06:44 | mnaser | the vm does migrate if the cluster is accessible, and then all further operations like resize/migrate/etc are broken | |
| 19:06:47 | sean-k-mooney | and it shoudl abort the migration | |
| 19:06:57 | mnaser | so it goes into a user-facing broken state | |
| 19:07:05 | sean-k-mooney | az are not fault domain | |
| 19:07:12 | sean-k-mooney | or isolated segments | |
| 19:07:36 | sean-k-mooney | mnaser: i do not belive you will get into a user facing broken state for live migration | |
| 19:07:58 | mnaser | you will.. if both ceph clusters are accessible, then the further operations will try to use the fsid of the target vm | |
| 19:08:12 | mnaser | i can ask to get tracebacks and logs from teh customer | |
| 19:08:39 | mnaser | but it makes sense since now its trying to use the _new_ cluster fsid, but doesnt find the volume, since its attached from the old cluster fsid | |
| 19:08:43 | sean-k-mooney | if both are accsabel and you only have ceph cred for one of them on the compute host then qemu will not be able to conenct | |
| 19:09:24 | sean-k-mooney | mnaser: that sound like they are trying to use the same user/keyring between both clusters | |
| 19:09:29 | mnaser | ok, assume one cluster with different pools when you're using ceph then | |
| 19:09:31 | sean-k-mooney | which is incorect | |
| 19:09:33 | mnaser | i havent dug that deep into their stuff | |
| 19:10:03 | mnaser | now when nova tries to do things it'll do it on the new pool but cant find that _disk image | |
| 19:10:44 | sean-k-mooney | which will fail when we try to create the qemu instance on the dest | |
| 19:10:49 | sean-k-mooney | but the migraiton shoudl abort then | |
| 19:11:03 | mnaser | isnt the old xml get transferred | |
| 19:11:05 | sean-k-mooney | and the vm shoudl stay runing on the souce node in actie | |
| 19:11:06 | mnaser | so it successfully completes? | |
| 19:11:16 | mnaser | s/isnt/doesnt/ | |
| 19:11:23 | sean-k-mooney | no the vm get created really really early on the dest | |
| 19:11:34 | mnaser | i dont think we rebuild xml from scratch on target but rather rely on shipping the xml from the old libvirt to the new one? | |
| 19:11:36 | sean-k-mooney | we have to create the vm on the dest so that the ram can be copied | |
| 19:11:54 | sean-k-mooney | mnaser: we generate a new xml on the souce for the dest | |
| 19:12:03 | mnaser | ok something is not adding up then | |
| 19:12:08 | sean-k-mooney | so my expectation is that it shoudl use the old cluster | |
| 19:12:19 | sean-k-mooney | so you woudl have cross az traffic | |
| 19:12:22 | mnaser | oh ok right yes, it would add up nevermind | |
| 19:12:29 | mnaser | if we generate xml on source for the dest it'll have the old | |
| 19:12:30 | sean-k-mooney | what might break is a hard reboot after that | |
| 19:12:38 | mnaser | yes exactly, or resize, etc | |
| 19:13:00 | sean-k-mooney | right but thats a complete differnt issue | |
| 19:13:13 | sean-k-mooney | we do not supprot move operations across diffent stroagge backends at all | |
| 19:13:30 | sean-k-mooney | and preventing that is left to the operator today and it has alyas been that way in nova | |
| 19:13:52 | mnaser | so as someone whos trying to get people to use openstack, giving them a big gun to shoot themselves in the foot | |
| 19:14:12 | mnaser | and then when they do that because it doesnt seem very trivial and obvious that what they did is wrong | |
| 19:14:18 | sean-k-mooney | mnaser: the simpelr approch si to use cells | |
| 19:14:20 | mnaser | when they went ahead, created az, aggregates, etc | |
| 19:14:36 | sean-k-mooney | we do not allwo cross cell live migration | |
| 19:14:37 | mnaser | that's a really good point | |
| 19:14:57 | mnaser | so ensure same storage backend inside a cell | |
| 19:15:03 | mnaser | seems like pretty sane advice | |
| 19:15:04 | sean-k-mooney | yes | |
| 19:15:16 | sean-k-mooney | with all that said we coudl work on a feature to adress this | |
| 19:15:45 | sean-k-mooney | but it would be a new feature and it would have to still allow usecase wehre cross az move operations make sense | |
| 19:16:14 | sean-k-mooney | mnaser: for example we recently added a similar feature for neutron routed networks | |
| 19:16:47 | sean-k-mooney | https://specs.openstack.org/openstack/nova-specs/specs/wallaby/implemented/routed-networks-scheduling.html | |
| 19:16:51 | mnaser | sometimes i really feel letting users create az's was a massive mistake lol | |
| 19:17:02 | sean-k-mooney | well users cant | |
| 19:17:04 | mnaser | it was always so loose and there's so many people who get shot in the foot with it | |
| 19:17:08 | mnaser | nah i mean from an operator perspective | |
| 19:17:11 | sean-k-mooney | its admin only unless you change the policy | |
| 19:17:26 | mnaser | people build out something and then it almost never gives them what they want | |
| 19:17:29 | sean-k-mooney | oh well the issue is peopel consufe nova az with aws | |
| 19:17:35 | sean-k-mooney | and they are nothign like each other | |
| 19:17:55 | mnaser | yeah | |
| 19:18:04 | sean-k-mooney | so before wallaybe tehre was no schduler supprot for route l3 networks | |
| 19:18:07 | mnaser | aws has a strong presence so its natural to think of it that way | |
| 19:18:22 | sean-k-mooney | i.e. there was nothign preventing you form cold/live migrating to a host where that ip coudl not be routed | |
| 19:18:36 | sean-k-mooney | https://specs.openstack.org/openstack/nova-specs/specs/wallaby/implemented/routed-networks-scheduling.html added support for this | |
| 19:19:11 | sean-k-mooney | it woudl not be unreasonable to have a similer feature for nova stoage | |
| 19:19:40 | sean-k-mooney | for example if we use the ceph fsid to create a placement aggrate containing all host that were configured to use that ceph cluster | |
| 19:19:58 | sean-k-mooney | and then recoded that in the isntance_system_metadata and schduled based on that if set | |
| 19:20:44 | sean-k-mooney | we would jsut need to do member_of=<fsid> in the pacement query | |
| 19:20:59 | mnaser | yeah, that seems like a handy simple way to track that for ceph | |
| 19:21:11 | sean-k-mooney | if rbd_fsid was in instance_system_metadata | |
| 19:21:34 | mnaser | i guess we would technically toss that into block device mapping data | |
| 19:21:44 | mnaser | i cant remember if nova uses that for its own storage | |
| 19:21:56 | sean-k-mooney | ish | |
| 19:22:02 | sean-k-mooney | we do in weird ways | |
| 19:22:10 | mnaser | maybe we should add a warning to the doc https://docs.openstack.org/nova/latest/admin/availability-zones.html about looking into using cells if you want to have full isolation and not allow migrations from one az to another | |
| 19:22:27 | sean-k-mooney | but this would be for the root disk really althoguh you could map cinder voluems to placment aggreats in a simialr way | |
| 19:23:32 | sean-k-mooney | bauzas: when you ahve time reading back over ^ would be good | |
| 19:23:48 | mnaser | ill push a PR to add some details about migrations and bring up cells | |
| 19:24:00 | sean-k-mooney | mnaser: cells are still not full isolation but ya. | |
| 19:24:11 | mnaser | i have to be honest in my ability of providing help, spec + new feature discussion + all that is a bit too far of a reach for this | |
| 19:24:12 | sean-k-mooney | mnaser: the other approch woudl be to have a weigher | |
| 19:24:19 | sean-k-mooney | so an az affinity weigher | |
| 19:24:29 | mnaser | hmm | |
| 19:24:33 | mnaser | i could do that out of tree i guess | |
| 19:24:42 | mnaser | as i dont think nova particlarly would wnat to carry that | |
| 19:25:06 | sean-k-mooney | we would need to pass the instance current cell to the sheduler and then the weigher could prefer to say in the same az | |
| 19:25:21 | sean-k-mooney | am i would not be against having it | |
| 19:25:57 | sean-k-mooney | we woudl need to modify the destination object and add a prefered az filed or something | |
| 19:26:55 | mnaser | i guess it can be a filter too but it would be very ugly | |
| 19:27:18 | mnaser | cause it would have to check if this is a reschedule (aka instance exists and we can find it) or first time (ignore) | |
| 19:27:27 | sean-k-mooney | well it should not be a filter because corss az move operations are valid | |
| 19:28:07 | mnaser | ah yes also addressing that | |
| 19:28:13 | sean-k-mooney | mnaser: basically we coudl add a "current_az" field here https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L1092-L1122 | |
| 19:29:19 | mnaser | this starts to enter the domain of requiring more resources/time than i have so trying to see how i can be them most useful with the little resource i can spend on this 😅 | |
| 19:29:19 | mnaser | this starts to enter the domain of requiring more resources/time than i have so trying to see how i can be them most useful with the little resource i can spend on this 😅 | |
| 19:29:21 | sean-k-mooney | thats used in a few places but we baskcialy woudl just need to get the instnace.az and pass it on | |
| 19:29:52 | sean-k-mooney | well simple solution is docs patch + ptg topic | |
| 19:30:15 | sean-k-mooney | and i can raise it as a "operator pain point" internally and see if there is interst in adressing it | |