Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-21
18:10:15 sean-k-mooney bauzas: elodilles can we prioritise review of this if possibel https://review.opendev.org/c/openstack/nova/+/874547
18:10:30 opendevreview Takashi Natsume proposed openstack/placement master: Move implemented specs for Xena and Yoga release https://review.opendev.org/c/openstack/placement/+/853730
18:11:18 sean-k-mooney this will help us fix our downstream ci
18:13:16 bauzas done but I leave you +W as I don't have a lot of context
18:15:39 sean-k-mooney tl;dr we use bindep in our downstream jobs to install deps before runing tox but rhel 8 nolonger has python-devel
18:15:47 gmann dansmith: +W on 'mysql memory reduction flags for ceph job'
18:15:55 sean-k-mooney i wanted to check with elodilles to make sure they were ok with the stable-only change
18:16:00 dansmith gmann: cool
18:16:04 dansmith thanks
18:24:10 dansmith gmann: oh jeez, I didn't realize the mysql periodic thing hadn't landed yet
18:24:32 dansmith gmann: do you think we should wait for that to soak for a bit?
18:24:47 dansmith I know the devstack one did, and I guess I misread that the tempest one hadn't yet
18:26:05 gmann dansmith: I also did not realize it when I checked that patch. but I think it is ok to enable it in ceph job and see.
18:26:25 gmann we can always revert it if it fail and make delay things during release time
18:26:34 dansmith okay, that's my preference too
18:44:46 mnaser i got a fun one. it looks like by default nova saves the az of the vm in the cell db, but it doesn't update the request_spec, but when we do migrations, we pass the request_spec to the scheduler (which contains az=null) which then moves you from one az to another in the migration
18:45:05 mnaser since .. https://github.com/openstack/nova/blob/90e2a5e50fbf08e62a1aedd5e176845ee22d96c9/nova/scheduler/request_filter.py#L138-L166 checks for request_spec az
18:45:40 sean-k-mooney mnaser: this was changed recently
18:45:56 mnaser this is in a scenario where an operator wants to make vms stick to their az if a user doesnt specify one
18:46:22 sean-k-mooney right so we spent a lot fo time trying to decide what the sematics shoudl be
18:46:56 sean-k-mooney im trying to find the spec
18:47:32 sean-k-mooney https://specs.openstack.org/openstack/nova-specs/specs/zed/implemented/unshelve-to-host.html
18:47:38 sean-k-mooney i guess this was for unshleve
18:48:05 sean-k-mooney mnaser: we epect tht the request spec would not have the az by the way if the user did not request one
18:48:44 mnaser makes sense cause that's their request
18:49:12 mnaser i understand it might not be everyone that wants this, but maybe for live migration use case it can cause issues if nova ends up doing cross-az migrations
18:49:54 sean-k-mooney for move operations where we supprot specifying a AZ it would be ok in some cases to set it in the request spec
18:50:13 sean-k-mooney mnaser: but we would want to have the same beahivor as in the unselve spec
18:50:29 sean-k-mooney i dont recal if we fixed the other move operatiosn to be consitent with that when we did this
18:50:35 sean-k-mooney Uggla: do you recall
18:51:36 sean-k-mooney mnaser: wiht unshelve to a specific az if you set it in the unshelve request and it was not set in the orgainl request spec it will be set after
18:52:52 sean-k-mooney mnaser: live migraiton does not currently supprot an az
18:52:55 mnaser sean-k-mooney: esentailly im thinking this is where this can be changed https://github.com/openstack/nova/blob/f01a90ccb85ab254236f84009cd432d03ce12ebb/nova/compute/api.py#L5499-L5500
18:53:20 mnaser cause live migrating from one az to another could pretty much fail, and we can just have it as an option i guess if we dont want to change default behaviour
18:53:22 sean-k-mooney nor does migrate
18:53:35 mnaser in most worlds migrate or live migrate will fail across az's
18:53:40 mnaser esp if you're using different storage backends for example
18:53:43 sean-k-mooney mnaser: this would be an api change and need a spec
18:54:21 sean-k-mooney in general live migration betwen AZ will either work on not work depending on yoru deployment. in general i would expect it to work in most cases
18:55:17 sean-k-mooney it just comes down to if you have exchanged ssh keys such that the hyperviors can comunicate and if you are using az with cinder or not
18:55:25 sean-k-mooney and cross_az attach
18:56:00 mnaser Maybe we can make it so that if cross az attach = false then it would update the request spec to match?
18:56:07 sean-k-mooney no
18:56:15 sean-k-mooney no config drvent api behavior
18:56:20 sean-k-mooney this is not a bug
18:56:38 sean-k-mooney if we want to supprot move operation to target an AZ or change the request spec this is an api change
18:56:54 mnaser No it’s not to target an AZ
18:56:54 mnaser No it’s not to target an AZ
18:57:17 sean-k-mooney i know you want to prefer to keep affinity
18:57:18 mnaser it’s for that it stays in the same AZ, or otherwise the live migration will fail
18:57:18 mnaser it’s for that it stays in the same AZ, or otherwise the live migration will fail
18:57:23 sean-k-mooney liek a weigher or filter
18:57:32 sean-k-mooney however that is not what the end user asked for
18:58:21 mnaser if nova allows you to live migrate from one az to another for a vm with cross_az_attach set to false is that a bug ?
18:58:45 sean-k-mooney not a schduler bug
18:59:01 sean-k-mooney it will fail in pre-live-migrate
18:59:16 sean-k-mooney and the vm will stay in active on the host
18:59:21 sean-k-mooney (source host)
18:59:51 mnaser now if you’re using rbd for images_type and you have 2 clusters with each az using different cluster
18:59:51 mnaser now if you’re using rbd for images_type and you have 2 clusters with each az using different cluster
19:00:23 mnaser And you do a live migrate and end up with vm running on the other side and but using it’s original storage
19:00:23 mnaser And you do a live migrate and end up with vm running on the other side and but using it’s original storage
19:00:27 sean-k-mooney then you need to configure your filters to ensure that you target the vsm to spcific cluster using a flavor or simialr
19:00:45 mnaser And then on resize ops it blows up horribly because it’s trying to use the destination cluster id
19:00:45 mnaser And then on resize ops it blows up horribly because it’s trying to use the destination cluster id
19:01:11 sean-k-mooney yep that operator error if they did not configure things properly to prevent this
19:01:37 sean-k-mooney adressing theses usecase is somethign that could be done but it would be a feature not a bug
19:01:37 mnaser How? So if you have 3 azs you create 3 flavors?
19:01:43 sean-k-mooney yep
19:01:53 mnaser Do you think that’s user friendly at all
19:01:53 mnaser Do you think that’s user friendly at all
19:02:13 sean-k-mooney nope but its how its currently desigined
19:02:23 sean-k-mooney and fixing it would not eb a bug fix
19:02:41 mnaser So really what you’re saying is nova will do live migrations that will break your vm
19:02:41 mnaser So really what you’re saying is nova will do live migrations that will break your vm
19:02:45 mnaser And that’s not a bug
19:02:45 mnaser And that’s not a bug
19:02:58 sean-k-mooney nope
19:03:19 sean-k-mooney nova check if it can attach the volcumes to the select host before it live migrates
19:03:27 sean-k-mooney so it will pass the schduler but fail in pre live migrate
19:03:33 mnaser ok, lets put that aside and talk about the users who use images_type=rbd
19:03:36 mnaser with different az's
19:03:44 mnaser it will break thoes vms
19:03:45 sean-k-mooney also live migrate is an admin only api and we allow you as an admin to select the host
19:04:06 mnaser ok when we're deploying openstack for customers they don't expect to sit and decide which host they are going to move things into at scale
19:04:15 sean-k-mooney mnaser: not if you use cross_az_atch=false
19:04:27 mnaser if i tell them 'sorry, openstack is kinda silly, it picks the wrong hosts, you just pick the right host yourself instead'
19:04:53 mnaser non-bfv, images_type=rbd, 2 az's with ceph cluster each will result in broken live migrations
19:04:54 sean-k-mooney if you want to propsoe a new feature for this im open to review that
19:05:16 sean-k-mooney what i do not think woudl be corerct it considerign this a bug when we previously declared it out of scope and backproting this
19:05:53 sean-k-mooney mnaser: it would break if the ceph cluster was inaccable yes
19:06:09 sean-k-mooney although i belvie
19:06:20 sean-k-mooney the vm would stay runnign on the souce host in active
19:06:24 sean-k-mooney with the migration in error
19:06:30 mnaser and any reasonable operator would make a sane assumption that the cloud would not live migrate across az's
19:06:42 sean-k-mooney libvirt will detect teh qemu instance was not able to connect
19:06:44 mnaser the vm does migrate if the cluster is accessible, and then all further operations like resize/migrate/etc are broken
19:06:47 sean-k-mooney and it shoudl abort the migration
19:06:57 mnaser so it goes into a user-facing broken state
19:07:05 sean-k-mooney az are not fault domain

Earlier   Later