| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-09 | |||
| 12:24:48 | sean-k-mooney | but if im not at my work laptop i somethimes use teh eleemnt client form my personal laptop or ipad to chat via teh matrix.org bridge | |
| 12:25:44 | sean-k-mooney | so ya if you keep messanges relitivly short (3-4 lines) it works fine i havent hit the lenght limit personlly | |
| 12:26:10 | sean-k-mooney | of the irc alternivies i have used matrix is really the only one i tollerate | |
| 12:27:16 | sean-k-mooney | if the element desktop clinet ever get the ablity to sign into two matix accounts at once it might even be something i would consider as a replacemnt for weechat | |
| 12:27:22 | frickler | sean-k-mooney: there have also been issues where the bridge disconnects but you do not notice on the matrix side, so my personal suggest is still to not use this, ymmv | |
| 12:29:54 | sean-k-mooney | ya i have not had that issue but i still use irc as my primary interface and matix as what i use when im traveling or not working from my normal location | |
| 12:30:12 | sean-k-mooney | so i porably would not notice if there were tempoiry issues | |
| 12:33:36 | admin1 | i have a vm which is always in a pause state in the hypervisor .. trying to unpause using virsh gives error: Timed out during operation: cannot acquire state change lock (held by monitor=remoteDispatchDomainCreateWithFlags) .. the vm is backed by volume on ceph, but ceph is fine and there are no locks | |
| 12:33:45 | admin1 | what can i do to check/troubleshoot this issue | |
| 12:33:56 | admin1 | i rebooted the hypervisor as well, no luck | |
| 12:34:50 | sean-k-mooney | this might be a lock crated by qemu | |
| 12:35:06 | sean-k-mooney | have you tried stopping the vm and staring it | |
| 12:35:14 | sean-k-mooney | e.g. via a hard reboot | |
| 12:35:47 | admin1 | when i do a vrish destroy, it disappears from virsh list --all | |
| 12:36:06 | admin1 | when i start again (horizon/cli) appears back | |
| 12:36:09 | admin1 | with a paused state | |
| 12:36:55 | sean-k-mooney | ack | |
| 12:37:19 | sean-k-mooney | did you check the qemu instance log for any errors | |
| 12:37:35 | sean-k-mooney | this does not sound like a nova issue by the way | |
| 12:37:55 | sean-k-mooney | this sound like an issue at the qemu/libvirt level and or perhaps the ceph interaction | |
| 12:38:19 | sean-k-mooney | you dont happen to have a kvm error in the instance log do you? | |
| 12:38:59 | sean-k-mooney | we hit an issue with ubutu 22.04 where libvirt incorrectly detected the cpu model | |
| 12:39:11 | sean-k-mooney | it enabled amd cpu flags in the domain on an intel host | |
| 12:39:25 | sean-k-mooney | that left teh vm in a paused state | |
| 12:39:59 | sean-k-mooney | although that would not expaling the lock message but i woudl check the qemu instance log in anycase | |
| 12:40:13 | admin1 | sean-k-mooney thanks . i know what to check for now | |
| 12:40:51 | admin1 | where does qemu/libvirt read the ceph connectioon details like mon addresses ? | |
| 12:40:57 | admin1 | from /etc/ceph/ceph.conf ? | |
| 12:41:51 | admin1 | or is it internally somewhere else | |
| 12:43:43 | sean-k-mooney | we get them form the cinder attachment connection info and then store them in our db and pass it to libvirt | |
| 12:43:57 | sean-k-mooney | so no not from the ceph.conf | |
| 12:45:10 | sean-k-mooney | in recent release of openstack (xena+) we have a nova manage command to refresh the atachment info | |
| 12:45:12 | sean-k-mooney | https://docs.openstack.org/nova/latest/cli/nova-manage.html#volume-attachment-refresh | |
| 12:45:29 | admin1 | this one is not xena yet | |
| 12:45:40 | admin1 | i want to remove 2x mons and use only 1 remaining mon | |
| 12:45:46 | admin1 | how do I update/edit this db ? | |
| 12:46:00 | sean-k-mooney | with great pain and care | |
| 12:46:21 | sean-k-mooney | so we added this command to nova-manage because this is sotred in a json blob in the db | |
| 12:46:41 | sean-k-mooney | while it can be modifed its a pain to do | |
| 12:47:15 | sean-k-mooney | admin1: one option woudl be to grab a xena contaiern or create a xena virtual env and just run nova manage | |
| 12:47:58 | sean-k-mooney | i belive this is implemented such that if you have the new version of nova manage and point it to an old cloud it can work but im not 100% certin of that | |
| 12:48:00 | admin1 | you mean have binaries of xena but connect to existing db to manage/manipulate the entries ? | |
| 12:48:08 | sean-k-mooney | ya | |
| 12:48:24 | sean-k-mooney | so bauzas gibi correct me if im wrong be we have had customer do that right^ | |
| 12:48:57 | sean-k-mooney | use the updated contaienr with this command ot repair old dbs when connection infor is out of date | |
| 12:49:34 | sean-k-mooney | admin1: i think we have a downstream backport of this by the way to some release which is why im not 100% sure how we used this downstream with train | |
| 12:51:45 | sean-k-mooney | admin1: ya so we have it backported downstream to train in our 16.2 product | |
| 12:52:15 | sean-k-mooney | and i think we have had custoemr use the 16.2 contaienr to fix this on queens/osp 13 | |
| 12:52:19 | admin1 | i am on osa tag 23.1.2 | |
| 12:53:25 | admin1 | wallaby | |
| 12:54:00 | sean-k-mooney | we cannot backport db/object/rpc change even downstream so the fact it works on train implies this is very self contaiend meanign you should be able to use it with wallaby | |
| 12:54:47 | admin1 | : invalid choice: 'volume_attachment' on this | |
| 12:54:56 | admin1 | i have to boot a new container, point to the existing one and try from there | |
| 12:55:21 | sean-k-mooney | yep | |
| 13:46:02 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806 | |
| 13:46:03 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055 | |
| 13:55:54 | dvo-plv | Hello, everyone, Could tou please review our comments on the next blueprint: https://review.opendev.org/c/openstack/nova-specs/+/859290 | |
| 14:19:50 | admin1 | sean-k-mooney,is this a libvirt-secrets-gone thing or a ceph thing ? https://gist.githubusercontent.com/a1git/67cc7dab45f9bff536296670ab6ce65d/raw/450a2e31125e01826a90b1f133bc9b4821f807e8/gistfile1.txt | |
| 14:20:15 | admin1 | my mons were deleted completely .. i recreated those from osds | |
| 14:20:25 | admin1 | and cinder client was added with the same keys | |
| 14:20:33 | admin1 | most vms started, a few come with this error | |
| 14:22:46 | sean-k-mooney | did the mon ips change | |
| 14:23:18 | sean-k-mooney | presumable the secret is the same | |
| 14:23:50 | sean-k-mooney | but ya it could be that either the secret is msisign or the user aut info changed | |
| 14:23:59 | sean-k-mooney | the sechre has the ceph keyring inside | |
| 14:24:33 | sean-k-mooney | i dont have a ceph deployment to check but i belive that is tied a a spcific pool/user uuid | |
| 14:24:55 | sean-k-mooney | im not really shoudl how you recover on teh cecph side form all mons going away | |
| 14:25:24 | sean-k-mooney | but if any of the uuid chaged then you might need to get new keyrings and update the secret | |
| 14:26:03 | sean-k-mooney | changing the mon ips is not supproted in an openstack env since it requried bd surgery to fix | |
| 14:26:07 | admin1 | the mons were gone totally, but the ips did not change | |
| 14:26:13 | admin1 | all 3 mons are back in qorum | |
| 14:26:21 | admin1 | ceph is healthy and most of the vms started OK | |
| 14:26:29 | sean-k-mooney | ok the ips are cluster uuid and secrete are teh main things | |
| 14:26:30 | admin1 | there are 2 i know of that show this behaviour | |
| 14:26:32 | admin1 | there is no lock | |
| 14:27:00 | sean-k-mooney | so if the ips are the same no need to update the nova db unless the cluster id changed | |
| 14:27:15 | admin1 | cluster name /fsid all is same | |
| 14:27:23 | sean-k-mooney | ya fsid was what i ment | |
| 14:27:29 | admin1 | fsid is the same | |
| 14:27:52 | sean-k-mooney | so if that the same then provide the keyring in the secret is still valid you are porably ok | |
| 14:28:18 | sean-k-mooney | have you tried using that to list the volumes on the pool | |
| 14:32:03 | admin1 | got it | |
| 15:50:26 | clarkb | melwitt: thanks. It does look like bfv is tested, but it does also appear that the image used for rescuing is modified to set its bus and device types? I wonder if that is what we are missing here. The rescue command itself doesn't appear to take those arguments so these would need to be specified before hand on a special image? I guess that lends more weight to having | |
| 15:50:28 | clarkb | dedicated images as a required part of the rescue process? | |
| 16:07:41 | melwitt | clarkb: can be image properties or expressed as extra specs in the flavor. I can test out a change that would add a flavor create setting bus and device types in the test | |
| 16:09:02 | clarkb | melwitt: either way it is something that the cloud or cloud user would need to be aware of. Currently the default rescue behavior is to reuse the same image the rescued node booted off of. This is problematic because of the label specifier collisions, but also because if the image itself is broken you'd still be broken in a rescue. That leads users to using another image, but | |
| 16:09:04 | clarkb | there isn't any clear indication to me as a user that I need to use a special image. | |
| 16:10:15 | opendevreview | Dan Smith proposed openstack/nova master: Test ceph-multistore with a real image https://review.opendev.org/c/openstack/nova/+/860864 | |
| 16:11:05 | clarkb | I suspect the solution here is to make it clear to cloud operators that rescue has requirements x y z (I don't know what they all are yet) and that they should provide an image that meets those requirements | |
| 16:11:57 | melwitt | clarkb: hm yeah you are probably right it's only image properties, this section doesn't mention using a flavor to do it https://docs.openstack.org/nova/latest/user/rescue.html#stable-device-instance-rescue | |
| 16:12:42 | clarkb | also I wonder if nova should drop the default behavior or reusing the running image and instead force people to explicitly provide one | |
| 16:13:07 | clarkb | I suspect there are scenarios where reusing the image would work, but in the vast majority it seems unlikely | |
| 16:13:24 | clarkb | and that would help provide signal that something different is required here | |
| 16:16:46 | melwitt | I think we could do that in a new API microversion to avoid breaking anyone who is using it the old way and succeeding ... but the fact that openstackclient defaults to lowest microversion makes it more difficult to signal imho | |
| 16:18:27 | clarkb | ya and I think users could manually specify the same image if they really did need/want that | |
| 16:18:38 | clarkb | it just wouldn't be provided as a dfeault (which I think users expect to work) | |
| 16:19:27 | melwitt | yeah, I think that makes sense | |
| 17:19:55 | opendevreview | Merged openstack/nova stable/yoga: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/861872 | |
| 17:31:39 | opendevreview | Amit Uniyal proposed openstack/nova stable/ussuri: add regression test case for bug 1978983 https://review.opendev.org/c/openstack/nova/+/862603 | |
| 17:31:40 | opendevreview | Amit Uniyal proposed openstack/nova stable/ussuri: For evacuation, ignore if task_state is not None https://review.opendev.org/c/openstack/nova/+/862604 | |