| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-03-09 | |||
| 14:43:00 | lyarwood | sean-k-mooney: there's nothing obvious in the logs but I wonder if it could be related to https://bugs.launchpad.net/cinder/+bug/1917750 | |
| 14:43:01 | openstack | Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New] | |
| 14:43:07 | sean-k-mooney | ~300 hits in 30 days | |
| 14:43:52 | sean-k-mooney | im seeing some libvirt issue on the contoler too not the node with teh detach issue | |
| 14:44:47 | lyarwood | do you have an example to hand? | |
| 14:45:19 | sean-k-mooney | am its hitting my vdpa pataches but also neutron let me get one | |
| 14:47:54 | sean-k-mooney | so ya https://review.opendev.org/c/openstack/nova/+/778350/4 https://zuul.opendev.org/t/openstack/build/fb643b53835341ac8589afeadfa7044d/logs | |
| 14:49:08 | sean-k-mooney | its showing up in neutron too https://review.opendev.org/c/openstack/neutron/+/777785 | |
| 14:49:50 | sean-k-mooney | in https://zuul.opendev.org/t/openstack/build/32f3dd64008b469eb9fb8b13ed33f137 | |
| 14:50:04 | sean-k-mooney | so i think this is just an issue with master in general | |
| 14:50:27 | sean-k-mooney | it could be related to https://bugs.launchpad.net/cinder/+bug/1917750 maybe havent looked at it yet | |
| 14:50:28 | openstack | Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New] | |
| 14:57:49 | lyarwood | sean-k-mooney: yeah I think it's related | |
| 15:01:19 | lyarwood | gah | |
| 15:01:23 | lyarwood | yeah it's that | |
| 15:02:15 | lyarwood | so we end up in a situation where one test attaches a volume to the host as /dev/sda from c-vol LVM/iSCSI backend #1 | |
| 15:02:39 | lyarwood | another test then attaches another volume with the same WWN to the hsot as /dev/sdb from c-vol LVM/iSCSI backend #2 | |
| 15:03:07 | lyarwood | the first test finishes and removes what it thinks is the first volume attachment | |
| 15:03:13 | lyarwood | but it's actually the second | |
| 15:03:36 | lyarwood | leaving libvirt unable to detach the device from the instance as I assume it can't flush | |
| 15:04:02 | lyarwood | I guess our current code swallows or ignores the failure from the libvirt and just retries? | |
| 15:04:27 | sean-k-mooney | porbaly ya | |
| 15:04:34 | sean-k-mooney | we dont use the events yet | |
| 15:04:51 | sean-k-mooney | so we need to revert running those n parrallel | |
| 15:05:11 | sean-k-mooney | although is this not a os-brick bug | |
| 15:05:26 | sean-k-mooney | i mean the concurrent attach should work | |
| 15:05:29 | lyarwood | yeah it's a job topology bug | |
| 15:05:58 | lyarwood | it `works` but the first test will end up sending I/O to the volume from the second test | |
| 15:06:21 | lyarwood | something has changed recently somewhere in the stack to allow both backends to return the same WWN tbh | |
| 15:06:27 | bauzas | stephenfin: apologies for forgetting that a string is immutable | |
| 15:06:37 | lyarwood | unless people have always missed these failures | |
| 15:06:37 | sean-k-mooney | lyarwood: would not not be possible thouhg in general | |
| 15:06:52 | sean-k-mooney | for them to have the same WWN | |
| 15:07:56 | lyarwood | not with a real world backend | |
| 15:08:37 | lyarwood | https://en.wikipedia.org/wiki/World_Wide_Name | |
| 15:12:19 | sean-k-mooney | lyarwood: ya i tought if you have iscsi portals from different sotrage backend thw WWN was only uniq within any given portal | |
| 15:13:26 | sean-k-mooney | lyarwood: i dont think cinder is choosing the WWN for the volume | |
| 15:16:41 | lyarwood | right yeah that would make sense | |
| 15:22:29 | stephenfin | bauzas: nw. I've replied to the rest of your comments also | |
| 15:42:15 | bauzas | yeah I'll review again | |
| 15:42:36 | gibi | lyarwood: do I understand correclty from the scrollback that making some cinder test serial would help merging patches? | |
| 15:42:55 | lyarwood | gibi: removing c-vol from the computes would | |
| 15:43:33 | lyarwood | gibi: if we really need to run another c-vol service then it needs to be on the controller alongside the original | |
| 15:43:41 | lyarwood | gibi: just building an env now to prove things | |
| 15:43:43 | stephenfin | lyarwood, gibi, sean-k-mooney: any aversion to adding mypy to pre-commit? | |
| 15:43:59 | sean-k-mooney | not speciriclly no | |
| 15:44:00 | stephenfin | assuming I can do it while respecting mypy-files | |
| 15:44:04 | lyarwood | stephenfin: against the files that are changing? | |
| 15:44:10 | gibi | lyarwood: thanks for looking into that, sign me up for review where there is something I can push | |
| 15:44:12 | sean-k-mooney | i have been bitten by flake8 passed by pep8 didnt | |
| 15:44:14 | stephenfin | yeah | |
| 15:44:30 | gibi | lyarwood: does having c-vol along with n-cpu is an invalid config? | |
| 15:44:33 | lyarwood | stephenfin: no issues assuming it isn't adding a huge amount of delay | |
| 15:44:34 | sean-k-mooney | i have gotten used to relying on pre-commit to do that form me | |
| 15:44:35 | stephenfin | sean-k-mooney: me too :( I ran it on HEAD but it turns out I broke something then fixed it in the next change | |
| 15:44:46 | stephenfin | me too * 2 :) | |
| 15:44:53 | stephenfin | pre-commit FTW | |
| 15:44:59 | gibi | stephenfin: no problem for me I dont use pre-commit :) | |
| 15:45:18 | stephenfin | gibi: You're missing out. It's pretty great :) | |
| 15:45:45 | lyarwood | gibi: no, the issue here is with the default c-vol backend, LVM/iSCSI, that you can't have >1 running version of in an env AFAICT | |
| 15:46:06 | gibi | stephenfin: I do run fast8 py39 and functional-py39 on every commit I push (except on old stable branches where there is no py39 or py38 support) | |
| 15:46:10 | bauzas | wife is sick and I need to taxi her to the doctor | |
| 15:46:10 | lyarwood | gibi: we end up mapping different volumes from different backends to the computes with the same WWN (world wide name) that should be unqiue | |
| 15:46:37 | gibi | lyarwood: so we cannot have multiple backend per compute? | |
| 15:46:50 | stephenfin | gibi: That's better than me. I run what I think are relevant tests and then let the CI do the rest. Our tests take too long to run locally | |
| 15:46:58 | lyarwood | gibi: we can't have multiple LVM/iSCSI backed c-vol's in the same env | |
| 15:47:22 | gibi | stephenfin: I have a beefy blade in a lab to run unit and func test (and devstack on baremetal to test sriov) | |
| 15:47:48 | lyarwood | it's not great but better than nothing | |
| 15:47:51 | stephenfin | I've a four year old laptop 0:) | |
| 15:48:01 | lyarwood | refresh is 3 soooooooooooo ;) | |
| 15:48:16 | gibi | still a laptop is slow, you should ask for a lab :) | |
| 15:49:40 | gibi | lyarwood: so in a real deployment there can only one c-vol service using the LVM/iSCSI backend? | |
| 15:49:59 | gibi | lyarwood: sorry that I'm slow to understant this :) | |
| 15:50:21 | lyarwood | gibi: I don't think that has ever been enforced but the LVM/iSCSI backend is only ever used for testing | |
| 15:50:48 | gibi | OK, then I stop worrying :) | |
| 15:50:53 | gibi | it is just test :) | |
| 15:52:19 | lyarwood | I'll just caveat all of the above with the fact that no one from Cinder has agreed with any of that in the bug as yet | |
| 15:52:26 | lyarwood | so I might be missing the point entirely | |
| 15:52:49 | lyarwood | but two volumes with the same WWN connected to the same host smells like something that will bork devicemapper | |
| 15:53:21 | gibi | if it fixes the CI and let us merge the api db compaction before FF then I'm happy to take the hit if we need to revert the change later | |
| 15:53:24 | gibi | :) | |
| 15:53:54 | lyarwood | gibi: is that stuck in a recheck loop? | |
| 15:54:01 | gibi | pretty much yes | |
| 15:54:04 | lyarwood | gah okay | |
| 15:54:11 | lyarwood | let me confirm and then I can push some changes | |
| 15:54:22 | gibi | the Queens one is recheckd through the whole weekend and still bouncing back | |
| 15:54:36 | gibi | mostly with the detach issue | |
| 15:54:46 | gibi | but also kernel panic, and recenlty with POST_FAILURe | |
| 16:07:32 | stephenfin | sean-k-mooney: because I didn't write it down, can you remind me again how you were suggesting me map a project to a hypervisor in '/os-hypervisors'? Was it metadata? | |
| 16:07:59 | stephenfin | sean-k-mooney: to clarify, if I say "get all hypervisors relevant to this user", what should I be filtering on? | |
| 16:10:58 | sean-k-mooney | use the existing metadtaa keys for tenant isolation | |
| 16:11:15 | sean-k-mooney | one sec ill get it | |
| 16:12:05 | sean-k-mooney | https://docs.openstack.org/nova/latest/admin/aggregates.html#tenant-isolation-with-placement | |
| 16:12:26 | sean-k-mooney | you would be looking for filter_tenant_id* | |
| 16:12:30 | lyarwood | sigh, why is devstack writing out /etc/cinder/cinder-api-uwsgi.ini when we only deploy c-vol | |
| 16:12:58 | stephenfin | sean-k-mooney: great | |
| 16:14:30 | sean-k-mooney | basically if the host is not a member of an aggreate with filter_tenant_id* then anyone coudl view it | |
| 16:14:48 | sean-k-mooney | if it is then only does listed in that can view it | |
| 16:16:11 | sean-k-mooney | the way i was suggsing doing it was look for all host wiht filter_tenant_id=<my-project> and if that is none then allow all hosts | |
| 16:16:33 | sean-k-mooney | you coudl do it other ways fo course but i think that is what i suggested in the past | |