Earlier  
Posted Nick Remark
#openstack-nova - 2021-03-09
13:44:32 lyarwood unless again it's something weird with the image
13:44:47 sean-k-mooney jrosser: which image did you set the rescue bus on?
13:45:02 jrosser the one i'm specifying with --image
13:45:05 sean-k-mooney if you dont specify it it would have to be on the original image
13:45:20 sean-k-mooney ya ok that should be the image whos metadata we use
13:46:55 jrosser in this case turns out those options are set on both images i've tried as the rescue image now
13:47:12 jrosser one of which is the original instance image
14:13:52 jrosser lyarwood: bug reports done
14:14:42 jrosser thankyou again for your help, i have something usable now with the usb bus and understanding the need for a different rescue image
14:17:00 lyarwood jrosser: np and thanks for the bugs, I'll try to get them resolved after feature freeze later this week
14:36:40 sean-k-mooney lyarwood: by the way do you know why the volumn detach is sometime failing in the live migration job
14:40:05 sean-k-mooney looks like its hitting nova.exception.DeviceDetachFailed: Device detach failed for vdb: Unable to detach the device from the live config.
14:41:53 lyarwood sean-k-mooney: no, I've been trying to push gibi's rework along to see if that resolved it tbh
14:42:24 sean-k-mooney http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22Unable%20to%20detach%20the%20device%20from%20the%20live%20config%5C%22%20AND%20loglevel%3A%20ERROR
14:43:00 lyarwood sean-k-mooney: there's nothing obvious in the logs but I wonder if it could be related to https://bugs.launchpad.net/cinder/+bug/1917750
14:43:01 openstack Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New]
14:43:07 sean-k-mooney ~300 hits in 30 days
14:43:52 sean-k-mooney im seeing some libvirt issue on the contoler too not the node with teh detach issue
14:44:47 lyarwood do you have an example to hand?
14:45:19 sean-k-mooney am its hitting my vdpa pataches but also neutron let me get one
14:47:54 sean-k-mooney so ya https://review.opendev.org/c/openstack/nova/+/778350/4 https://zuul.opendev.org/t/openstack/build/fb643b53835341ac8589afeadfa7044d/logs
14:49:08 sean-k-mooney its showing up in neutron too https://review.opendev.org/c/openstack/neutron/+/777785
14:49:50 sean-k-mooney in https://zuul.opendev.org/t/openstack/build/32f3dd64008b469eb9fb8b13ed33f137
14:50:04 sean-k-mooney so i think this is just an issue with master in general
14:50:27 sean-k-mooney it could be related to https://bugs.launchpad.net/cinder/+bug/1917750 maybe havent looked at it yet
14:50:28 openstack Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New]
14:57:49 lyarwood sean-k-mooney: yeah I think it's related
15:01:19 lyarwood gah
15:01:23 lyarwood yeah it's that
15:02:15 lyarwood so we end up in a situation where one test attaches a volume to the host as /dev/sda from c-vol LVM/iSCSI backend #1
15:02:39 lyarwood another test then attaches another volume with the same WWN to the hsot as /dev/sdb from c-vol LVM/iSCSI backend #2
15:03:07 lyarwood the first test finishes and removes what it thinks is the first volume attachment
15:03:13 lyarwood but it's actually the second
15:03:36 lyarwood leaving libvirt unable to detach the device from the instance as I assume it can't flush
15:04:02 lyarwood I guess our current code swallows or ignores the failure from the libvirt and just retries?
15:04:27 sean-k-mooney porbaly ya
15:04:34 sean-k-mooney we dont use the events yet
15:04:51 sean-k-mooney so we need to revert running those n parrallel
15:05:11 sean-k-mooney although is this not a os-brick bug
15:05:26 sean-k-mooney i mean the concurrent attach should work
15:05:29 lyarwood yeah it's a job topology bug
15:05:58 lyarwood it `works` but the first test will end up sending I/O to the volume from the second test
15:06:21 lyarwood something has changed recently somewhere in the stack to allow both backends to return the same WWN tbh
15:06:27 bauzas stephenfin: apologies for forgetting that a string is immutable
15:06:37 lyarwood unless people have always missed these failures
15:06:37 sean-k-mooney lyarwood: would not not be possible thouhg in general
15:06:52 sean-k-mooney for them to have the same WWN
15:07:56 lyarwood not with a real world backend
15:08:37 lyarwood https://en.wikipedia.org/wiki/World_Wide_Name
15:12:19 sean-k-mooney lyarwood: ya i tought if you have iscsi portals from different sotrage backend thw WWN was only uniq within any given portal
15:13:26 sean-k-mooney lyarwood: i dont think cinder is choosing the WWN for the volume
15:16:41 lyarwood right yeah that would make sense
15:22:29 stephenfin bauzas: nw. I've replied to the rest of your comments also
15:42:15 bauzas yeah I'll review again
15:42:36 gibi lyarwood: do I understand correclty from the scrollback that making some cinder test serial would help merging patches?
15:42:55 lyarwood gibi: removing c-vol from the computes would
15:43:33 lyarwood gibi: if we really need to run another c-vol service then it needs to be on the controller alongside the original
15:43:41 lyarwood gibi: just building an env now to prove things
15:43:43 stephenfin lyarwood, gibi, sean-k-mooney: any aversion to adding mypy to pre-commit?
15:43:59 sean-k-mooney not speciriclly no
15:44:00 stephenfin assuming I can do it while respecting mypy-files
15:44:04 lyarwood stephenfin: against the files that are changing?
15:44:10 gibi lyarwood: thanks for looking into that, sign me up for review where there is something I can push
15:44:12 sean-k-mooney i have been bitten by flake8 passed by pep8 didnt
15:44:14 stephenfin yeah
15:44:30 gibi lyarwood: does having c-vol along with n-cpu is an invalid config?
15:44:33 lyarwood stephenfin: no issues assuming it isn't adding a huge amount of delay
15:44:34 sean-k-mooney i have gotten used to relying on pre-commit to do that form me
15:44:35 stephenfin sean-k-mooney: me too :( I ran it on HEAD but it turns out I broke something then fixed it in the next change
15:44:46 stephenfin me too * 2 :)
15:44:53 stephenfin pre-commit FTW
15:44:59 gibi stephenfin: no problem for me I dont use pre-commit :)
15:45:18 stephenfin gibi: You're missing out. It's pretty great :)
15:45:45 lyarwood gibi: no, the issue here is with the default c-vol backend, LVM/iSCSI, that you can't have >1 running version of in an env AFAICT
15:46:06 gibi stephenfin: I do run fast8 py39 and functional-py39 on every commit I push (except on old stable branches where there is no py39 or py38 support)
15:46:10 bauzas wife is sick and I need to taxi her to the doctor
15:46:10 lyarwood gibi: we end up mapping different volumes from different backends to the computes with the same WWN (world wide name) that should be unqiue
15:46:37 gibi lyarwood: so we cannot have multiple backend per compute?
15:46:50 stephenfin gibi: That's better than me. I run what I think are relevant tests and then let the CI do the rest. Our tests take too long to run locally
15:46:58 lyarwood gibi: we can't have multiple LVM/iSCSI backed c-vol's in the same env
15:47:22 gibi stephenfin: I have a beefy blade in a lab to run unit and func test (and devstack on baremetal to test sriov)
15:47:48 lyarwood it's not great but better than nothing
15:47:51 stephenfin I've a four year old laptop 0:)
15:48:01 lyarwood refresh is 3 soooooooooooo ;)
15:48:16 gibi still a laptop is slow, you should ask for a lab :)
15:49:40 gibi lyarwood: so in a real deployment there can only one c-vol service using the LVM/iSCSI backend?
15:49:59 gibi lyarwood: sorry that I'm slow to understant this :)
15:50:21 lyarwood gibi: I don't think that has ever been enforced but the LVM/iSCSI backend is only ever used for testing
15:50:48 gibi OK, then I stop worrying :)
15:50:53 gibi it is just test :)
15:52:19 lyarwood I'll just caveat all of the above with the fact that no one from Cinder has agreed with any of that in the bug as yet
15:52:26 lyarwood so I might be missing the point entirely
15:52:49 lyarwood but two volumes with the same WWN connected to the same host smells like something that will bork devicemapper
15:53:21 gibi if it fixes the CI and let us merge the api db compaction before FF then I'm happy to take the hit if we need to revert the change later
15:53:24 gibi :)
15:53:54 lyarwood gibi: is that stuck in a recheck loop?
15:54:01 gibi pretty much yes
15:54:04 lyarwood gah okay
15:54:11 lyarwood let me confirm and then I can push some changes
15:54:22 gibi the Queens one is recheckd through the whole weekend and still bouncing back

Earlier   Later