| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-03-09 | |||
| 13:31:35 | jrosser | theres a bunch of trap doors for the unwary here :) | |
| 13:32:11 | sean-k-mooney | for what its worth i have generally not had issues with this. its typeically just worked | |
| 13:32:23 | sean-k-mooney | even before the stable rescuse work | |
| 13:35:43 | jrosser | heres what i get if i don't pass --image http://paste.openstack.org/show/803379/ | |
| 13:36:24 | lyarwood | jrosser: kk that's a bug | |
| 13:36:34 | lyarwood | jrosser: but glad the original issue is resolved at least | |
| 13:38:50 | lyarwood | https://github.com/openstack/nova/blob/31889ce296d1e1a62fe5825292479009118ddfab/nova/compute/manager.py#L4123-L4130 doesn't look right | |
| 13:38:52 | jrosser | lyarwood: would you expect hw_rescue_bus=scsi to work? | |
| 13:39:17 | lyarwood | jrosser: with a different image and an instance that already had a disk attached via SCSI yes | |
| 13:39:51 | jrosser | feels like something else there as i get "No Bootable device" in the instance console | |
| 13:44:00 | lyarwood | jrosser: would you mind raising a bug for that and the API error above when --image is missing? | |
| 13:44:17 | lyarwood | I'm not sure about the SCSI failure to find a boot device tbh | |
| 13:44:32 | lyarwood | unless again it's something weird with the image | |
| 13:44:47 | sean-k-mooney | jrosser: which image did you set the rescue bus on? | |
| 13:45:02 | jrosser | the one i'm specifying with --image | |
| 13:45:05 | sean-k-mooney | if you dont specify it it would have to be on the original image | |
| 13:45:20 | sean-k-mooney | ya ok that should be the image whos metadata we use | |
| 13:46:55 | jrosser | in this case turns out those options are set on both images i've tried as the rescue image now | |
| 13:47:12 | jrosser | one of which is the original instance image | |
| 14:13:52 | jrosser | lyarwood: bug reports done | |
| 14:14:42 | jrosser | thankyou again for your help, i have something usable now with the usb bus and understanding the need for a different rescue image | |
| 14:17:00 | lyarwood | jrosser: np and thanks for the bugs, I'll try to get them resolved after feature freeze later this week | |
| 14:36:40 | sean-k-mooney | lyarwood: by the way do you know why the volumn detach is sometime failing in the live migration job | |
| 14:40:05 | sean-k-mooney | looks like its hitting nova.exception.DeviceDetachFailed: Device detach failed for vdb: Unable to detach the device from the live config. | |
| 14:41:53 | lyarwood | sean-k-mooney: no, I've been trying to push gibi's rework along to see if that resolved it tbh | |
| 14:42:24 | sean-k-mooney | http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22Unable%20to%20detach%20the%20device%20from%20the%20live%20config%5C%22%20AND%20loglevel%3A%20ERROR | |
| 14:43:00 | lyarwood | sean-k-mooney: there's nothing obvious in the logs but I wonder if it could be related to https://bugs.launchpad.net/cinder/+bug/1917750 | |
| 14:43:01 | openstack | Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New] | |
| 14:43:07 | sean-k-mooney | ~300 hits in 30 days | |
| 14:43:52 | sean-k-mooney | im seeing some libvirt issue on the contoler too not the node with teh detach issue | |
| 14:44:47 | lyarwood | do you have an example to hand? | |
| 14:45:19 | sean-k-mooney | am its hitting my vdpa pataches but also neutron let me get one | |
| 14:47:54 | sean-k-mooney | so ya https://review.opendev.org/c/openstack/nova/+/778350/4 https://zuul.opendev.org/t/openstack/build/fb643b53835341ac8589afeadfa7044d/logs | |
| 14:49:08 | sean-k-mooney | its showing up in neutron too https://review.opendev.org/c/openstack/neutron/+/777785 | |
| 14:49:50 | sean-k-mooney | in https://zuul.opendev.org/t/openstack/build/32f3dd64008b469eb9fb8b13ed33f137 | |
| 14:50:04 | sean-k-mooney | so i think this is just an issue with master in general | |
| 14:50:27 | sean-k-mooney | it could be related to https://bugs.launchpad.net/cinder/+bug/1917750 maybe havent looked at it yet | |
| 14:50:28 | openstack | Launchpad bug 1917750 in Cinder "Running parallel iSCSI/LVM c-vol backends is causing random failures in CI" [Undecided,New] | |
| 14:57:49 | lyarwood | sean-k-mooney: yeah I think it's related | |
| 15:01:19 | lyarwood | gah | |
| 15:01:23 | lyarwood | yeah it's that | |
| 15:02:15 | lyarwood | so we end up in a situation where one test attaches a volume to the host as /dev/sda from c-vol LVM/iSCSI backend #1 | |
| 15:02:39 | lyarwood | another test then attaches another volume with the same WWN to the hsot as /dev/sdb from c-vol LVM/iSCSI backend #2 | |
| 15:03:07 | lyarwood | the first test finishes and removes what it thinks is the first volume attachment | |
| 15:03:13 | lyarwood | but it's actually the second | |
| 15:03:36 | lyarwood | leaving libvirt unable to detach the device from the instance as I assume it can't flush | |
| 15:04:02 | lyarwood | I guess our current code swallows or ignores the failure from the libvirt and just retries? | |
| 15:04:27 | sean-k-mooney | porbaly ya | |
| 15:04:34 | sean-k-mooney | we dont use the events yet | |
| 15:04:51 | sean-k-mooney | so we need to revert running those n parrallel | |
| 15:05:11 | sean-k-mooney | although is this not a os-brick bug | |
| 15:05:26 | sean-k-mooney | i mean the concurrent attach should work | |
| 15:05:29 | lyarwood | yeah it's a job topology bug | |
| 15:05:58 | lyarwood | it `works` but the first test will end up sending I/O to the volume from the second test | |
| 15:06:21 | lyarwood | something has changed recently somewhere in the stack to allow both backends to return the same WWN tbh | |
| 15:06:27 | bauzas | stephenfin: apologies for forgetting that a string is immutable | |
| 15:06:37 | lyarwood | unless people have always missed these failures | |
| 15:06:37 | sean-k-mooney | lyarwood: would not not be possible thouhg in general | |
| 15:06:52 | sean-k-mooney | for them to have the same WWN | |
| 15:07:56 | lyarwood | not with a real world backend | |
| 15:08:37 | lyarwood | https://en.wikipedia.org/wiki/World_Wide_Name | |
| 15:12:19 | sean-k-mooney | lyarwood: ya i tought if you have iscsi portals from different sotrage backend thw WWN was only uniq within any given portal | |
| 15:13:26 | sean-k-mooney | lyarwood: i dont think cinder is choosing the WWN for the volume | |
| 15:16:41 | lyarwood | right yeah that would make sense | |
| 15:22:29 | stephenfin | bauzas: nw. I've replied to the rest of your comments also | |
| 15:42:15 | bauzas | yeah I'll review again | |
| 15:42:36 | gibi | lyarwood: do I understand correclty from the scrollback that making some cinder test serial would help merging patches? | |
| 15:42:55 | lyarwood | gibi: removing c-vol from the computes would | |
| 15:43:33 | lyarwood | gibi: if we really need to run another c-vol service then it needs to be on the controller alongside the original | |
| 15:43:41 | lyarwood | gibi: just building an env now to prove things | |
| 15:43:43 | stephenfin | lyarwood, gibi, sean-k-mooney: any aversion to adding mypy to pre-commit? | |
| 15:43:59 | sean-k-mooney | not speciriclly no | |
| 15:44:00 | stephenfin | assuming I can do it while respecting mypy-files | |
| 15:44:04 | lyarwood | stephenfin: against the files that are changing? | |
| 15:44:10 | gibi | lyarwood: thanks for looking into that, sign me up for review where there is something I can push | |
| 15:44:12 | sean-k-mooney | i have been bitten by flake8 passed by pep8 didnt | |
| 15:44:14 | stephenfin | yeah | |
| 15:44:30 | gibi | lyarwood: does having c-vol along with n-cpu is an invalid config? | |
| 15:44:33 | lyarwood | stephenfin: no issues assuming it isn't adding a huge amount of delay | |
| 15:44:34 | sean-k-mooney | i have gotten used to relying on pre-commit to do that form me | |
| 15:44:35 | stephenfin | sean-k-mooney: me too :( I ran it on HEAD but it turns out I broke something then fixed it in the next change | |
| 15:44:46 | stephenfin | me too * 2 :) | |
| 15:44:53 | stephenfin | pre-commit FTW | |
| 15:44:59 | gibi | stephenfin: no problem for me I dont use pre-commit :) | |
| 15:45:18 | stephenfin | gibi: You're missing out. It's pretty great :) | |
| 15:45:45 | lyarwood | gibi: no, the issue here is with the default c-vol backend, LVM/iSCSI, that you can't have >1 running version of in an env AFAICT | |
| 15:46:06 | gibi | stephenfin: I do run fast8 py39 and functional-py39 on every commit I push (except on old stable branches where there is no py39 or py38 support) | |
| 15:46:10 | bauzas | wife is sick and I need to taxi her to the doctor | |
| 15:46:10 | lyarwood | gibi: we end up mapping different volumes from different backends to the computes with the same WWN (world wide name) that should be unqiue | |
| 15:46:37 | gibi | lyarwood: so we cannot have multiple backend per compute? | |
| 15:46:50 | stephenfin | gibi: That's better than me. I run what I think are relevant tests and then let the CI do the rest. Our tests take too long to run locally | |
| 15:46:58 | lyarwood | gibi: we can't have multiple LVM/iSCSI backed c-vol's in the same env | |
| 15:47:22 | gibi | stephenfin: I have a beefy blade in a lab to run unit and func test (and devstack on baremetal to test sriov) | |
| 15:47:48 | lyarwood | it's not great but better than nothing | |
| 15:47:51 | stephenfin | I've a four year old laptop 0:) | |
| 15:48:01 | lyarwood | refresh is 3 soooooooooooo ;) | |
| 15:48:16 | gibi | still a laptop is slow, you should ask for a lab :) | |
| 15:49:40 | gibi | lyarwood: so in a real deployment there can only one c-vol service using the LVM/iSCSI backend? | |
| 15:49:59 | gibi | lyarwood: sorry that I'm slow to understant this :) | |
| 15:50:21 | lyarwood | gibi: I don't think that has ever been enforced but the LVM/iSCSI backend is only ever used for testing | |