| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-25 | |||
| 17:20:28 | sean-k-mooney | im going to go get dinner. i might be around later but im mostly done for today | |
| 17:50:23 | dansmith | mmm, ceph job appears to be failing again in a similar way.. hope we don't have more work to do | |
| 18:31:21 | bauzas | dansmith: which patch are you checking for the job runs ? | |
| 18:31:46 | bauzas | so I can try to look over it tomorrow morning | |
| 18:32:05 | dansmith | https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315 | |
| 18:32:31 | bauzas | ack, will target it tomorrow morning | |
| 18:32:39 | dansmith | it only failed six tests this time instead of a timeout, so maybe it's better than I thought | |
| 18:32:49 | dansmith | but six is still a lot, and I haven't gone through the latest logs yet | |
| 18:33:35 | bauzas | I can try to dig into those later | |
| 18:44:26 | dansmith | the fails look all volume detach related | |
| 18:44:31 | dansmith | so perhaps it's not really a ceph problem | |
| 18:44:38 | dansmith | but it seems like a large number for a single run, so I'm not sure | |
| 20:31:17 | dansmith | melwitt: can you +W this? https://review.opendev.org/c/openstack/nova/+/881409/2 | |
| 20:38:44 | dansmith | eharney: gouthamr: Well, it installs and "works" on jammy, but something isn't happy: https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315?tab=change-view-tab-header-zuul-results-summary | |
| 20:41:21 | dansmith | eharney: gouthamr I don't really know what I'm looking at, but I don't see any errors in the cinder or ceph stuff that I recognize, and just some "busy" messages from rbd around the failed detaches in the n-cpu log | |
| 20:42:25 | gouthamr | hey dansmith - /me is late to the party | |
| 20:43:06 | dansmith | gouthamr: we need to drop focal from the jobs and our gate has been blocked for two days because some did it early..we've reverted those things for the moment, but we need to get the ceph job working on jammy | |
| 20:43:37 | dansmith | gouthamr: the above patch unpins the jobs to let them run on jammy and they get pretty far, but some volume/ceph related failures are showing that something is not happy | |
| 20:44:31 | dansmith | gouthamr: are you the right person to get that working? | |
| 20:44:49 | melwitt | dansmith: just to confirm, you still want to remove after neutron has reverted? https://review.opendev.org/c/openstack/neutron/+/881430 | |
| 20:45:27 | dansmith | melwitt: yeah the neutron failure was a couple failures ago, and not even the only problem.. but as noted in the original patch, it was intended to only live for antelope and then be reverted, so we need to do it anyway | |
| 20:45:27 | gouthamr | dansmith: probably not; i'm not an expert on rbd or cinder... | |
| 20:45:41 | dansmith | gouthamr: oh.. who is then? | |
| 20:46:10 | gouthamr | eharney is my go to guy, probably jbernard | |
| 20:46:48 | dansmith | gouthamr: okay he said earlier today that he was not likely the guy to ask (unless I misunderstood) | |
| 20:47:42 | gouthamr | ah :) let me look at the logs and see if something pops out | |
| 20:48:23 | gouthamr | we've been burnt before by using distro packages for ceph because fixes took forever to land - so we shied away from them and looked upstream.. | |
| 20:49:08 | gouthamr | but, like you've discussed, the ceph community hasn't built jammy packages for the latest release (quincy) - they meant to, they lost people/mindshare in the recent months | |
| 20:49:20 | dansmith | gouthamr: yeah, but last I checked, there were not packages from ceph themselves for jammy | |
| 20:49:34 | dansmith | gouthamr: and the cephadm job is even more broken and marked n-v so I assume it's not healthier | |
| 20:49:48 | gouthamr | on the manila jobs, we pivoted to use centos-stream-9 because ceph folks continue to publish packages there | |
| 20:49:54 | gouthamr | s/there/for it | |
| 20:50:29 | gouthamr | dansmith: yep; on that change, i see there's a problem with podman... | |
| 20:50:35 | dansmith | gouthamr: yeah, but stream breaks us constantly | |
| 20:50:40 | gouthamr | oh | |
| 20:50:54 | dansmith | there's very little chance we're going to be able to have this job run on stream :) | |
| 20:51:15 | dansmith | gouthamr: yeah I see the podman thing, but the job is also marked non-voting | |
| 20:51:26 | dansmith | so I assume it doesn't have a long track record of stability :) | |
| 20:52:03 | gouthamr | yes; i don't think we've run the job long enough to test for stability: https://zuul.opendev.org/t/openstack/builds?job_name=devstack-plugin-ceph-cephfs-nfs | |
| 20:52:15 | dansmith | presumably the cephadm approach means we could run on jammy but with upstream ceph fixes | |
| 20:52:26 | gouthamr | yep | |
| 20:52:37 | dansmith | gouthamr: that's the cephfs job, but I assume that's not what nova needs | |
| 20:52:57 | dansmith | I'm looking at devstack-plugin-ceph-tempest-cephadm | |
| 20:53:01 | gouthamr | yes; just pointing out that job because it uses centos-9-stream | |
| 20:53:07 | dansmith | ah okay | |
| 20:53:14 | gouthamr | ack; | |
| 20:53:32 | gouthamr | "devstack-plugin-ceph-tempest-cephadm" on focal-fossa used a third party repo to get podman | |
| 20:54:02 | dansmith | ah, and podman is in jammy itself I think right? | |
| 20:54:12 | gouthamr | by the looks of it, yes | |
| 20:54:30 | dansmith | although it seems broken :) | |
| 20:55:08 | dansmith | anyway, I thought there was also some concern that the cephadm job didn't expose the ceph config that nova needed or something like that, but I heard that like 20th hand | |
| 20:55:27 | gouthamr | shouldn't be the case | |
| 20:55:55 | dansmith | okay | |
| 20:56:23 | gouthamr | i think we tested some of this without tempest in the picture - but that job ("devstack-plugin-ceph-tempest-cephadm") has never passed; we assumed someone working on nova/cinder/glance would help looking at it at some point | |
| 20:57:32 | gouthamr | sorry this feels disjointed - conversations happened on irc and gerrit, ptg and the ML iirc.. but its time to reprise this because it's urgent.. | |
| 20:57:49 | gouthamr | that's a tangent though, let me see if i can spot an issue with the package based job you're looking to fix | |
| 20:58:03 | dansmith | ugh, never passed? that's no good.. I wonder if for the same reason the non-cephadm job is failing/ | |
| 20:58:21 | dansmith | I can try to get podman working on this to see if it's otherwise the same | |
| 20:58:26 | gouthamr | ++ | |
| 20:58:41 | dansmith | gouthamr: this is definitely disjointed, and I feel like I'm just flailing because nobody else is :/ | |
| 20:59:25 | gouthamr | you're doing godly work :D | |
| 20:59:53 | dansmith | $deitly work you mean :) | |
| 21:00:07 | dansmith | er, $deityly .. or soething | |
| 21:00:18 | gouthamr | :P | |
| 21:00:32 | dansmith | okay I just pushed something that might get podman working based on that error message, so we'll see | |
| 21:01:34 | dansmith | melwitt: thanks for the +W.. I would have just removed that neutron reference and changed to "effing everything" but didn't want to have to make another trip through the jobs, as you can probably imagine :) | |
| 21:06:04 | dansmith | gouthamr: the cephadm that has never passed.. is that always on quincy (i.e. newer than what we were running in focal) or what? | |
| 21:06:08 | dansmith | I | |
| 21:06:29 | melwitt | dansmith: understandable :) | |
| 21:07:27 | gouthamr | dansmith: 5/6 failures are on volume detach timeouts; and the request never got to cinder afaict .... https://zuul.opendev.org/t/openstack/build/9ebda7c1ebf843209e57ef0eac13814f/log/controller/logs/screen-n-cpu.txt#61276-61329 | |
| 21:07:53 | dansmith | gouthamr: yeah but you see the rbd busy messages right? | |
| 21:08:18 | dansmith | I commented on an earlier patch | |
| 21:08:27 | gouthamr | ah; no i missed those | |
| 21:08:48 | dansmith | I'm not sure the cinder detach would have happened by this point by the way, because we haven't gotten the guest to let go yet | |
| 21:09:19 | gouthamr | yes | |
| 21:10:24 | gouthamr | might be my browser, but i don't see "rbd.ImageBusy: [errno 16] RBD image is busy (error removing image)" in the latest n-cpu logs | |
| 21:10:30 | gouthamr | should i be looking elsewhere? | |
| 21:11:09 | dansmith | yeah actually I don't think I see them in the latest either.. but that was just a recheck | |
| 21:11:23 | dansmith | almost identical set of failures though | |
| 21:11:28 | dansmith | so yeah.. weird | |
| 21:13:43 | dansmith | gouthamr: here's an example from the previous run: https://zuul.opendev.org/t/openstack/build/f2ecbdd78616419cb5c8c2b3f4a8b71a/log/controller/logs/screen-n-cpu.txt#55439 | |
| 21:15:19 | dansmith | that's all over those logs and absent from the latest.. bizarre | |
| 21:17:14 | dansmith | gouthamr: cephadm job made it past cephadm install phase, so.. progress I think | |
| 21:17:23 | gouthamr | very nice | |
| 21:20:57 | dansmith | and finished pool setup (ignorant guess from the commands it ran) | |
| 21:22:43 | gouthamr | https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/devstack/files/debs/devstack-plugin-ceph breaks focal fossa jobs (cephfs-native) though - but, we can fix that up with some os version annotation, correct? | |
| 21:23:27 | dansmith | we can fix it by putting it in the code instead of those package lists | |
| 21:23:42 | dansmith | I just jammed it in there because it was hard to mess up and will make sure we get it installed from the distro | |
| 21:24:04 | gouthamr | ack; either that or i can just fix the native cephfs job to use jammy too | |
| 21:24:09 | dansmith | the current install_podman thing does checks for focal, so I'd just extend that | |
| 21:24:32 | dansmith | gouthamr: yeah, although the ceph jobs on stable have to use the plugin without branches don't they? | |
| 21:24:46 | gouthamr | no this repo is branched | |
| 21:24:54 | dansmith | ah okay | |
| 21:25:30 | dansmith | well, either way.. if we move this to only >=jammy for everything (along with the PTI for 2023.2) then we can just use this debs list thing and remove the focal-specific install stuff.. whatever you ant | |
| 21:25:33 | dansmith | I just want it to work :) | |
| 21:25:50 | gouthamr | agree; lets see this work | |
| 21:27:14 | gouthamr | if the "rbd remove" thing fails again on this run with the error, i would suggest reporting a bug - we could _try_ this thing on centos-9-stream and see whether there's some weirdness in the distro packages | |
| 21:27:58 | gouthamr | but, i am nervous about that job's future with ubuntu since we've learned about the ceph community's stance | |
| 21:28:03 | dansmith | ack, well, based on how this is working, I'm hoping the cephadm will either "work" or "fail the same way" and then we can discount distro packages | |
| 21:28:11 | gouthamr | ++ | |