Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-25
20:42:25 gouthamr hey dansmith - /me is late to the party
20:43:06 dansmith gouthamr: we need to drop focal from the jobs and our gate has been blocked for two days because some did it early..we've reverted those things for the moment, but we need to get the ceph job working on jammy
20:43:37 dansmith gouthamr: the above patch unpins the jobs to let them run on jammy and they get pretty far, but some volume/ceph related failures are showing that something is not happy
20:44:31 dansmith gouthamr: are you the right person to get that working?
20:44:49 melwitt dansmith: just to confirm, you still want to remove after neutron has reverted? https://review.opendev.org/c/openstack/neutron/+/881430
20:45:27 dansmith melwitt: yeah the neutron failure was a couple failures ago, and not even the only problem.. but as noted in the original patch, it was intended to only live for antelope and then be reverted, so we need to do it anyway
20:45:27 gouthamr dansmith: probably not; i'm not an expert on rbd or cinder...
20:45:41 dansmith gouthamr: oh.. who is then?
20:46:10 gouthamr eharney is my go to guy, probably jbernard
20:46:48 dansmith gouthamr: okay he said earlier today that he was not likely the guy to ask (unless I misunderstood)
20:47:42 gouthamr ah :) let me look at the logs and see if something pops out
20:48:23 gouthamr we've been burnt before by using distro packages for ceph because fixes took forever to land - so we shied away from them and looked upstream..
20:49:08 gouthamr but, like you've discussed, the ceph community hasn't built jammy packages for the latest release (quincy) - they meant to, they lost people/mindshare in the recent months
20:49:20 dansmith gouthamr: yeah, but last I checked, there were not packages from ceph themselves for jammy
20:49:34 dansmith gouthamr: and the cephadm job is even more broken and marked n-v so I assume it's not healthier
20:49:48 gouthamr on the manila jobs, we pivoted to use centos-stream-9 because ceph folks continue to publish packages there
20:49:54 gouthamr s/there/for it
20:50:29 gouthamr dansmith: yep; on that change, i see there's a problem with podman...
20:50:35 dansmith gouthamr: yeah, but stream breaks us constantly
20:50:40 gouthamr oh
20:50:54 dansmith there's very little chance we're going to be able to have this job run on stream :)
20:51:15 dansmith gouthamr: yeah I see the podman thing, but the job is also marked non-voting
20:51:26 dansmith so I assume it doesn't have a long track record of stability :)
20:52:03 gouthamr yes; i don't think we've run the job long enough to test for stability: https://zuul.opendev.org/t/openstack/builds?job_name=devstack-plugin-ceph-cephfs-nfs
20:52:15 dansmith presumably the cephadm approach means we could run on jammy but with upstream ceph fixes
20:52:26 gouthamr yep
20:52:37 dansmith gouthamr: that's the cephfs job, but I assume that's not what nova needs
20:52:57 dansmith I'm looking at devstack-plugin-ceph-tempest-cephadm
20:53:01 gouthamr yes; just pointing out that job because it uses centos-9-stream
20:53:07 dansmith ah okay
20:53:14 gouthamr ack;
20:53:32 gouthamr "devstack-plugin-ceph-tempest-cephadm" on focal-fossa used a third party repo to get podman
20:54:02 dansmith ah, and podman is in jammy itself I think right?
20:54:12 gouthamr by the looks of it, yes
20:54:30 dansmith although it seems broken :)
20:55:08 dansmith anyway, I thought there was also some concern that the cephadm job didn't expose the ceph config that nova needed or something like that, but I heard that like 20th hand
20:55:27 gouthamr shouldn't be the case
20:55:55 dansmith okay
20:56:23 gouthamr i think we tested some of this without tempest in the picture - but that job ("devstack-plugin-ceph-tempest-cephadm") has never passed; we assumed someone working on nova/cinder/glance would help looking at it at some point
20:57:32 gouthamr sorry this feels disjointed - conversations happened on irc and gerrit, ptg and the ML iirc.. but its time to reprise this because it's urgent..
20:57:49 gouthamr that's a tangent though, let me see if i can spot an issue with the package based job you're looking to fix
20:58:03 dansmith ugh, never passed? that's no good.. I wonder if for the same reason the non-cephadm job is failing/
20:58:21 dansmith I can try to get podman working on this to see if it's otherwise the same
20:58:26 gouthamr ++
20:58:41 dansmith gouthamr: this is definitely disjointed, and I feel like I'm just flailing because nobody else is :/
20:59:25 gouthamr you're doing godly work :D
20:59:53 dansmith $deitly work you mean :)
21:00:07 dansmith er, $deityly .. or soething
21:00:18 gouthamr :P
21:00:32 dansmith okay I just pushed something that might get podman working based on that error message, so we'll see
21:01:34 dansmith melwitt: thanks for the +W.. I would have just removed that neutron reference and changed to "effing everything" but didn't want to have to make another trip through the jobs, as you can probably imagine :)
21:06:04 dansmith gouthamr: the cephadm that has never passed.. is that always on quincy (i.e. newer than what we were running in focal) or what?
21:06:08 dansmith I
21:06:29 melwitt dansmith: understandable :)
21:07:27 gouthamr dansmith: 5/6 failures are on volume detach timeouts; and the request never got to cinder afaict .... https://zuul.opendev.org/t/openstack/build/9ebda7c1ebf843209e57ef0eac13814f/log/controller/logs/screen-n-cpu.txt#61276-61329
21:07:53 dansmith gouthamr: yeah but you see the rbd busy messages right?
21:08:18 dansmith I commented on an earlier patch
21:08:27 gouthamr ah; no i missed those
21:08:48 dansmith I'm not sure the cinder detach would have happened by this point by the way, because we haven't gotten the guest to let go yet
21:09:19 gouthamr yes
21:10:24 gouthamr might be my browser, but i don't see "rbd.ImageBusy: [errno 16] RBD image is busy (error removing image)" in the latest n-cpu logs
21:10:30 gouthamr should i be looking elsewhere?
21:11:09 dansmith yeah actually I don't think I see them in the latest either.. but that was just a recheck
21:11:23 dansmith almost identical set of failures though
21:11:28 dansmith so yeah.. weird
21:13:43 dansmith gouthamr: here's an example from the previous run: https://zuul.opendev.org/t/openstack/build/f2ecbdd78616419cb5c8c2b3f4a8b71a/log/controller/logs/screen-n-cpu.txt#55439
21:15:19 dansmith that's all over those logs and absent from the latest.. bizarre
21:17:14 dansmith gouthamr: cephadm job made it past cephadm install phase, so.. progress I think
21:17:23 gouthamr very nice
21:20:57 dansmith and finished pool setup (ignorant guess from the commands it ran)
21:22:43 gouthamr https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/devstack/files/debs/devstack-plugin-ceph breaks focal fossa jobs (cephfs-native) though - but, we can fix that up with some os version annotation, correct?
21:23:27 dansmith we can fix it by putting it in the code instead of those package lists
21:23:42 dansmith I just jammed it in there because it was hard to mess up and will make sure we get it installed from the distro
21:24:04 gouthamr ack; either that or i can just fix the native cephfs job to use jammy too
21:24:09 dansmith the current install_podman thing does checks for focal, so I'd just extend that
21:24:32 dansmith gouthamr: yeah, although the ceph jobs on stable have to use the plugin without branches don't they?
21:24:46 gouthamr no this repo is branched
21:24:54 dansmith ah okay
21:25:30 dansmith well, either way.. if we move this to only >=jammy for everything (along with the PTI for 2023.2) then we can just use this debs list thing and remove the focal-specific install stuff.. whatever you ant
21:25:33 dansmith I just want it to work :)
21:25:50 gouthamr agree; lets see this work
21:27:14 gouthamr if the "rbd remove" thing fails again on this run with the error, i would suggest reporting a bug - we could _try_ this thing on centos-9-stream and see whether there's some weirdness in the distro packages
21:27:58 gouthamr but, i am nervous about that job's future with ubuntu since we've learned about the ceph community's stance
21:28:03 dansmith ack, well, based on how this is working, I'm hoping the cephadm will either "work" or "fail the same way" and then we can discount distro packages
21:28:11 gouthamr ++
21:28:27 dansmith it's running tempest now and looking identical to the distro version so far (i.e. hasn't failed but hasn't run volume tests yet)
21:28:46 dansmith so that's majorly better. failing with upstream bits is a minor win over failing with distro bits :P
21:29:03 eharney i have some work in flight currently around rbd ImageBusy errors... i wonder if this job is running some tests that were previously disabled?
21:29:31 dansmith eharney: shouldn't be, and one run failed with a bunch of them and then a single recheck failed similarly, but with no rbdbusy errors
21:31:40 gouthamr eharney: interesting, are there librbd changes that you're having to work around?
21:32:15 eharney gouthamr: not new ones, just working on fixing the class of errors around rbd images that can't be deleted that we've always had
21:34:57 gouthamr eharney: oh.. this error seems to have occurred multiple times in the libvirt rbd "remove image" call..
21:36:04 gouthamr i'm hoping no openstack code needs to change; we're hoping to support quincy with stable/wallaby downstream :D based on some testing of this stuff elsewhere
21:36:36 dansmith gouthamr: yeah I was going to ask earlier.. are we using quincy downstream such that we know it works? this set of failures has me concerned that it's something fundamental of course
21:37:14 gouthamr (same tests, different OS, full blown ceph cluster as opposed to our aio .. etc etc -- so there could be a number of things being issues)
21:38:41 gouthamr dansmith: yes, we're not testing quincy with openstack's trunk though... we trail downstream; but this stuff is working with zed last i checked, with the same tempest tests passing there
21:38:53 dansmith okay that's good
21:56:15 dansmith well, it's passing some volume tests at least
21:56:29 gouthamr ++
21:57:22 gouthamr a couple of things we wanted to try in this cephadm job in the past:

Earlier   Later