Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-25
20:47:42 gouthamr ah :) let me look at the logs and see if something pops out
20:48:23 gouthamr we've been burnt before by using distro packages for ceph because fixes took forever to land - so we shied away from them and looked upstream..
20:49:08 gouthamr but, like you've discussed, the ceph community hasn't built jammy packages for the latest release (quincy) - they meant to, they lost people/mindshare in the recent months
20:49:20 dansmith gouthamr: yeah, but last I checked, there were not packages from ceph themselves for jammy
20:49:34 dansmith gouthamr: and the cephadm job is even more broken and marked n-v so I assume it's not healthier
20:49:48 gouthamr on the manila jobs, we pivoted to use centos-stream-9 because ceph folks continue to publish packages there
20:49:54 gouthamr s/there/for it
20:50:29 gouthamr dansmith: yep; on that change, i see there's a problem with podman...
20:50:35 dansmith gouthamr: yeah, but stream breaks us constantly
20:50:40 gouthamr oh
20:50:54 dansmith there's very little chance we're going to be able to have this job run on stream :)
20:51:15 dansmith gouthamr: yeah I see the podman thing, but the job is also marked non-voting
20:51:26 dansmith so I assume it doesn't have a long track record of stability :)
20:52:03 gouthamr yes; i don't think we've run the job long enough to test for stability: https://zuul.opendev.org/t/openstack/builds?job_name=devstack-plugin-ceph-cephfs-nfs
20:52:15 dansmith presumably the cephadm approach means we could run on jammy but with upstream ceph fixes
20:52:26 gouthamr yep
20:52:37 dansmith gouthamr: that's the cephfs job, but I assume that's not what nova needs
20:52:57 dansmith I'm looking at devstack-plugin-ceph-tempest-cephadm
20:53:01 gouthamr yes; just pointing out that job because it uses centos-9-stream
20:53:07 dansmith ah okay
20:53:14 gouthamr ack;
20:53:32 gouthamr "devstack-plugin-ceph-tempest-cephadm" on focal-fossa used a third party repo to get podman
20:54:02 dansmith ah, and podman is in jammy itself I think right?
20:54:12 gouthamr by the looks of it, yes
20:54:30 dansmith although it seems broken :)
20:55:08 dansmith anyway, I thought there was also some concern that the cephadm job didn't expose the ceph config that nova needed or something like that, but I heard that like 20th hand
20:55:27 gouthamr shouldn't be the case
20:55:55 dansmith okay
20:56:23 gouthamr i think we tested some of this without tempest in the picture - but that job ("devstack-plugin-ceph-tempest-cephadm") has never passed; we assumed someone working on nova/cinder/glance would help looking at it at some point
20:57:32 gouthamr sorry this feels disjointed - conversations happened on irc and gerrit, ptg and the ML iirc.. but its time to reprise this because it's urgent..
20:57:49 gouthamr that's a tangent though, let me see if i can spot an issue with the package based job you're looking to fix
20:58:03 dansmith ugh, never passed? that's no good.. I wonder if for the same reason the non-cephadm job is failing/
20:58:21 dansmith I can try to get podman working on this to see if it's otherwise the same
20:58:26 gouthamr ++
20:58:41 dansmith gouthamr: this is definitely disjointed, and I feel like I'm just flailing because nobody else is :/
20:59:25 gouthamr you're doing godly work :D
20:59:53 dansmith $deitly work you mean :)
21:00:07 dansmith er, $deityly .. or soething
21:00:18 gouthamr :P
21:00:32 dansmith okay I just pushed something that might get podman working based on that error message, so we'll see
21:01:34 dansmith melwitt: thanks for the +W.. I would have just removed that neutron reference and changed to "effing everything" but didn't want to have to make another trip through the jobs, as you can probably imagine :)
21:06:04 dansmith gouthamr: the cephadm that has never passed.. is that always on quincy (i.e. newer than what we were running in focal) or what?
21:06:08 dansmith I
21:06:29 melwitt dansmith: understandable :)
21:07:27 gouthamr dansmith: 5/6 failures are on volume detach timeouts; and the request never got to cinder afaict .... https://zuul.opendev.org/t/openstack/build/9ebda7c1ebf843209e57ef0eac13814f/log/controller/logs/screen-n-cpu.txt#61276-61329
21:07:53 dansmith gouthamr: yeah but you see the rbd busy messages right?
21:08:18 dansmith I commented on an earlier patch
21:08:27 gouthamr ah; no i missed those
21:08:48 dansmith I'm not sure the cinder detach would have happened by this point by the way, because we haven't gotten the guest to let go yet
21:09:19 gouthamr yes
21:10:24 gouthamr might be my browser, but i don't see "rbd.ImageBusy: [errno 16] RBD image is busy (error removing image)" in the latest n-cpu logs
21:10:30 gouthamr should i be looking elsewhere?
21:11:09 dansmith yeah actually I don't think I see them in the latest either.. but that was just a recheck
21:11:23 dansmith almost identical set of failures though
21:11:28 dansmith so yeah.. weird
21:13:43 dansmith gouthamr: here's an example from the previous run: https://zuul.opendev.org/t/openstack/build/f2ecbdd78616419cb5c8c2b3f4a8b71a/log/controller/logs/screen-n-cpu.txt#55439
21:15:19 dansmith that's all over those logs and absent from the latest.. bizarre
21:17:14 dansmith gouthamr: cephadm job made it past cephadm install phase, so.. progress I think
21:17:23 gouthamr very nice
21:20:57 dansmith and finished pool setup (ignorant guess from the commands it ran)
21:22:43 gouthamr https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/devstack/files/debs/devstack-plugin-ceph breaks focal fossa jobs (cephfs-native) though - but, we can fix that up with some os version annotation, correct?
21:23:27 dansmith we can fix it by putting it in the code instead of those package lists
21:23:42 dansmith I just jammed it in there because it was hard to mess up and will make sure we get it installed from the distro
21:24:04 gouthamr ack; either that or i can just fix the native cephfs job to use jammy too
21:24:09 dansmith the current install_podman thing does checks for focal, so I'd just extend that
21:24:32 dansmith gouthamr: yeah, although the ceph jobs on stable have to use the plugin without branches don't they?
21:24:46 gouthamr no this repo is branched
21:24:54 dansmith ah okay
21:25:30 dansmith well, either way.. if we move this to only >=jammy for everything (along with the PTI for 2023.2) then we can just use this debs list thing and remove the focal-specific install stuff.. whatever you ant
21:25:33 dansmith I just want it to work :)
21:25:50 gouthamr agree; lets see this work
21:27:14 gouthamr if the "rbd remove" thing fails again on this run with the error, i would suggest reporting a bug - we could _try_ this thing on centos-9-stream and see whether there's some weirdness in the distro packages
21:27:58 gouthamr but, i am nervous about that job's future with ubuntu since we've learned about the ceph community's stance
21:28:03 dansmith ack, well, based on how this is working, I'm hoping the cephadm will either "work" or "fail the same way" and then we can discount distro packages
21:28:11 gouthamr ++
21:28:27 dansmith it's running tempest now and looking identical to the distro version so far (i.e. hasn't failed but hasn't run volume tests yet)
21:28:46 dansmith so that's majorly better. failing with upstream bits is a minor win over failing with distro bits :P
21:29:03 eharney i have some work in flight currently around rbd ImageBusy errors... i wonder if this job is running some tests that were previously disabled?
21:29:31 dansmith eharney: shouldn't be, and one run failed with a bunch of them and then a single recheck failed similarly, but with no rbdbusy errors
21:31:40 gouthamr eharney: interesting, are there librbd changes that you're having to work around?
21:32:15 eharney gouthamr: not new ones, just working on fixing the class of errors around rbd images that can't be deleted that we've always had
21:34:57 gouthamr eharney: oh.. this error seems to have occurred multiple times in the libvirt rbd "remove image" call..
21:36:04 gouthamr i'm hoping no openstack code needs to change; we're hoping to support quincy with stable/wallaby downstream :D based on some testing of this stuff elsewhere
21:36:36 dansmith gouthamr: yeah I was going to ask earlier.. are we using quincy downstream such that we know it works? this set of failures has me concerned that it's something fundamental of course
21:37:14 gouthamr (same tests, different OS, full blown ceph cluster as opposed to our aio .. etc etc -- so there could be a number of things being issues)
21:38:41 gouthamr dansmith: yes, we're not testing quincy with openstack's trunk though... we trail downstream; but this stuff is working with zed last i checked, with the same tempest tests passing there
21:38:53 dansmith okay that's good
21:56:15 dansmith well, it's passing some volume tests at least
21:56:29 gouthamr ++
21:57:22 gouthamr a couple of things we wanted to try in this cephadm job in the past:
21:57:44 dansmith gouthamr: so if this works magically, you're okay just making this drop support for focal as long as the other jobs here are set to run on jammy?
21:57:46 gouthamr (1) revert to default test concurrency --- the concurrency was set to 1 because we saw resource contention
21:58:00 gouthamr dansmith: yes
21:58:02 dansmith resource contention like memory/
21:58:10 gouthamr yes, and disk
21:58:26 dansmith so there's a tweak in devstack for memory that has been helping a lot of jobs, and we run with it enabled in the nova ceph jobs
21:58:39 gouthamr oh? i'd love to know!
21:58:43 dansmith drops mysql usage by about half, which is ~400MiB on these jobs
21:58:47 gouthamr nice
21:58:55 dansmith gouthamr: problem is you have to pay me a royalty per job execution to use it

Earlier   Later