Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-25
21:25:30 dansmith well, either way.. if we move this to only >=jammy for everything (along with the PTI for 2023.2) then we can just use this debs list thing and remove the focal-specific install stuff.. whatever you ant
21:25:33 dansmith I just want it to work :)
21:25:50 gouthamr agree; lets see this work
21:27:14 gouthamr if the "rbd remove" thing fails again on this run with the error, i would suggest reporting a bug - we could _try_ this thing on centos-9-stream and see whether there's some weirdness in the distro packages
21:27:58 gouthamr but, i am nervous about that job's future with ubuntu since we've learned about the ceph community's stance
21:28:03 dansmith ack, well, based on how this is working, I'm hoping the cephadm will either "work" or "fail the same way" and then we can discount distro packages
21:28:11 gouthamr ++
21:28:27 dansmith it's running tempest now and looking identical to the distro version so far (i.e. hasn't failed but hasn't run volume tests yet)
21:28:46 dansmith so that's majorly better. failing with upstream bits is a minor win over failing with distro bits :P
21:29:03 eharney i have some work in flight currently around rbd ImageBusy errors... i wonder if this job is running some tests that were previously disabled?
21:29:31 dansmith eharney: shouldn't be, and one run failed with a bunch of them and then a single recheck failed similarly, but with no rbdbusy errors
21:31:40 gouthamr eharney: interesting, are there librbd changes that you're having to work around?
21:32:15 eharney gouthamr: not new ones, just working on fixing the class of errors around rbd images that can't be deleted that we've always had
21:34:57 gouthamr eharney: oh.. this error seems to have occurred multiple times in the libvirt rbd "remove image" call..
21:36:04 gouthamr i'm hoping no openstack code needs to change; we're hoping to support quincy with stable/wallaby downstream :D based on some testing of this stuff elsewhere
21:36:36 dansmith gouthamr: yeah I was going to ask earlier.. are we using quincy downstream such that we know it works? this set of failures has me concerned that it's something fundamental of course
21:37:14 gouthamr (same tests, different OS, full blown ceph cluster as opposed to our aio .. etc etc -- so there could be a number of things being issues)
21:38:41 gouthamr dansmith: yes, we're not testing quincy with openstack's trunk though... we trail downstream; but this stuff is working with zed last i checked, with the same tempest tests passing there
21:38:53 dansmith okay that's good
21:56:15 dansmith well, it's passing some volume tests at least
21:56:29 gouthamr ++
21:57:22 gouthamr a couple of things we wanted to try in this cephadm job in the past:
21:57:44 dansmith gouthamr: so if this works magically, you're okay just making this drop support for focal as long as the other jobs here are set to run on jammy?
21:57:46 gouthamr (1) revert to default test concurrency --- the concurrency was set to 1 because we saw resource contention
21:58:00 gouthamr dansmith: yes
21:58:02 dansmith resource contention like memory/
21:58:10 gouthamr yes, and disk
21:58:26 dansmith so there's a tweak in devstack for memory that has been helping a lot of jobs, and we run with it enabled in the nova ceph jobs
21:58:39 gouthamr oh? i'd love to know!
21:58:43 dansmith drops mysql usage by about half, which is ~400MiB on these jobs
21:58:47 gouthamr nice
21:58:55 dansmith gouthamr: problem is you have to pay me a royalty per job execution to use it
21:59:17 dansmith gouthamr: https://github.com/openstack/nova/blob/master/.zuul.yaml#L422
21:59:33 gouthamr :D i'd be poor if i was betting on ceph using less memory ever
21:59:38 dansmith haha
22:00:03 gouthamr (2) there's also a way to turn off "cephadm" after the install -- we should set that option on the CI - there's no use to keep it running
22:00:06 dansmith so based on our success with that, if you're cool with it, I'd say we turn that on for these jobs anyway
22:00:24 dansmith gouthamr: meaning before we start running tempest?
22:01:30 gouthamr yes, the plugin will do that for you after the ceph cluster deployment is done: https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/cephadm#L28
22:02:12 dansmith oh so why is that not being set?
22:02:33 dansmith I don't even know what that means.. cephadm is a tool I thought, but you're saying it stays running even though the services are started or something?
22:02:36 gouthamr (here's an example from the only good cephadm job at the moment: https://github.com/openstack/manila-tempest-plugin/blob/ad0db6359c8c51c4521ac6660c8014981b2f1dea/zuul.d/manila-tempest-jobs.yaml#L413)
22:02:49 dansmith ack
22:02:56 gouthamr yes; we need it for day2 operations on the cluster
22:03:07 gouthamr so if you're running this on your local devstack, it's useful
22:03:13 dansmith ah, sure, okay
22:03:24 dansmith we should set that in these jobs that everyone inherits from
22:03:36 gouthamr ++
22:03:42 dansmith so maybe a follow-on to this to set that, the mysql thing, and make this job voting
22:04:00 dansmith you know, if and when it starts working :)
22:04:40 gouthamr ^^ +1
22:06:02 dansmith hmm, gouthamr I just noticed that we're still set to release=pacific for the distro-package-based job
22:06:11 dansmith is that maybe setting stuff that is non-ideal for quincy that could be related?
22:06:38 gouthamr i think the plugin ignores it, let me check
22:06:43 dansmith okay
22:06:59 dansmith https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/.zuul.yaml#46
22:08:43 dansmith okay just used for package repos anyway it looks like
22:08:51 gouthamr ack; we should get that opt out of the job because its confusing
22:08:53 gouthamr https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/ceph#L981-L989
22:09:43 dansmith ack, so I assume we're configuring a repo but just not installing anything from it on jammy, right?
22:10:02 dansmith I'll add it to my follow-on patch to add the optimization devstack vars
22:10:08 dansmith add .. the removal of it, I mean :)
22:10:13 gouthamr ack ty
22:10:20 gouthamr i may have messed this up in my last patch
22:10:21 gouthamr https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/ceph#L1067-L1094
22:10:47 gouthamr we're not invoking that method at all for ubuntu anymore..
22:11:27 dansmith ack
22:11:45 dansmith well, there's a bunch of focal-specific stuff to clean up in there regardless
22:14:58 dansmith hmm, I think it's about to fail a test
22:15:11 dansmith been stopped for going on five minutes, assuming retrying a detach
22:15:38 gouthamr yes.. it was a head-scratcher; think vkmc and i noticed that our override to use download.ceph.com stopped working with focal at some point since ubuntu default-enabled the ubuntu ceph repos.. so dropping it made no difference, we ended up using/testing with the distro provided packages
22:15:43 gouthamr oh
22:16:19 gouthamr test_rebuild_server_with_volume_attached?
22:16:25 dansmith dunno yet
22:16:40 gouthamr ah
22:16:43 dansmith but it's in the rebuld group
22:16:46 dansmith yep
22:17:14 dansmith test_rebuild_server_with_volume_attached [430.305393s] ... FAILED
22:17:18 dansmith ugh
22:18:19 dansmith so the other variable here is the version of qemu and qemu's block-rbd driver are different in jammy of course, compared to what we've been testing
22:18:42 dansmith so could be a bug in one of those, especially since it's related to the detach in the guest
22:19:35 gouthamr ack; another thing to try would be to bump the ceph image to the latest quincy: https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/cephadm#L32
22:20:02 dansmith okay
22:20:17 gouthamr they've published v17.2.6 today, and v17.2.5 a month ago
22:21:01 gouthamr https://quay.io/repository/ceph/ceph?tab=tags
22:21:45 dansmith ack, I hate this sort of "version minesweeper" game.. if we're that sensitive to version, it feels like we're doing something wrong
22:22:39 gouthamr agreed; but since this stuff hasn't worked before on our ci, its worth a try
22:22:58 dansmith yeah for sure
22:23:18 dansmith volume resize passed
22:23:32 dansmith maybe we'll find that this is fewer fails or something
22:25:18 dansmith actually, that one didn't fail before
22:28:07 gouthamr okay, this may be good news? devstack-plugin-ceph-tempest-py3 is going to pass
22:28:34 dansmith no, really?
22:28:52 dansmith third recheck's a charm?
22:28:55 gouthamr :D
22:42:28 dansmith more fails on this cephadm job
22:42:48 dansmith so I guess it's not something fundamental, but maybe just massively less stable or we're hitting some race easier?
22:43:26 dansmith maybe it is memory-related and we're stressed more here
22:43:41 dansmith maybe I should try turning on the two optimizations here to see if that makes things more stable
22:43:54 gouthamr ++

Earlier   Later