Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-25
20:58:03 dansmith ugh, never passed? that's no good.. I wonder if for the same reason the non-cephadm job is failing/
20:58:21 dansmith I can try to get podman working on this to see if it's otherwise the same
20:58:26 gouthamr ++
20:58:41 dansmith gouthamr: this is definitely disjointed, and I feel like I'm just flailing because nobody else is :/
20:59:25 gouthamr you're doing godly work :D
20:59:53 dansmith $deitly work you mean :)
21:00:07 dansmith er, $deityly .. or soething
21:00:18 gouthamr :P
21:00:32 dansmith okay I just pushed something that might get podman working based on that error message, so we'll see
21:01:34 dansmith melwitt: thanks for the +W.. I would have just removed that neutron reference and changed to "effing everything" but didn't want to have to make another trip through the jobs, as you can probably imagine :)
21:06:04 dansmith gouthamr: the cephadm that has never passed.. is that always on quincy (i.e. newer than what we were running in focal) or what?
21:06:08 dansmith I
21:06:29 melwitt dansmith: understandable :)
21:07:27 gouthamr dansmith: 5/6 failures are on volume detach timeouts; and the request never got to cinder afaict .... https://zuul.opendev.org/t/openstack/build/9ebda7c1ebf843209e57ef0eac13814f/log/controller/logs/screen-n-cpu.txt#61276-61329
21:07:53 dansmith gouthamr: yeah but you see the rbd busy messages right?
21:08:18 dansmith I commented on an earlier patch
21:08:27 gouthamr ah; no i missed those
21:08:48 dansmith I'm not sure the cinder detach would have happened by this point by the way, because we haven't gotten the guest to let go yet
21:09:19 gouthamr yes
21:10:24 gouthamr might be my browser, but i don't see "rbd.ImageBusy: [errno 16] RBD image is busy (error removing image)" in the latest n-cpu logs
21:10:30 gouthamr should i be looking elsewhere?
21:11:09 dansmith yeah actually I don't think I see them in the latest either.. but that was just a recheck
21:11:23 dansmith almost identical set of failures though
21:11:28 dansmith so yeah.. weird
21:13:43 dansmith gouthamr: here's an example from the previous run: https://zuul.opendev.org/t/openstack/build/f2ecbdd78616419cb5c8c2b3f4a8b71a/log/controller/logs/screen-n-cpu.txt#55439
21:15:19 dansmith that's all over those logs and absent from the latest.. bizarre
21:17:14 dansmith gouthamr: cephadm job made it past cephadm install phase, so.. progress I think
21:17:23 gouthamr very nice
21:20:57 dansmith and finished pool setup (ignorant guess from the commands it ran)
21:22:43 gouthamr https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/devstack/files/debs/devstack-plugin-ceph breaks focal fossa jobs (cephfs-native) though - but, we can fix that up with some os version annotation, correct?
21:23:27 dansmith we can fix it by putting it in the code instead of those package lists
21:23:42 dansmith I just jammed it in there because it was hard to mess up and will make sure we get it installed from the distro
21:24:04 gouthamr ack; either that or i can just fix the native cephfs job to use jammy too
21:24:09 dansmith the current install_podman thing does checks for focal, so I'd just extend that
21:24:32 dansmith gouthamr: yeah, although the ceph jobs on stable have to use the plugin without branches don't they?
21:24:46 gouthamr no this repo is branched
21:24:54 dansmith ah okay
21:25:30 dansmith well, either way.. if we move this to only >=jammy for everything (along with the PTI for 2023.2) then we can just use this debs list thing and remove the focal-specific install stuff.. whatever you ant
21:25:33 dansmith I just want it to work :)
21:25:50 gouthamr agree; lets see this work
21:27:14 gouthamr if the "rbd remove" thing fails again on this run with the error, i would suggest reporting a bug - we could _try_ this thing on centos-9-stream and see whether there's some weirdness in the distro packages
21:27:58 gouthamr but, i am nervous about that job's future with ubuntu since we've learned about the ceph community's stance
21:28:03 dansmith ack, well, based on how this is working, I'm hoping the cephadm will either "work" or "fail the same way" and then we can discount distro packages
21:28:11 gouthamr ++
21:28:27 dansmith it's running tempest now and looking identical to the distro version so far (i.e. hasn't failed but hasn't run volume tests yet)
21:28:46 dansmith so that's majorly better. failing with upstream bits is a minor win over failing with distro bits :P
21:29:03 eharney i have some work in flight currently around rbd ImageBusy errors... i wonder if this job is running some tests that were previously disabled?
21:29:31 dansmith eharney: shouldn't be, and one run failed with a bunch of them and then a single recheck failed similarly, but with no rbdbusy errors
21:31:40 gouthamr eharney: interesting, are there librbd changes that you're having to work around?
21:32:15 eharney gouthamr: not new ones, just working on fixing the class of errors around rbd images that can't be deleted that we've always had
21:34:57 gouthamr eharney: oh.. this error seems to have occurred multiple times in the libvirt rbd "remove image" call..
21:36:04 gouthamr i'm hoping no openstack code needs to change; we're hoping to support quincy with stable/wallaby downstream :D based on some testing of this stuff elsewhere
21:36:36 dansmith gouthamr: yeah I was going to ask earlier.. are we using quincy downstream such that we know it works? this set of failures has me concerned that it's something fundamental of course
21:37:14 gouthamr (same tests, different OS, full blown ceph cluster as opposed to our aio .. etc etc -- so there could be a number of things being issues)
21:38:41 gouthamr dansmith: yes, we're not testing quincy with openstack's trunk though... we trail downstream; but this stuff is working with zed last i checked, with the same tempest tests passing there
21:38:53 dansmith okay that's good
21:56:15 dansmith well, it's passing some volume tests at least
21:56:29 gouthamr ++
21:57:22 gouthamr a couple of things we wanted to try in this cephadm job in the past:
21:57:44 dansmith gouthamr: so if this works magically, you're okay just making this drop support for focal as long as the other jobs here are set to run on jammy?
21:57:46 gouthamr (1) revert to default test concurrency --- the concurrency was set to 1 because we saw resource contention
21:58:00 gouthamr dansmith: yes
21:58:02 dansmith resource contention like memory/
21:58:10 gouthamr yes, and disk
21:58:26 dansmith so there's a tweak in devstack for memory that has been helping a lot of jobs, and we run with it enabled in the nova ceph jobs
21:58:39 gouthamr oh? i'd love to know!
21:58:43 dansmith drops mysql usage by about half, which is ~400MiB on these jobs
21:58:47 gouthamr nice
21:58:55 dansmith gouthamr: problem is you have to pay me a royalty per job execution to use it
21:59:17 dansmith gouthamr: https://github.com/openstack/nova/blob/master/.zuul.yaml#L422
21:59:33 gouthamr :D i'd be poor if i was betting on ceph using less memory ever
21:59:38 dansmith haha
22:00:03 gouthamr (2) there's also a way to turn off "cephadm" after the install -- we should set that option on the CI - there's no use to keep it running
22:00:06 dansmith so based on our success with that, if you're cool with it, I'd say we turn that on for these jobs anyway
22:00:24 dansmith gouthamr: meaning before we start running tempest?
22:01:30 gouthamr yes, the plugin will do that for you after the ceph cluster deployment is done: https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/cephadm#L28
22:02:12 dansmith oh so why is that not being set?
22:02:33 dansmith I don't even know what that means.. cephadm is a tool I thought, but you're saying it stays running even though the services are started or something?
22:02:36 gouthamr (here's an example from the only good cephadm job at the moment: https://github.com/openstack/manila-tempest-plugin/blob/ad0db6359c8c51c4521ac6660c8014981b2f1dea/zuul.d/manila-tempest-jobs.yaml#L413)
22:02:49 dansmith ack
22:02:56 gouthamr yes; we need it for day2 operations on the cluster
22:03:07 gouthamr so if you're running this on your local devstack, it's useful
22:03:13 dansmith ah, sure, okay
22:03:24 dansmith we should set that in these jobs that everyone inherits from
22:03:36 gouthamr ++
22:03:42 dansmith so maybe a follow-on to this to set that, the mysql thing, and make this job voting
22:04:00 dansmith you know, if and when it starts working :)
22:04:40 gouthamr ^^ +1
22:06:02 dansmith hmm, gouthamr I just noticed that we're still set to release=pacific for the distro-package-based job
22:06:11 dansmith is that maybe setting stuff that is non-ideal for quincy that could be related?
22:06:38 gouthamr i think the plugin ignores it, let me check
22:06:43 dansmith okay
22:06:59 dansmith https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/.zuul.yaml#46
22:08:43 dansmith okay just used for package repos anyway it looks like
22:08:51 gouthamr ack; we should get that opt out of the job because its confusing
22:08:53 gouthamr https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/ceph#L981-L989
22:09:43 dansmith ack, so I assume we're configuring a repo but just not installing anything from it on jammy, right?
22:10:02 dansmith I'll add it to my follow-on patch to add the optimization devstack vars
22:10:08 dansmith add .. the removal of it, I mean :)
22:10:13 gouthamr ack ty

Earlier   Later