| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 13:16:53 | bauzas | dansmith: I can also help for the rescue action | |
| 13:17:01 | bauzas | so you can look at the ceph failures | |
| 13:17:15 | bauzas | I'm done with my small feature now | |
| 13:18:02 | dansmith | bauzas: no, they're related | |
| 13:18:30 | dansmith | bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job | |
| 13:19:07 | bauzas | dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6 | |
| 13:19:16 | bauzas | -ETOOMANYFAILURES | |
| 13:19:36 | dansmith | bauzas: right that's what I'm talking about | |
| 13:20:21 | bauzas | hmmm, then I'm lost out of context | |
| 13:20:36 | sean-k-mooney | so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point | |
| 13:20:55 | sean-k-mooney | basically im almost at a point where context switing for a day or two would be ok | |
| 13:20:57 | bauzas | I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob | |
| 13:21:19 | dansmith | my patch for the ceph job has gotten us almost all of the way there, | |
| 13:21:52 | dansmith | just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on | |
| 13:22:09 | dansmith | if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases | |
| 13:23:27 | bauzas | dansmith: ok, you're prioritary (because of the blocking gate), so cool with me | |
| 13:24:10 | dansmith | the gate is unblocked, so I think the pressure is off for the moment, but we obviously have to get this resolved ASAP | |
| 13:24:35 | dansmith | bauzas: maybe you could re+W these: https://review.opendev.org/c/openstack/nova/+/880632/5 | |
| 13:24:45 | dansmith | the bottom patch already merged, they just need a kick to get into the gate | |
| 13:25:26 | bauzas | dansmith: I did this kickass | |
| 13:26:45 | dansmith | thanks | |
| 13:27:27 | bauzas | fwiw, I'll send a status email saying then that the gate is back | |
| 13:28:06 | bauzas | are we sure all the py38-removal releases were reverted ? | |
| 13:29:45 | dansmith | well, we're passing jobs so, I think that's all we need to know | |
| 13:29:56 | frickler | pretty sure oslo bumps were not reverted yet. but only oslo.db was released | |
| 13:30:02 | dansmith | there have been other releases but they are not yet in requirements, so we're kinda on the edge | |
| 13:30:06 | dansmith | yeah that^ | |
| 14:44:59 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 15:01:48 | dansmith | yeah, so the rescue tests work for me locally, which I guess is good | |
| 15:01:56 | dansmith | however, the teardown does fail for me | |
| 15:02:29 | dansmith | which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation | |
| 15:03:50 | dansmith | sorry, s/rescue/rebuild/ | |
| 15:07:13 | bauzas | dansmith: bravo | |
| 15:09:47 | dansmith | locally that includes the validations, which is disabled in the job (for some reason) so I pushed up a patch to set those to enabled, | |
| 15:10:02 | dansmith | which will both likely slow it down but also maybe catch some other issues | |
| 15:10:53 | dansmith | that causes us to at least wait for the server to be created and sshable in some of the other cases, before we might try to attach a volume, but like I say, the failures I'm seeing don't seem like they'd be impacted there | |
| 15:11:15 | dansmith | and the regular jobs have those enabled, and we see detach failures there still | |
| 15:11:17 | dansmith | albeit not 100% | |
| 15:23:34 | dansmith | bauzas: this needs re-+W as well: https://review.opendev.org/c/openstack/nova/+/880633 | |
| 15:23:39 | dansmith | it's about to pass its recheck | |
| 15:24:07 | bauzas | dansmith: done | |
| 15:24:14 | dansmith | thanks | |
| 15:37:44 | opendevreview | Merged openstack/nova master: Remove silent failure to find a node on rebuild https://review.opendev.org/c/openstack/nova/+/880632 | |
| 15:38:15 | bauzas | :) | |
| 15:53:56 | dansmith | well, the distro-based ceph job just passed with validations turned on | |
| 15:54:20 | dansmith | the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that | |
| 16:09:36 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| 16:19:19 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| 16:21:17 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| 16:24:28 | dansmith | okay it timed out, but legitimately, it was making progress just taking much longer because of the validations I think | |
| 16:24:48 | dansmith | so we might need to trim that job down again or we can try with the concurrency upped to 2 | |
| 16:26:41 | dansmith | ah, the base job is not concurrency=1, so that's probably the difference there.. might be getting close then! | |
| 16:29:18 | dansmith | gouthamr: ^ | |
| 16:52:31 | opendevreview | ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| 17:01:43 | opendevreview | Merged openstack/nova master: Stop ignoring missing compute nodes in claims https://review.opendev.org/c/openstack/nova/+/880633 | |
| 17:46:12 | dansmith | eharney: gouthamr: the cephadm run just passed.. the distro package one failed with errors that look very different than before, but related to ImageBusy type things | |
| 17:46:52 | dansmith | the difference here is enabling validations on both, which means we try to ssh to instances before we do certain things to them, which helps a lot with race conditions. I dunno why those were disabled on these jobs before, but we thought they were enabled | |
| 17:47:47 | dansmith | so my plan is to: (1) switch the cephadm job to be the voting primary job and make the distro one non-voting, (2) set the nova DNM to run against cephadm, and get another run of everything | |
| 17:48:31 | dansmith | is there anything else you want to change before we could/should merge this? the cephfs job is still on focal (because it's defined elsewhere) and thus failing because of packages, but I could hard-code that to jammy here to "fix" that if you want | |
| 17:48:44 | dansmith | it's non-voting so we should be able to switch it externally separately, but let me know what you want | |
| 17:50:22 | dansmith | oh another change on the cephadm job was removing the concurrency limit, which was causing it to run too slow to finish when validations was turned on. We need that to run in parallel anyway, so it's good that it passed in that form | |
| 17:53:58 | dansmith | the cephadm job also finished pretty quickly, which is also a good sign that it was pretty healthy | |
| 18:15:11 | dansmith | ah, the ceph-osd was oom-killed in the non-cephadm job, so that explains that failure I think | |
| 18:16:37 | sean-k-mooney | that would do it | |
| 18:16:55 | sean-k-mooney | it makes snese why you are siing rbd busy messages | |
| 18:17:21 | dansmith | yeah | |
| 18:17:48 | sean-k-mooney | is this with the mariadb and ceph options to reduce mememory set | |
| 18:18:03 | sean-k-mooney | and idealy the job configred for 8GB of swap? | |
| 18:18:31 | dansmith | it does include the mysql and ceph tweaks, not sure about swap, | |
| 18:18:45 | sean-k-mooney | i think we default to 2GB instead of 8 | |
| 18:18:49 | dansmith | but what I care about is the cephadm job at this point I think, so I kinda want another data point before we change too much else | |
| 18:19:01 | sean-k-mooney | ack | |
| 18:19:03 | dansmith | I know | |
| 18:19:27 | dansmith | I'm waiting for the nova one to finish and see if it's the same, | |
| 18:19:38 | dansmith | but also to switch it to cephadm and get another set of data | |
| 18:19:49 | dansmith | but this is the first time I've ever seen the cephadm job pass, so I'm encouraged | |
| 18:20:46 | sean-k-mooney | if memory continue to be an issue we have 3 paths forward, 1.) increase swap. 2.) use the nested virt nodeset, 3.) wait for t he qemu cache bluepitn to be merged and set it back to 32mb instead of 1G per vm | |
| 18:21:29 | dansmith | we can also try reducing concurrency.. =1 times out because it takes too long, but we're at =4 right now | |
| 18:21:46 | dansmith | the cephadm job ran at =4 at top speed with no failures this time | |
| 18:22:00 | sean-k-mooney | ya settign it to 2 or 3 might sticke a better ballance for jobs with ceph | |
| 18:22:16 | sean-k-mooney | cool | |
| 18:23:09 | sean-k-mooney | your waiting fr this to complete ya https://zuul.openstack.org/stream/677a38a3278c4ced9deee595bb616997?logfile=console.log | |
| 18:23:19 | dansmith | yeah | |
| 18:23:30 | dansmith | I want the logs so I can look while I run it against cephadm as the base | |
| 18:23:47 | dansmith | hoping it's also the same issue as the other one because they look very similar, just watching them | |
| 18:23:58 | sean-k-mooney | an hour and 19 minuts is not bad | |
| 18:24:08 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/7493255ead1b4c7abbc2ba45c8f1d9d4 | |
| 18:24:22 | dansmith | right, it's super fast actually | |
| 18:24:33 | sean-k-mooney | its also singel node which helps | |
| 18:24:44 | dansmith | so I'm hoping that the newer ceph is somehow making that smoother/faster/leaner :) | |
| 18:24:50 | dansmith | compared to the distro jobs | |
| 18:25:12 | sean-k-mooney | well going to python 3.10 wil add about 30% to the openstack performace | |
| 18:25:20 | sean-k-mooney | vs 3.8 i think | |
| 18:25:29 | sean-k-mooney | oh | |
| 18:25:44 | sean-k-mooney | you are comparing to disto packages on 22.04 not focal | |
| 18:25:51 | dansmith | yeah, so that might be helping the speed when it doesn't fail | |
| 18:25:57 | dansmith | sean-k-mooney: correct | |
| 18:26:15 | dansmith | the cephadm job takes ceph via podman containers from ceph directly, the nova and -py3 jobs are using jammy's ceph packages | |
| 18:26:33 | sean-k-mooney | cool | |
| 18:26:50 | sean-k-mooney | so wait are you finding containerisation useful for something :P | |