| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-25 | |||
| 22:20:17 | gouthamr | they've published v17.2.6 today, and v17.2.5 a month ago | |
| 22:21:01 | gouthamr | https://quay.io/repository/ceph/ceph?tab=tags | |
| 22:21:45 | dansmith | ack, I hate this sort of "version minesweeper" game.. if we're that sensitive to version, it feels like we're doing something wrong | |
| 22:22:39 | gouthamr | agreed; but since this stuff hasn't worked before on our ci, its worth a try | |
| 22:22:58 | dansmith | yeah for sure | |
| 22:23:18 | dansmith | volume resize passed | |
| 22:23:32 | dansmith | maybe we'll find that this is fewer fails or something | |
| 22:25:18 | dansmith | actually, that one didn't fail before | |
| 22:28:07 | gouthamr | okay, this may be good news? devstack-plugin-ceph-tempest-py3 is going to pass | |
| 22:28:34 | dansmith | no, really? | |
| 22:28:52 | dansmith | third recheck's a charm? | |
| 22:28:55 | gouthamr | :D | |
| 22:42:28 | dansmith | more fails on this cephadm job | |
| 22:42:48 | dansmith | so I guess it's not something fundamental, but maybe just massively less stable or we're hitting some race easier? | |
| 22:43:26 | dansmith | maybe it is memory-related and we're stressed more here | |
| 22:43:41 | dansmith | maybe I should try turning on the two optimizations here to see if that makes things more stable | |
| 22:43:54 | gouthamr | ++ | |
| 22:45:47 | gouthamr | "DISABLE_CEPHADM_POST_DEPLOY: true" and "MYSQL_REDUCE_MEMORY: true" for the rescue; we can iterate after with the concurrency if these failures reduce | |
| 22:45:54 | gouthamr | or go away | |
| 22:46:11 | opendevreview | Artom Lifshitz proposed openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967 | |
| 22:46:12 | opendevreview | Artom Lifshitz proposed openstack/nova master: Save cell socket correctly when updating host NUMA topology https://review.opendev.org/c/openstack/nova/+/862964 | |
| 22:46:21 | dansmith | gouthamr: ack, will put those in here and see | |
| 22:53:15 | opendevreview | Merged openstack/nova master: Remove focal job for 2023.2 https://review.opendev.org/c/openstack/nova/+/881409 | |
| 23:25:39 | dansmith | gouthamr: okay, it's off and running with those flags | |
| 23:25:53 | dansmith | I'm burnt out so I'll circle back tomorrow | |
| 23:27:32 | gouthamr | dansmith++ works; good evening! :) | |
| #openstack-nova - 2023-04-26 | |||
| 00:01:40 | opendevreview | Merged openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967 | |
| 07:55:12 | gibi | auniyal, bauzas, sean-k-mooney my recent finding about test not waiting for sshable before attach/detach are in https://bugs.launchpad.net/nova/+bug/1998148 but I have no cycles to push tempest changes so feel free to jump on it | |
| 07:56:33 | bauzas | gibi: ack, but I'm atm still digesting dansmith's and gouthamr efforts on ceph/Jammy jobs | |
| 08:38:25 | gibi | sure. I just wanted to be explict about the fact that I won't push a solution | |
| 08:51:41 | bauzas | gibi: I'll try to help this afternoon then | |
| 08:51:49 | bauzas | thanks for the findings again | |
| 11:31:16 | ykarel | Hi bauzas | |
| 11:31:29 | ykarel | reported https://blueprints.launchpad.net/nova/+spec/libvirt-tb-cache-size as discussed yesterday, can you please check | |
| 11:32:34 | bauzas | ykarel: ack, approving it then | |
| 11:34:00 | bauzas | and done | |
| 11:34:41 | ykarel | Thanks bauzas | |
| 11:34:55 | bauzas | np | |
| 11:37:23 | sean-k-mooney | ykarel: once we get the ceph/py38 issues resolved im fine with prioritsing getting this landed too | |
| 11:37:59 | ykarel | sean-k-mooney, ack | |
| 11:38:42 | opendevreview | Elod Illes proposed openstack/nova master: Drop py38 based zuul jobs https://review.opendev.org/c/openstack/nova/+/881339 | |
| 12:53:06 | dansmith | gibi: I'm going to look into that today | |
| 13:09:22 | dansmith | gibi: we're failing that rescue test 100% of the time with the new ceph job I'm trying to get working, so perhaps we're losing that race more often there | |
| 13:09:48 | sean-k-mooney | odd | |
| 13:09:56 | sean-k-mooney | i wonder what makes rescue speciel | |
| 13:10:19 | dansmith | oh we're failing plenty of other things 100% of the time | |
| 13:10:27 | dansmith | but yeah, I don't know | |
| 13:10:31 | dansmith | I'm working on it today | |
| 13:13:33 | gibi | I think something is wrongly set up in the rescuce tests regarding the sshable waiter | |
| 13:13:43 | gibi | hence the more frequent failure there | |
| 13:13:52 | gibi | dansmith: thanks for looking into it | |
| 13:15:04 | dansmith | gibi: yeah, some of the other failures are clearly after we have already attached and detached once, so I'm sure they're not sshable related, | |
| 13:15:16 | dansmith | but I haven't dug into the rescue ones yet | |
| 13:16:13 | opendevreview | Sylvain Bauza proposed openstack/nova master: Add a new policy for cold-migrate with host https://review.opendev.org/c/openstack/nova/+/881562 | |
| 13:16:53 | bauzas | dansmith: I can also help for the rescue action | |
| 13:17:01 | bauzas | so you can look at the ceph failures | |
| 13:17:15 | bauzas | I'm done with my small feature now | |
| 13:18:02 | dansmith | bauzas: no, they're related | |
| 13:18:30 | dansmith | bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job | |
| 13:19:07 | bauzas | dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6 | |
| 13:19:16 | bauzas | -ETOOMANYFAILURES | |
| 13:19:36 | dansmith | bauzas: right that's what I'm talking about | |
| 13:20:21 | bauzas | hmmm, then I'm lost out of context | |
| 13:20:36 | sean-k-mooney | so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point | |
| 13:20:55 | sean-k-mooney | basically im almost at a point where context switing for a day or two would be ok | |
| 13:20:57 | bauzas | I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob | |
| 13:21:19 | dansmith | my patch for the ceph job has gotten us almost all of the way there, | |
| 13:21:52 | dansmith | just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on | |
| 13:22:09 | dansmith | if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases | |
| 13:23:27 | bauzas | dansmith: ok, you're prioritary (because of the blocking gate), so cool with me | |
| 13:24:10 | dansmith | the gate is unblocked, so I think the pressure is off for the moment, but we obviously have to get this resolved ASAP | |
| 13:24:35 | dansmith | bauzas: maybe you could re+W these: https://review.opendev.org/c/openstack/nova/+/880632/5 | |
| 13:24:45 | dansmith | the bottom patch already merged, they just need a kick to get into the gate | |
| 13:25:26 | bauzas | dansmith: I did this kickass | |
| 13:26:45 | dansmith | thanks | |
| 13:27:27 | bauzas | fwiw, I'll send a status email saying then that the gate is back | |
| 13:28:06 | bauzas | are we sure all the py38-removal releases were reverted ? | |
| 13:29:45 | dansmith | well, we're passing jobs so, I think that's all we need to know | |
| 13:29:56 | frickler | pretty sure oslo bumps were not reverted yet. but only oslo.db was released | |
| 13:30:02 | dansmith | there have been other releases but they are not yet in requirements, so we're kinda on the edge | |
| 13:30:06 | dansmith | yeah that^ | |
| 14:44:59 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 15:01:48 | dansmith | yeah, so the rescue tests work for me locally, which I guess is good | |
| 15:01:56 | dansmith | however, the teardown does fail for me | |
| 15:02:29 | dansmith | which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation | |
| 15:03:50 | dansmith | sorry, s/rescue/rebuild/ | |
| 15:07:13 | bauzas | dansmith: bravo | |
| 15:09:47 | dansmith | locally that includes the validations, which is disabled in the job (for some reason) so I pushed up a patch to set those to enabled, | |
| 15:10:02 | dansmith | which will both likely slow it down but also maybe catch some other issues | |
| 15:10:53 | dansmith | that causes us to at least wait for the server to be created and sshable in some of the other cases, before we might try to attach a volume, but like I say, the failures I'm seeing don't seem like they'd be impacted there | |
| 15:11:15 | dansmith | and the regular jobs have those enabled, and we see detach failures there still | |
| 15:11:17 | dansmith | albeit not 100% | |
| 15:23:34 | dansmith | bauzas: this needs re-+W as well: https://review.opendev.org/c/openstack/nova/+/880633 | |
| 15:23:39 | dansmith | it's about to pass its recheck | |
| 15:24:07 | bauzas | dansmith: done | |
| 15:24:14 | dansmith | thanks | |
| 15:37:44 | opendevreview | Merged openstack/nova master: Remove silent failure to find a node on rebuild https://review.opendev.org/c/openstack/nova/+/880632 | |
| 15:38:15 | bauzas | :) | |
| 15:53:56 | dansmith | well, the distro-based ceph job just passed with validations turned on | |
| 15:54:20 | dansmith | the cephadm based job is definitely quite a bit slower for some reason, so we're still waiting for that | |