| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-25 | |||
| 22:06:38 | gouthamr | i think the plugin ignores it, let me check | |
| 22:06:43 | dansmith | okay | |
| 22:06:59 | dansmith | https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/865315/11/.zuul.yaml#46 | |
| 22:08:43 | dansmith | okay just used for package repos anyway it looks like | |
| 22:08:51 | gouthamr | ack; we should get that opt out of the job because its confusing | |
| 22:08:53 | gouthamr | https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/ceph#L981-L989 | |
| 22:09:43 | dansmith | ack, so I assume we're configuring a repo but just not installing anything from it on jammy, right? | |
| 22:10:02 | dansmith | I'll add it to my follow-on patch to add the optimization devstack vars | |
| 22:10:08 | dansmith | add .. the removal of it, I mean :) | |
| 22:10:13 | gouthamr | ack ty | |
| 22:10:20 | gouthamr | i may have messed this up in my last patch | |
| 22:10:21 | gouthamr | https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/ceph#L1067-L1094 | |
| 22:10:47 | gouthamr | we're not invoking that method at all for ubuntu anymore.. | |
| 22:11:27 | dansmith | ack | |
| 22:11:45 | dansmith | well, there's a bunch of focal-specific stuff to clean up in there regardless | |
| 22:14:58 | dansmith | hmm, I think it's about to fail a test | |
| 22:15:11 | dansmith | been stopped for going on five minutes, assuming retrying a detach | |
| 22:15:38 | gouthamr | yes.. it was a head-scratcher; think vkmc and i noticed that our override to use download.ceph.com stopped working with focal at some point since ubuntu default-enabled the ubuntu ceph repos.. so dropping it made no difference, we ended up using/testing with the distro provided packages | |
| 22:15:43 | gouthamr | oh | |
| 22:16:19 | gouthamr | test_rebuild_server_with_volume_attached? | |
| 22:16:25 | dansmith | dunno yet | |
| 22:16:40 | gouthamr | ah | |
| 22:16:43 | dansmith | but it's in the rebuld group | |
| 22:16:46 | dansmith | yep | |
| 22:17:14 | dansmith | test_rebuild_server_with_volume_attached [430.305393s] ... FAILED | |
| 22:17:18 | dansmith | ugh | |
| 22:18:19 | dansmith | so the other variable here is the version of qemu and qemu's block-rbd driver are different in jammy of course, compared to what we've been testing | |
| 22:18:42 | dansmith | so could be a bug in one of those, especially since it's related to the detach in the guest | |
| 22:19:35 | gouthamr | ack; another thing to try would be to bump the ceph image to the latest quincy: https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/cephadm#L32 | |
| 22:20:02 | dansmith | okay | |
| 22:20:17 | gouthamr | they've published v17.2.6 today, and v17.2.5 a month ago | |
| 22:21:01 | gouthamr | https://quay.io/repository/ceph/ceph?tab=tags | |
| 22:21:45 | dansmith | ack, I hate this sort of "version minesweeper" game.. if we're that sensitive to version, it feels like we're doing something wrong | |
| 22:22:39 | gouthamr | agreed; but since this stuff hasn't worked before on our ci, its worth a try | |
| 22:22:58 | dansmith | yeah for sure | |
| 22:23:18 | dansmith | volume resize passed | |
| 22:23:32 | dansmith | maybe we'll find that this is fewer fails or something | |
| 22:25:18 | dansmith | actually, that one didn't fail before | |
| 22:28:07 | gouthamr | okay, this may be good news? devstack-plugin-ceph-tempest-py3 is going to pass | |
| 22:28:34 | dansmith | no, really? | |
| 22:28:52 | dansmith | third recheck's a charm? | |
| 22:28:55 | gouthamr | :D | |
| 22:42:28 | dansmith | more fails on this cephadm job | |
| 22:42:48 | dansmith | so I guess it's not something fundamental, but maybe just massively less stable or we're hitting some race easier? | |
| 22:43:26 | dansmith | maybe it is memory-related and we're stressed more here | |
| 22:43:41 | dansmith | maybe I should try turning on the two optimizations here to see if that makes things more stable | |
| 22:43:54 | gouthamr | ++ | |
| 22:45:47 | gouthamr | "DISABLE_CEPHADM_POST_DEPLOY: true" and "MYSQL_REDUCE_MEMORY: true" for the rescue; we can iterate after with the concurrency if these failures reduce | |
| 22:45:54 | gouthamr | or go away | |
| 22:46:11 | opendevreview | Artom Lifshitz proposed openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967 | |
| 22:46:12 | opendevreview | Artom Lifshitz proposed openstack/nova master: Save cell socket correctly when updating host NUMA topology https://review.opendev.org/c/openstack/nova/+/862964 | |
| 22:46:21 | dansmith | gouthamr: ack, will put those in here and see | |
| 22:53:15 | opendevreview | Merged openstack/nova master: Remove focal job for 2023.2 https://review.opendev.org/c/openstack/nova/+/881409 | |
| 23:25:39 | dansmith | gouthamr: okay, it's off and running with those flags | |
| 23:25:53 | dansmith | I'm burnt out so I'll circle back tomorrow | |
| 23:27:32 | gouthamr | dansmith++ works; good evening! :) | |
| #openstack-nova - 2023-04-26 | |||
| 00:01:40 | opendevreview | Merged openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967 | |
| 07:55:12 | gibi | auniyal, bauzas, sean-k-mooney my recent finding about test not waiting for sshable before attach/detach are in https://bugs.launchpad.net/nova/+bug/1998148 but I have no cycles to push tempest changes so feel free to jump on it | |
| 07:56:33 | bauzas | gibi: ack, but I'm atm still digesting dansmith's and gouthamr efforts on ceph/Jammy jobs | |
| 08:38:25 | gibi | sure. I just wanted to be explict about the fact that I won't push a solution | |
| 08:51:41 | bauzas | gibi: I'll try to help this afternoon then | |
| 08:51:49 | bauzas | thanks for the findings again | |
| 11:31:16 | ykarel | Hi bauzas | |
| 11:31:29 | ykarel | reported https://blueprints.launchpad.net/nova/+spec/libvirt-tb-cache-size as discussed yesterday, can you please check | |
| 11:32:34 | bauzas | ykarel: ack, approving it then | |
| 11:34:00 | bauzas | and done | |
| 11:34:41 | ykarel | Thanks bauzas | |
| 11:34:55 | bauzas | np | |
| 11:37:23 | sean-k-mooney | ykarel: once we get the ceph/py38 issues resolved im fine with prioritsing getting this landed too | |
| 11:37:59 | ykarel | sean-k-mooney, ack | |
| 11:38:42 | opendevreview | Elod Illes proposed openstack/nova master: Drop py38 based zuul jobs https://review.opendev.org/c/openstack/nova/+/881339 | |
| 12:53:06 | dansmith | gibi: I'm going to look into that today | |
| 13:09:22 | dansmith | gibi: we're failing that rescue test 100% of the time with the new ceph job I'm trying to get working, so perhaps we're losing that race more often there | |
| 13:09:48 | sean-k-mooney | odd | |
| 13:09:56 | sean-k-mooney | i wonder what makes rescue speciel | |
| 13:10:19 | dansmith | oh we're failing plenty of other things 100% of the time | |
| 13:10:27 | dansmith | but yeah, I don't know | |
| 13:10:31 | dansmith | I'm working on it today | |
| 13:13:33 | gibi | I think something is wrongly set up in the rescuce tests regarding the sshable waiter | |
| 13:13:43 | gibi | hence the more frequent failure there | |
| 13:13:52 | gibi | dansmith: thanks for looking into it | |
| 13:15:04 | dansmith | gibi: yeah, some of the other failures are clearly after we have already attached and detached once, so I'm sure they're not sshable related, | |
| 13:15:16 | dansmith | but I haven't dug into the rescue ones yet | |
| 13:16:13 | opendevreview | Sylvain Bauza proposed openstack/nova master: Add a new policy for cold-migrate with host https://review.opendev.org/c/openstack/nova/+/881562 | |
| 13:16:53 | bauzas | dansmith: I can also help for the rescue action | |
| 13:17:01 | bauzas | so you can look at the ceph failures | |
| 13:17:15 | bauzas | I'm done with my small feature now | |
| 13:18:02 | dansmith | bauzas: no, they're related | |
| 13:18:30 | dansmith | bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job | |
| 13:19:07 | bauzas | dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6 | |
| 13:19:16 | bauzas | -ETOOMANYFAILURES | |
| 13:19:36 | dansmith | bauzas: right that's what I'm talking about | |
| 13:20:21 | bauzas | hmmm, then I'm lost out of context | |
| 13:20:36 | sean-k-mooney | so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point | |
| 13:20:55 | sean-k-mooney | basically im almost at a point where context switing for a day or two would be ok | |
| 13:20:57 | bauzas | I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob | |
| 13:21:19 | dansmith | my patch for the ceph job has gotten us almost all of the way there, | |
| 13:21:52 | dansmith | just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on | |
| 13:22:09 | dansmith | if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases | |
| 13:23:27 | bauzas | dansmith: ok, you're prioritary (because of the blocking gate), so cool with me | |