Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-25
22:15:38 gouthamr yes.. it was a head-scratcher; think vkmc and i noticed that our override to use download.ceph.com stopped working with focal at some point since ubuntu default-enabled the ubuntu ceph repos.. so dropping it made no difference, we ended up using/testing with the distro provided packages
22:15:43 gouthamr oh
22:16:19 gouthamr test_rebuild_server_with_volume_attached?
22:16:25 dansmith dunno yet
22:16:40 gouthamr ah
22:16:43 dansmith but it's in the rebuld group
22:16:46 dansmith yep
22:17:14 dansmith test_rebuild_server_with_volume_attached [430.305393s] ... FAILED
22:17:18 dansmith ugh
22:18:19 dansmith so the other variable here is the version of qemu and qemu's block-rbd driver are different in jammy of course, compared to what we've been testing
22:18:42 dansmith so could be a bug in one of those, especially since it's related to the detach in the guest
22:19:35 gouthamr ack; another thing to try would be to bump the ceph image to the latest quincy: https://github.com/openstack/devstack-plugin-ceph/blob/563cb5deeb21815ce0c62fa30249e85e886c783a/devstack/lib/cephadm#L32
22:20:02 dansmith okay
22:20:17 gouthamr they've published v17.2.6 today, and v17.2.5 a month ago
22:21:01 gouthamr https://quay.io/repository/ceph/ceph?tab=tags
22:21:45 dansmith ack, I hate this sort of "version minesweeper" game.. if we're that sensitive to version, it feels like we're doing something wrong
22:22:39 gouthamr agreed; but since this stuff hasn't worked before on our ci, its worth a try
22:22:58 dansmith yeah for sure
22:23:18 dansmith volume resize passed
22:23:32 dansmith maybe we'll find that this is fewer fails or something
22:25:18 dansmith actually, that one didn't fail before
22:28:07 gouthamr okay, this may be good news? devstack-plugin-ceph-tempest-py3 is going to pass
22:28:34 dansmith no, really?
22:28:52 dansmith third recheck's a charm?
22:28:55 gouthamr :D
22:42:28 dansmith more fails on this cephadm job
22:42:48 dansmith so I guess it's not something fundamental, but maybe just massively less stable or we're hitting some race easier?
22:43:26 dansmith maybe it is memory-related and we're stressed more here
22:43:41 dansmith maybe I should try turning on the two optimizations here to see if that makes things more stable
22:43:54 gouthamr ++
22:45:47 gouthamr "DISABLE_CEPHADM_POST_DEPLOY: true" and "MYSQL_REDUCE_MEMORY: true" for the rescue; we can iterate after with the concurrency if these failures reduce
22:45:54 gouthamr or go away
22:46:11 opendevreview Artom Lifshitz proposed openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967
22:46:12 opendevreview Artom Lifshitz proposed openstack/nova master: Save cell socket correctly when updating host NUMA topology https://review.opendev.org/c/openstack/nova/+/862964
22:46:21 dansmith gouthamr: ack, will put those in here and see
22:53:15 opendevreview Merged openstack/nova master: Remove focal job for 2023.2 https://review.opendev.org/c/openstack/nova/+/881409
23:25:39 dansmith gouthamr: okay, it's off and running with those flags
23:25:53 dansmith I'm burnt out so I'll circle back tomorrow
23:27:32 gouthamr dansmith++ works; good evening! :)
#openstack-nova - 2023-04-26
00:01:40 opendevreview Merged openstack/nova master: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/862967
07:55:12 gibi auniyal, bauzas, sean-k-mooney my recent finding about test not waiting for sshable before attach/detach are in https://bugs.launchpad.net/nova/+bug/1998148 but I have no cycles to push tempest changes so feel free to jump on it
07:56:33 bauzas gibi: ack, but I'm atm still digesting dansmith's and gouthamr efforts on ceph/Jammy jobs
08:38:25 gibi sure. I just wanted to be explict about the fact that I won't push a solution
08:51:41 bauzas gibi: I'll try to help this afternoon then
08:51:49 bauzas thanks for the findings again
11:31:16 ykarel Hi bauzas
11:31:29 ykarel reported https://blueprints.launchpad.net/nova/+spec/libvirt-tb-cache-size as discussed yesterday, can you please check
11:32:34 bauzas ykarel: ack, approving it then
11:34:00 bauzas and done
11:34:41 ykarel Thanks bauzas
11:34:55 bauzas np
11:37:23 sean-k-mooney ykarel: once we get the ceph/py38 issues resolved im fine with prioritsing getting this landed too
11:37:59 ykarel sean-k-mooney, ack
11:38:42 opendevreview Elod Illes proposed openstack/nova master: Drop py38 based zuul jobs https://review.opendev.org/c/openstack/nova/+/881339
12:53:06 dansmith gibi: I'm going to look into that today
13:09:22 dansmith gibi: we're failing that rescue test 100% of the time with the new ceph job I'm trying to get working, so perhaps we're losing that race more often there
13:09:48 sean-k-mooney odd
13:09:56 sean-k-mooney i wonder what makes rescue speciel
13:10:19 dansmith oh we're failing plenty of other things 100% of the time
13:10:27 dansmith but yeah, I don't know
13:10:31 dansmith I'm working on it today
13:13:33 gibi I think something is wrongly set up in the rescuce tests regarding the sshable waiter
13:13:43 gibi hence the more frequent failure there
13:13:52 gibi dansmith: thanks for looking into it
13:15:04 dansmith gibi: yeah, some of the other failures are clearly after we have already attached and detached once, so I'm sure they're not sshable related,
13:15:16 dansmith but I haven't dug into the rescue ones yet
13:16:13 opendevreview Sylvain Bauza proposed openstack/nova master: Add a new policy for cold-migrate with host https://review.opendev.org/c/openstack/nova/+/881562
13:16:53 bauzas dansmith: I can also help for the rescue action
13:17:01 bauzas so you can look at the ceph failures
13:17:15 bauzas I'm done with my small feature now
13:18:02 dansmith bauzas: no, they're related
13:18:30 dansmith bauzas: they pass all the time (except for the occasional failure) on the focal ceph job, and 100% of the time on the jammy/quincy job
13:19:07 bauzas dansmith: sorry I meant about https://bugs.launchpad.net/nova/+bug/1998148/comments/6
13:19:16 bauzas -ETOOMANYFAILURES
13:19:36 dansmith bauzas: right that's what I'm talking about
13:20:21 bauzas hmmm, then I'm lost out of context
13:20:36 sean-k-mooney so i have mostly got the downstream ting i was workign on in a mostly mergable state. i need to respond to some feedback but i can try and deploy with ceph later today or tormoow and take a look if its still blocking at that point
13:20:55 sean-k-mooney basically im almost at a point where context switing for a day or two would be ok
13:20:57 bauzas I eventually tried this morning to look at the ceph patch by the failing logs, but honestly I'm a noob
13:21:19 dansmith my patch for the ceph job has gotten us almost all of the way there,
13:21:52 dansmith just give me some time this morning to try to hack in an sshable wait to some of the tests that are failing 100% of the time (if applicable) and then we can go from there.. if they're not applicable, then we've got something worse going on
13:22:09 dansmith if they are, maybe we're just losing that race more often now and this will help the ceph and non-ceph cases
13:23:27 bauzas dansmith: ok, you're prioritary (because of the blocking gate), so cool with me
13:24:10 dansmith the gate is unblocked, so I think the pressure is off for the moment, but we obviously have to get this resolved ASAP
13:24:35 dansmith bauzas: maybe you could re+W these: https://review.opendev.org/c/openstack/nova/+/880632/5
13:24:45 dansmith the bottom patch already merged, they just need a kick to get into the gate
13:25:26 bauzas dansmith: I did this kickass
13:26:45 dansmith thanks
13:27:27 bauzas fwiw, I'll send a status email saying then that the gate is back
13:28:06 bauzas are we sure all the py38-removal releases were reverted ?
13:29:45 dansmith well, we're passing jobs so, I think that's all we need to know
13:29:56 frickler pretty sure oslo bumps were not reverted yet. but only oslo.db was released
13:30:02 dansmith there have been other releases but they are not yet in requirements, so we're kinda on the edge
13:30:06 dansmith yeah that^
14:44:59 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
15:01:48 dansmith yeah, so the rescue tests work for me locally, which I guess is good
15:01:56 dansmith however, the teardown does fail for me
15:02:29 dansmith which tells me, again, that I don't think ssh'able is the solution for those, because we've already attached a volume and rebooted the instance (into rescue) and then rebooted it back, long after the original attach operation
15:03:50 dansmith sorry, s/rescue/rebuild/
15:07:13 bauzas dansmith: bravo

Earlier   Later