Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-26
22:49:40 dansmith yeah
23:48:35 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
#openstack-nova - 2023-04-27
01:29:53 dansmith ugh, failed 102 on that last one
01:36:19 dansmith er, no, I suck at numbers.. 9, and all scenario, so must be related to my tempest change
03:59:16 opendevreview yatin proposed openstack/nova master: [DNM] Check detatch issue with lower tb size https://review.opendev.org/c/openstack/nova/+/881690
08:17:21 sean-k-mooney gibi: bauzas: fun issue https://bugs.launchpad.net/os-vif/+bug/2017868
08:17:41 bauzas ack
08:20:11 opendevreview Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644
08:20:49 sean-k-mooney im still considering -2ing ^
10:02:46 opendevreview Sofia Enriquez proposed openstack/nova master: Implement is_luks_inside_qcow2 funtion https://review.opendev.org/c/openstack/nova/+/854030
10:03:18 opendevreview Sofia Enriquez proposed openstack/nova master: Implement encryption on backingStore https://review.opendev.org/c/openstack/nova/+/870012
14:30:55 opendevreview ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773
15:57:12 gouthamr o/ dansmith - i found a "rbd image is busy" error in a different change parented on your devstack-plugin-change: https://zuul.opendev.org/t/openstack/build/278d411484764c55829c26cec2860bee/log/controller/logs/screen-n-cpu.txt#10742
15:58:05 gouthamr im still looking, but, the cephfs job runs some scenario tests where it boots up nova VMs, attaches manila shares to them... these tests are failing to delete the VMs after
15:59:05 gouthamr the DELETE /servers/{id} call returns with 204, we wait in a loop, the server doesn't go away within 300 seconds.. and n-cpu logs don't have an error..
16:01:29 gouthamr i don't know if the "rbd.ImageBusy" exception has anything to do with this silent failure
16:02:46 gouthamr it might
16:08:53 dansmith gouthamr: yeah I think eharney mentioned that it's been a pervasive problem, but I usually see it like once or twice and not a huge raft of them in one job when there's not something else going wrong
16:09:50 gouthamr i think in this job it's been consistently failing..
16:10:16 gouthamr maybe i am missing a config opt here that you've set elsewhere
17:20:05 opendevreview sean mooney proposed openstack/os-vif master: [WIP] set default qos policy https://review.opendev.org/c/openstack/os-vif/+/881751
22:52:45 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
#openstack-nova - 2023-04-28
07:47:09 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
12:09:24 opendevreview Stephen Finucane proposed openstack/nova master: libvirt: Remove unnecessary arg https://review.opendev.org/c/openstack/nova/+/881817
12:14:33 opendevreview yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419
12:30:01 opendevreview Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517
12:30:28 opendevreview Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517
12:34:58 opendevreview Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517
15:59:30 opendevreview Alexey Stupnikov proposed openstack/nova master: Preserve cached base images for failed resize ops https://review.opendev.org/c/openstack/nova/+/877410
18:26:48 opendevreview Ghanshyam proposed openstack/nova-specs master: Re-propose "Policy service role spec" https://review.opendev.org/c/openstack/nova-specs/+/881880
19:57:46 dansmith gouthamr: I think I'm getting close.. only one failure on the nova patch this last time, and it's yet another volume test that doesn't wait for sshable
19:57:54 dansmith about out of steam for the week, but maybe we'll have some good news on monday
19:59:17 gouthamr dansmith: o/ that's awesome... that's gone down from double digits, so very encouraging to know!
19:59:33 dansmith also,
19:59:45 dansmith we've been chasing a bunch of volume attach failures unrelated to ceph for a while
19:59:58 dansmith so these wait-for-sshable fixes will hopefully improve the stability of the other jobs too
20:01:21 gouthamr ^^ ++ ; i was following the pattern; i think we have at least one "wait for ssh" workaround in some manila scenario tests that can be removed because you're fixing up the base class we're using
20:01:59 dansmith gouthamr: well, we had to abandon the "make them all do it all the time" approach for now at least for a variety of reasons,
20:02:07 dansmith so don't remove any of that :)
20:02:56 dansmith I resorted to just pushing a patch to cinder-tempest-plugin to fix the tests there, and the base patch just makes it possible to wait_until=SSHABLE for the scenario tests
20:04:41 gouthamr i see
20:05:03 dansmith not my preference, but gotta make incremental progress :/
20:05:16 gouthamr ack; i wanted to see if it was the same issue first too, so i set ENABLE_CEPH_NOVA=False, and the cephfs job was passing; now i'll pop that off and see if your fixes are good with those too
20:05:35 dansmith so the stack is currently nova->cinder-tempest->tempest->devstack-ceph-plugin ... it's thorny
20:06:17 gouthamr yes, and i extended the chain further :D with -->manila-tempest-plugin --> devstack-plugin-ceph
20:06:30 dansmith going for a world record? :)
20:07:00 dansmith life was so simple when everything fit on a single floppy ..
20:07:05 gouthamr lol :D if someone wasn't looking, they'd think we're building a new fancy feature
20:07:09 dansmith hehe
20:42:28 opendevreview Merged openstack/nova master: Fix a typo in this URL: https://docs.openstack.org/nova/latest/admin/availability-zones.html https://review.opendev.org/c/openstack/nova/+/878797
#openstack-nova - 2023-05-01
12:54:28 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:01:10 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:04:28 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:31:56 dansmith sean-k-mooney: so check this out: https://264a86dd518c5142b8a6-3b04393c7506ab488b4fde073fc22e36.ssl.cf1.rackcdn.com/881764/1/check/cinder-tempest-plugin-lvm-multiattach/d28afe7/testr_results.html
13:32:01 dansmith maybe kashyap too ^
13:32:25 dansmith that almost kinda looks like we attached the disk to the guest and the kernel crashed while probing partitions/filesystems, right?
13:33:10 dansmith in sysfs_add_file_mode_ns()
13:33:49 dansmith or perhaps while creating sysfs entries for the disk perhaps?
16:30:56 dansmith eharney: you around by chance?
16:45:20 eharney dansmith: yes
16:45:43 dansmith eharney: so, I seem to have gotten the nova ceph job down to a single repeatable failure
16:45:56 dansmith it is in the volume extend test, and it fails during cleanup
16:46:25 dansmith it's trying to, I guess, detach the volume from the server before deleting it and then before deleting the volume
16:46:49 dansmith it does *not* fail locally, so I don't think it's something fundamentally broken with new ceph or anything like that
16:46:54 dansmith https://955f32f8268e5d475e65-6c8f4c6e546a0854b4c11cc7c78829ca.ssl.cf5.rackcdn.com/881585/6/check/nova-ceph-multistore/1ecb09a/testr_results.html
16:47:25 eharney dansmith: let me take a look through the logs
16:47:49 eharney detach failing isn't something i'm familiar with (other than it being mentioned here the other day)
16:48:04 dansmith I've just been tracing through the code looking for how this works and it seems to me like everything is working, but perhaps it's just legitimately that the guest doesn't let go of the volume when we detach after the resize happened
16:48:17 dansmith okay, volume detach is without a doubt our most common failure
16:48:34 dansmith eharney: yeah appreciate if you could see if you spot anything in the logs
17:00:13 dansmith I see the guest saw the size change on vdb and also mentions that it's resizing the filesystem on it... that looks like more than just the block device size change,
17:00:41 dansmith so I wonder if it is literally doing a resize2fs on it and that is still happening when we try to issue the detach and that gets us stuck
17:01:02 dansmith because we start the detach less than half a second after the resize happens
17:03:32 eharney dansmith: yeah, i was also just looking down the path of whether the libvirt block resize call is synchronous or not (it looks like nova assumes it is?)
17:04:47 dansmith eharney: synchronous with what? It's synchronous to libvirt from compute, but I don't know that it waits to return until it's delivered to the guest, but definitely stuff like resizing filesystems would happen after that returns
17:05:08 dansmith even still, the test polls for completion of the operation as far as nova is concerned, but the guest stuff would all be async
17:05:34 eharney i see
17:05:56 dansmith I think maybe I should make the test ssh to the guest and see if it's mounted, and perhaps try to unmount it before the test ends or something, to avoid racing with the detach
17:11:10 dansmith the resize happens before cirros is even done with its startup stuff,
17:11:36 dansmith so I wonder on a slow emulated guest we mount all the filesystems we find, which means it's mounted in the guest when we resize, so it does the resize2fs activity automatically
17:11:58 dansmith but on a fast non-nested local run, it finishes startup before we do the attach and thus doesn't end up with it mounted during the resize
17:12:29 eharney is resize2fs etc triggered by qemu-guest-agent?
17:13:08 dansmith could be.. does cirros have the guest agent in it? I assumed not
17:13:22 eharney i don't think so
17:13:25 dansmith on regular systems I've never had resize2fs triggered automatically for me, so I'm not really sure why that would happen
17:13:30 dansmith but it seems like it is here
17:14:00 eharney i would have guessed that resize2fs doesn't happen in these test jobs
17:14:10 dansmith me too
17:14:12 dansmith you see it in the output though right?
17:14:26 eharney no, where is that?
17:14:27 dansmith [ 48.255755] EXT4-fs (vda1): resized filesystem to 259835
17:14:27 dansmith [ 48.156649] EXT4-fs (vda1): resizing filesystem from 25600 to 259835 blocks
17:14:40 dansmith in the guest console dump
17:14:45 dansmith oh damn
17:14:53 dansmith that is vda, nevermind!
17:15:08 dansmith that's cirros resizing its root disk on startup, not the attached volume
17:15:10 eharney ah, right
17:15:21 dansmith mah bad
17:17:19 dansmith so, without a doubt the most common failure in nova jobs is failing to detach volumes
17:17:33 dansmith we've been trying to get a handle on it for a long time,

Earlier   Later