| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 22:44:11 | dansmith | gmann: is it possible to make that create_server() always wait? | |
| 22:44:59 | gmann | this one right? https://github.com/openstack/cinder-tempest-plugin/blob/d6989d3c1a31f2bebf94f2a7a8dac1d9eb788b1b/cinder_tempest_plugin/scenario/test_volume_encrypted.py#L86 | |
| 22:45:24 | dansmith | gmann: yeah, but I'm wondering if create_server() itself can wait so anything that uses it will wait | |
| 22:45:45 | gmann | dansmith: I think it make sense to wait for SSH by default in scenario tests | |
| 22:45:48 | dansmith | gmann: the current thinking is that a lot of our volume detach problems come from trying to attach a volume before an instance is booted | |
| 22:45:58 | dansmith | gmann: ack, okay | |
| 22:46:33 | gmann | yeah, and with nature of most of scenario tests, booting server fully is always good before trying things on that | |
| 22:47:20 | dansmith | yeah, okay lemme try to do that | |
| 22:47:52 | gmann | dansmith: ok. we need to create/pass the validation resources from scenario manager which is needed for SSH/ping | |
| 22:48:46 | gmann | or may be via get_remote_client | |
| 22:49:40 | dansmith | yeah | |
| 23:48:35 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-27 | |||
| 01:29:53 | dansmith | ugh, failed 102 on that last one | |
| 01:36:19 | dansmith | er, no, I suck at numbers.. 9, and all scenario, so must be related to my tempest change | |
| 03:59:16 | opendevreview | yatin proposed openstack/nova master: [DNM] Check detatch issue with lower tb size https://review.opendev.org/c/openstack/nova/+/881690 | |
| 08:17:21 | sean-k-mooney | gibi: bauzas: fun issue https://bugs.launchpad.net/os-vif/+bug/2017868 | |
| 08:17:41 | bauzas | ack | |
| 08:20:11 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 08:20:49 | sean-k-mooney | im still considering -2ing ^ | |
| 10:02:46 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement is_luks_inside_qcow2 funtion https://review.opendev.org/c/openstack/nova/+/854030 | |
| 10:03:18 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement encryption on backingStore https://review.opendev.org/c/openstack/nova/+/870012 | |
| 14:30:55 | opendevreview | ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| 15:57:12 | gouthamr | o/ dansmith - i found a "rbd image is busy" error in a different change parented on your devstack-plugin-change: https://zuul.opendev.org/t/openstack/build/278d411484764c55829c26cec2860bee/log/controller/logs/screen-n-cpu.txt#10742 | |
| 15:58:05 | gouthamr | im still looking, but, the cephfs job runs some scenario tests where it boots up nova VMs, attaches manila shares to them... these tests are failing to delete the VMs after | |
| 15:59:05 | gouthamr | the DELETE /servers/{id} call returns with 204, we wait in a loop, the server doesn't go away within 300 seconds.. and n-cpu logs don't have an error.. | |
| 16:01:29 | gouthamr | i don't know if the "rbd.ImageBusy" exception has anything to do with this silent failure | |
| 16:02:46 | gouthamr | it might | |
| 16:08:53 | dansmith | gouthamr: yeah I think eharney mentioned that it's been a pervasive problem, but I usually see it like once or twice and not a huge raft of them in one job when there's not something else going wrong | |
| 16:09:50 | gouthamr | i think in this job it's been consistently failing.. | |
| 16:10:16 | gouthamr | maybe i am missing a config opt here that you've set elsewhere | |
| 17:20:05 | opendevreview | sean mooney proposed openstack/os-vif master: [WIP] set default qos policy https://review.opendev.org/c/openstack/os-vif/+/881751 | |
| 22:52:45 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-28 | |||
| 07:47:09 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 12:09:24 | opendevreview | Stephen Finucane proposed openstack/nova master: libvirt: Remove unnecessary arg https://review.opendev.org/c/openstack/nova/+/881817 | |
| 12:14:33 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 12:30:01 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 12:30:28 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 12:34:58 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 15:59:30 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Preserve cached base images for failed resize ops https://review.opendev.org/c/openstack/nova/+/877410 | |
| 18:26:48 | opendevreview | Ghanshyam proposed openstack/nova-specs master: Re-propose "Policy service role spec" https://review.opendev.org/c/openstack/nova-specs/+/881880 | |
| 19:57:46 | dansmith | gouthamr: I think I'm getting close.. only one failure on the nova patch this last time, and it's yet another volume test that doesn't wait for sshable | |
| 19:57:54 | dansmith | about out of steam for the week, but maybe we'll have some good news on monday | |
| 19:59:17 | gouthamr | dansmith: o/ that's awesome... that's gone down from double digits, so very encouraging to know! | |
| 19:59:33 | dansmith | also, | |
| 19:59:45 | dansmith | we've been chasing a bunch of volume attach failures unrelated to ceph for a while | |
| 19:59:58 | dansmith | so these wait-for-sshable fixes will hopefully improve the stability of the other jobs too | |
| 20:01:21 | gouthamr | ^^ ++ ; i was following the pattern; i think we have at least one "wait for ssh" workaround in some manila scenario tests that can be removed because you're fixing up the base class we're using | |
| 20:01:59 | dansmith | gouthamr: well, we had to abandon the "make them all do it all the time" approach for now at least for a variety of reasons, | |
| 20:02:07 | dansmith | so don't remove any of that :) | |
| 20:02:56 | dansmith | I resorted to just pushing a patch to cinder-tempest-plugin to fix the tests there, and the base patch just makes it possible to wait_until=SSHABLE for the scenario tests | |
| 20:04:41 | gouthamr | i see | |
| 20:05:03 | dansmith | not my preference, but gotta make incremental progress :/ | |
| 20:05:16 | gouthamr | ack; i wanted to see if it was the same issue first too, so i set ENABLE_CEPH_NOVA=False, and the cephfs job was passing; now i'll pop that off and see if your fixes are good with those too | |
| 20:05:35 | dansmith | so the stack is currently nova->cinder-tempest->tempest->devstack-ceph-plugin ... it's thorny | |
| 20:06:17 | gouthamr | yes, and i extended the chain further :D with -->manila-tempest-plugin --> devstack-plugin-ceph | |
| 20:06:30 | dansmith | going for a world record? :) | |
| 20:07:00 | dansmith | life was so simple when everything fit on a single floppy .. | |
| 20:07:05 | gouthamr | lol :D if someone wasn't looking, they'd think we're building a new fancy feature | |
| 20:07:09 | dansmith | hehe | |
| 20:42:28 | opendevreview | Merged openstack/nova master: Fix a typo in this URL: https://docs.openstack.org/nova/latest/admin/availability-zones.html https://review.opendev.org/c/openstack/nova/+/878797 | |
| #openstack-nova - 2023-05-01 | |||
| 12:54:28 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:01:10 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:04:28 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:31:56 | dansmith | sean-k-mooney: so check this out: https://264a86dd518c5142b8a6-3b04393c7506ab488b4fde073fc22e36.ssl.cf1.rackcdn.com/881764/1/check/cinder-tempest-plugin-lvm-multiattach/d28afe7/testr_results.html | |
| 13:32:01 | dansmith | maybe kashyap too ^ | |
| 13:32:25 | dansmith | that almost kinda looks like we attached the disk to the guest and the kernel crashed while probing partitions/filesystems, right? | |
| 13:33:10 | dansmith | in sysfs_add_file_mode_ns() | |
| 13:33:49 | dansmith | or perhaps while creating sysfs entries for the disk perhaps? | |
| 16:30:56 | dansmith | eharney: you around by chance? | |
| 16:45:20 | eharney | dansmith: yes | |
| 16:45:43 | dansmith | eharney: so, I seem to have gotten the nova ceph job down to a single repeatable failure | |
| 16:45:56 | dansmith | it is in the volume extend test, and it fails during cleanup | |
| 16:46:25 | dansmith | it's trying to, I guess, detach the volume from the server before deleting it and then before deleting the volume | |
| 16:46:49 | dansmith | it does *not* fail locally, so I don't think it's something fundamentally broken with new ceph or anything like that | |
| 16:46:54 | dansmith | https://955f32f8268e5d475e65-6c8f4c6e546a0854b4c11cc7c78829ca.ssl.cf5.rackcdn.com/881585/6/check/nova-ceph-multistore/1ecb09a/testr_results.html | |
| 16:47:25 | eharney | dansmith: let me take a look through the logs | |
| 16:47:49 | eharney | detach failing isn't something i'm familiar with (other than it being mentioned here the other day) | |
| 16:48:04 | dansmith | I've just been tracing through the code looking for how this works and it seems to me like everything is working, but perhaps it's just legitimately that the guest doesn't let go of the volume when we detach after the resize happened | |
| 16:48:17 | dansmith | okay, volume detach is without a doubt our most common failure | |
| 16:48:34 | dansmith | eharney: yeah appreciate if you could see if you spot anything in the logs | |
| 17:00:13 | dansmith | I see the guest saw the size change on vdb and also mentions that it's resizing the filesystem on it... that looks like more than just the block device size change, | |
| 17:00:41 | dansmith | so I wonder if it is literally doing a resize2fs on it and that is still happening when we try to issue the detach and that gets us stuck | |
| 17:01:02 | dansmith | because we start the detach less than half a second after the resize happens | |
| 17:03:32 | eharney | dansmith: yeah, i was also just looking down the path of whether the libvirt block resize call is synchronous or not (it looks like nova assumes it is?) | |
| 17:04:47 | dansmith | eharney: synchronous with what? It's synchronous to libvirt from compute, but I don't know that it waits to return until it's delivered to the guest, but definitely stuff like resizing filesystems would happen after that returns | |
| 17:05:08 | dansmith | even still, the test polls for completion of the operation as far as nova is concerned, but the guest stuff would all be async | |
| 17:05:34 | eharney | i see | |
| 17:05:56 | dansmith | I think maybe I should make the test ssh to the guest and see if it's mounted, and perhaps try to unmount it before the test ends or something, to avoid racing with the detach | |
| 17:11:10 | dansmith | the resize happens before cirros is even done with its startup stuff, | |
| 17:11:36 | dansmith | so I wonder on a slow emulated guest we mount all the filesystems we find, which means it's mounted in the guest when we resize, so it does the resize2fs activity automatically | |
| 17:11:58 | dansmith | but on a fast non-nested local run, it finishes startup before we do the attach and thus doesn't end up with it mounted during the resize | |
| 17:12:29 | eharney | is resize2fs etc triggered by qemu-guest-agent? | |
| 17:13:08 | dansmith | could be.. does cirros have the guest agent in it? I assumed not | |
| 17:13:22 | eharney | i don't think so | |
| 17:13:25 | dansmith | on regular systems I've never had resize2fs triggered automatically for me, so I'm not really sure why that would happen | |
| 17:13:30 | dansmith | but it seems like it is here | |
| 17:14:00 | eharney | i would have guessed that resize2fs doesn't happen in these test jobs | |
| 17:14:10 | dansmith | me too | |
| 17:14:12 | dansmith | you see it in the output though right? | |
| 17:14:26 | eharney | no, where is that? | |