| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-04-26 | |||
| 20:22:40 | dansmith | gouthamr: I have this all ready to go as soon as the nova job finishes to capture logs, so let me push it up as it is (with 8G) and then we can swap things around after that so I can see what the distro job does with more | |
| 20:22:57 | gouthamr | ack dansmith | |
| 20:23:26 | dansmith | gouthamr: yes of course.. I'm certainly not arguing to keep it in such that we need a vote or anything, so if you're actively hoping to drop that support from devstack or something I certainly won't argue against it | |
| 20:23:35 | dansmith | just want one job that works, is all :) | |
| 20:23:43 | gouthamr | ++ | |
| 20:41:38 | dansmith | wow, no oom on the nova job, but just constant fail until it timed out | |
| 20:43:55 | dansmith | and like no errors in n-cpu log | |
| 20:47:03 | dansmith | seems like all cinder fails | |
| 22:01:34 | dansmith | nova job failed again | |
| 22:34:02 | dansmith | gouthamr: here's the report from the nova job: https://2db12bf686954cafec50-571baf10d8fb9f8a9cb5a5e38315002b.ssl.cf5.rackcdn.com/881585/4/check/nova-ceph-multistore/4694eec/testr_results.html | |
| 22:34:07 | dansmith | not nearly as much fail as before | |
| 22:34:15 | dansmith | the second job is a failure waiting for cinder to detach | |
| 22:34:36 | dansmith | the first one is a nova detach, I need to check that test to see if it's doing an ssh wait, which might help | |
| 22:35:11 | dansmith | oh actually that first one is in the cinder tempest plugin, not one of ours | |
| 22:36:56 | dansmith | looks like it's probably not | |
| 22:41:37 | dansmith | gmann: around? | |
| 22:43:08 | gmann | dansmith: hi | |
| 22:43:40 | dansmith | gmann: looks like the scenario.manager create_server in tempest is not waiting for sshableness | |
| 22:43:57 | dansmith | gmann: and the cinder tempest plugin has a test that uses that and may be part of the volume detach issue(s) | |
| 22:44:11 | dansmith | gmann: is it possible to make that create_server() always wait? | |
| 22:44:59 | gmann | this one right? https://github.com/openstack/cinder-tempest-plugin/blob/d6989d3c1a31f2bebf94f2a7a8dac1d9eb788b1b/cinder_tempest_plugin/scenario/test_volume_encrypted.py#L86 | |
| 22:45:24 | dansmith | gmann: yeah, but I'm wondering if create_server() itself can wait so anything that uses it will wait | |
| 22:45:45 | gmann | dansmith: I think it make sense to wait for SSH by default in scenario tests | |
| 22:45:48 | dansmith | gmann: the current thinking is that a lot of our volume detach problems come from trying to attach a volume before an instance is booted | |
| 22:45:58 | dansmith | gmann: ack, okay | |
| 22:46:33 | gmann | yeah, and with nature of most of scenario tests, booting server fully is always good before trying things on that | |
| 22:47:20 | dansmith | yeah, okay lemme try to do that | |
| 22:47:52 | gmann | dansmith: ok. we need to create/pass the validation resources from scenario manager which is needed for SSH/ping | |
| 22:48:46 | gmann | or may be via get_remote_client | |
| 22:49:40 | dansmith | yeah | |
| 23:48:35 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-27 | |||
| 01:29:53 | dansmith | ugh, failed 102 on that last one | |
| 01:36:19 | dansmith | er, no, I suck at numbers.. 9, and all scenario, so must be related to my tempest change | |
| 03:59:16 | opendevreview | yatin proposed openstack/nova master: [DNM] Check detatch issue with lower tb size https://review.opendev.org/c/openstack/nova/+/881690 | |
| 08:17:21 | sean-k-mooney | gibi: bauzas: fun issue https://bugs.launchpad.net/os-vif/+bug/2017868 | |
| 08:17:41 | bauzas | ack | |
| 08:20:11 | opendevreview | Christophe Fontaine proposed openstack/os-vif master: OVS DPDK tx-steering mode support https://review.opendev.org/c/openstack/os-vif/+/881644 | |
| 08:20:49 | sean-k-mooney | im still considering -2ing ^ | |
| 10:02:46 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement is_luks_inside_qcow2 funtion https://review.opendev.org/c/openstack/nova/+/854030 | |
| 10:03:18 | opendevreview | Sofia Enriquez proposed openstack/nova master: Implement encryption on backingStore https://review.opendev.org/c/openstack/nova/+/870012 | |
| 14:30:55 | opendevreview | ribaudr proposed openstack/nova master: Fix live migrating to a host with cpu_shared_set configured will now update the VM's configuration accordingly. https://review.opendev.org/c/openstack/nova/+/877773 | |
| 15:57:12 | gouthamr | o/ dansmith - i found a "rbd image is busy" error in a different change parented on your devstack-plugin-change: https://zuul.opendev.org/t/openstack/build/278d411484764c55829c26cec2860bee/log/controller/logs/screen-n-cpu.txt#10742 | |
| 15:58:05 | gouthamr | im still looking, but, the cephfs job runs some scenario tests where it boots up nova VMs, attaches manila shares to them... these tests are failing to delete the VMs after | |
| 15:59:05 | gouthamr | the DELETE /servers/{id} call returns with 204, we wait in a loop, the server doesn't go away within 300 seconds.. and n-cpu logs don't have an error.. | |
| 16:01:29 | gouthamr | i don't know if the "rbd.ImageBusy" exception has anything to do with this silent failure | |
| 16:02:46 | gouthamr | it might | |
| 16:08:53 | dansmith | gouthamr: yeah I think eharney mentioned that it's been a pervasive problem, but I usually see it like once or twice and not a huge raft of them in one job when there's not something else going wrong | |
| 16:09:50 | gouthamr | i think in this job it's been consistently failing.. | |
| 16:10:16 | gouthamr | maybe i am missing a config opt here that you've set elsewhere | |
| 17:20:05 | opendevreview | sean mooney proposed openstack/os-vif master: [WIP] set default qos policy https://review.opendev.org/c/openstack/os-vif/+/881751 | |
| 22:52:45 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-04-28 | |||
| 07:47:09 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 12:09:24 | opendevreview | Stephen Finucane proposed openstack/nova master: libvirt: Remove unnecessary arg https://review.opendev.org/c/openstack/nova/+/881817 | |
| 12:14:33 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 12:30:01 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 12:30:28 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 12:34:58 | opendevreview | Artom Lifshitz proposed openstack/nova master: Fix pep8 errors with new hacking https://review.opendev.org/c/openstack/nova/+/874517 | |
| 15:59:30 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Preserve cached base images for failed resize ops https://review.opendev.org/c/openstack/nova/+/877410 | |
| 18:26:48 | opendevreview | Ghanshyam proposed openstack/nova-specs master: Re-propose "Policy service role spec" https://review.opendev.org/c/openstack/nova-specs/+/881880 | |
| 19:57:46 | dansmith | gouthamr: I think I'm getting close.. only one failure on the nova patch this last time, and it's yet another volume test that doesn't wait for sshable | |
| 19:57:54 | dansmith | about out of steam for the week, but maybe we'll have some good news on monday | |
| 19:59:17 | gouthamr | dansmith: o/ that's awesome... that's gone down from double digits, so very encouraging to know! | |
| 19:59:33 | dansmith | also, | |
| 19:59:45 | dansmith | we've been chasing a bunch of volume attach failures unrelated to ceph for a while | |
| 19:59:58 | dansmith | so these wait-for-sshable fixes will hopefully improve the stability of the other jobs too | |
| 20:01:21 | gouthamr | ^^ ++ ; i was following the pattern; i think we have at least one "wait for ssh" workaround in some manila scenario tests that can be removed because you're fixing up the base class we're using | |
| 20:01:59 | dansmith | gouthamr: well, we had to abandon the "make them all do it all the time" approach for now at least for a variety of reasons, | |
| 20:02:07 | dansmith | so don't remove any of that :) | |
| 20:02:56 | dansmith | I resorted to just pushing a patch to cinder-tempest-plugin to fix the tests there, and the base patch just makes it possible to wait_until=SSHABLE for the scenario tests | |
| 20:04:41 | gouthamr | i see | |
| 20:05:03 | dansmith | not my preference, but gotta make incremental progress :/ | |
| 20:05:16 | gouthamr | ack; i wanted to see if it was the same issue first too, so i set ENABLE_CEPH_NOVA=False, and the cephfs job was passing; now i'll pop that off and see if your fixes are good with those too | |
| 20:05:35 | dansmith | so the stack is currently nova->cinder-tempest->tempest->devstack-ceph-plugin ... it's thorny | |
| 20:06:17 | gouthamr | yes, and i extended the chain further :D with -->manila-tempest-plugin --> devstack-plugin-ceph | |
| 20:06:30 | dansmith | going for a world record? :) | |
| 20:07:00 | dansmith | life was so simple when everything fit on a single floppy .. | |
| 20:07:05 | gouthamr | lol :D if someone wasn't looking, they'd think we're building a new fancy feature | |
| 20:07:09 | dansmith | hehe | |
| 20:42:28 | opendevreview | Merged openstack/nova master: Fix a typo in this URL: https://docs.openstack.org/nova/latest/admin/availability-zones.html https://review.opendev.org/c/openstack/nova/+/878797 | |
| #openstack-nova - 2023-05-01 | |||
| 12:54:28 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:01:10 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:04:28 | opendevreview | sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912 | |
| 13:31:56 | dansmith | sean-k-mooney: so check this out: https://264a86dd518c5142b8a6-3b04393c7506ab488b4fde073fc22e36.ssl.cf1.rackcdn.com/881764/1/check/cinder-tempest-plugin-lvm-multiattach/d28afe7/testr_results.html | |
| 13:32:01 | dansmith | maybe kashyap too ^ | |
| 13:32:25 | dansmith | that almost kinda looks like we attached the disk to the guest and the kernel crashed while probing partitions/filesystems, right? | |
| 13:33:10 | dansmith | in sysfs_add_file_mode_ns() | |
| 13:33:49 | dansmith | or perhaps while creating sysfs entries for the disk perhaps? | |
| 16:30:56 | dansmith | eharney: you around by chance? | |
| 16:45:20 | eharney | dansmith: yes | |
| 16:45:43 | dansmith | eharney: so, I seem to have gotten the nova ceph job down to a single repeatable failure | |
| 16:45:56 | dansmith | it is in the volume extend test, and it fails during cleanup | |
| 16:46:25 | dansmith | it's trying to, I guess, detach the volume from the server before deleting it and then before deleting the volume | |
| 16:46:49 | dansmith | it does *not* fail locally, so I don't think it's something fundamentally broken with new ceph or anything like that | |
| 16:46:54 | dansmith | https://955f32f8268e5d475e65-6c8f4c6e546a0854b4c11cc7c78829ca.ssl.cf5.rackcdn.com/881585/6/check/nova-ceph-multistore/1ecb09a/testr_results.html | |
| 16:47:25 | eharney | dansmith: let me take a look through the logs | |
| 16:47:49 | eharney | detach failing isn't something i'm familiar with (other than it being mentioned here the other day) | |
| 16:48:04 | dansmith | I've just been tracing through the code looking for how this works and it seems to me like everything is working, but perhaps it's just legitimately that the guest doesn't let go of the volume when we detach after the resize happened | |
| 16:48:17 | dansmith | okay, volume detach is without a doubt our most common failure | |
| 16:48:34 | dansmith | eharney: yeah appreciate if you could see if you spot anything in the logs | |
| 17:00:13 | dansmith | I see the guest saw the size change on vdb and also mentions that it's resizing the filesystem on it... that looks like more than just the block device size change, | |