Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-18
18:17:56 artom sean-k-mooney, we could... I just don't see the point? It's already all info the guest has access to
18:18:00 artom With lspci, `ip`, etc
18:18:17 sean-k-mooney the point woudl jsut to make that part of the metadta nolonger optional
18:18:36 sean-k-mooney so you woudl not have to check if its aviabel or not in the ugest it will be
18:18:43 sean-k-mooney even if you could get that info elsewhere
18:24:07 opendevreview Balazs Gibizer proposed openstack/nova master: Clean up after ImportModulePoisonFixture https://review.opendev.org/c/openstack/nova/+/870993
18:24:51 gibi bauzas: while it is part of the the functional instability it is not the case and therefore this is not the fix for it, but it is related and it is a cleanup ^^
18:25:10 gibi s/it is not the case/it is not the root cause/
18:27:03 sean-k-mooney hum interesting
18:27:07 sean-k-mooney what were we leaking
18:27:18 sean-k-mooney ah the filter
18:28:04 sean-k-mooney ok so we were leakign the filters which could increae memory usage
18:28:17 sean-k-mooney but it is not sharing state or caussing ither issues
18:28:30 gibi yeah
18:28:33 sean-k-mooney it does proably contribute to the oom issues
18:28:51 gibi it is part of the functional instability realted to https://bugs.launchpad.net/nova/+bug/1946339
18:29:14 gibi as in the recent case it is that import poison that get called by the late eventlet
18:33:33 gibi so I looked at the poision and found a global state
18:34:05 sean-k-mooney ya so each executor has what about 1000 tests per worker
18:34:52 sean-k-mooney i would guess this is addign 1-2MB per run
18:34:59 sean-k-mooney at most
18:35:44 gibi yeah it is not realted the OOM case we saw in tempest
18:35:54 sean-k-mooney i have not check but i would be surpried if adding a filter allocates more then a KB or memory it should jsut be a few fucntion pointer and python objects
18:36:22 sean-k-mooney well didnt we also have OOM issues in teh funtional tests
18:36:47 sean-k-mooney wll not OOM python interpreter crahses
18:38:22 sean-k-mooney gibi: one question however
18:39:25 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/870993/1/nova/tests/fixtures/nova.py#1838
18:39:41 sean-k-mooney shoudl we install the filter in setUp not init
18:40:02 sean-k-mooney i assume the orgianl intent of doign it in init was to do it only once
18:40:13 sean-k-mooney for the life time of the fixture and resue the fixture
18:40:21 sean-k-mooney which is not how we are doign it
18:40:54 sean-k-mooney so if we are going to clean up in cleanup should we set it up in setUp
21:01:04 opendevreview Merged openstack/nova master: Use new get_rpc_client API from oslo.messaging https://review.opendev.org/c/openstack/nova/+/869900
21:01:25 opendevreview Merged openstack/nova master: FUP for the scheduler part of PCI in placement https://review.opendev.org/c/openstack/nova/+/862876
#openstack-nova - 2023-01-19
00:41:13 opendevreview Ghanshyam proposed openstack/nova master: Bump openstack-placement version in functional tox env https://review.opendev.org/c/openstack/nova/+/871011
01:27:03 gmann gibi: bauzas: placement is released and bumping the version in tox.ini https://review.opendev.org/c/openstack/nova/+/871011
08:33:25 gibi gmann: thanks
08:37:20 bauzas gmann: ack, looking
09:19:26 opendevreview Balazs Gibizer proposed openstack/nova master: Clean up after ImportModulePoisonFixture https://review.opendev.org/c/openstack/nova/+/870993
09:19:51 gibi sean-k-mooney: good point about __init__ vs setUp() in the fixture, I fixed it up ^^
09:21:07 gibi sean-k-mooney: we have python test executor crashes in functional but we don't have syslog saved in functional so I cannot confirm that it is due to OOM. I see python crashes locally with functional run couple of weeks back and that wasnt OOM but segfault from the interpreter
09:21:35 zigo https://salsa.debian.org/openstack-team/services/ceilometer-instance-poller/-/blob/debian/zed/ceilometer_instance_poller/instance_poller.py#L133
09:21:35 zigo which works well, but crashes here:
09:21:35 zigo https://salsa.debian.org/openstack-team/services/ceilometer-instance-poller/
09:21:35 zigo We wrote this:
09:21:35 zigo Hi there!
09:21:37 zigo if we're using the Ceph backend.
09:21:39 zigo Does anyone of you know how to set the libvirt Ceph secret so then libvirt/libguestfs knows how to use the Ceph backend?
09:37:37 bauzas zigo: sorry, not a expert at all on how we integrate with ceph
09:37:54 bauzas maybe look at how nova attaches rdb disks
09:37:59 bauzas rbd*
09:38:08 zigo Yeah, it's smoething like that! :)
09:38:30 zigo Got to find out how to tell libvirt/libguestfs what username and secret to use.
09:38:50 zigo Obviously, I've been search for a long time before asking here ...
09:40:09 bauzas gibi: weirdo
09:40:46 bauzas gibi: I find we only take 30secs for grabbing the image which sizes 85
09:40:48 bauzas 2023-01-18 16:24:58.358 127779 INFO tempest.api.compute.admin.test_aaa_volume [-] DNM bauzas: Time elapsed: 31.457598802999883 2023-01-18 16:24:58.379 127779 INFO tempest.api.compute.admin.test_aaa_volume [-] DNM bauzas: Size of custom img: 85
09:41:44 bauzas and in this run, it took 93 secs
09:41:56 bauzas have we changed the size of the default image ?
09:42:29 bauzas mmm, yeah
09:42:30 bauzas 2023-01-18 16:22:09.379365 | controller | ++ lib/tempest:configure_tempest:434 : iniset /opt/stack/tempest/etc/tempest.conf compute image_ref 7bac73d8-9596-43ae-900f-546a3d53de36 2023-01-18 16:22:09.411839 | controller | ++ lib/tempest:configure_tempest:435 : iniset /opt/stack/tempest/etc/tempest.conf compute image_ref_alt 7bac73d8-9596-43ae-900f-546a3d53de36
09:44:51 bauzas oh, nevermind
09:45:00 bauzas we return the image_id hence the very small size
10:11:39 opendevreview Merged openstack/nova master: Bump openstack-placement version in functional tox env https://review.opendev.org/c/openstack/nova/+/871011
10:19:33 gibi I cannot reproduce the functional failure locally with the same way I did in the past https://bugs.launchpad.net/nova/+bug/1946339 so I'm stuck a bit there
10:26:40 bauzas gibi: I'll try to run it
10:27:14 bauzas gibi: have you tried to run the functional tests by using the subunit ?
10:27:57 bauzas gibi: using $ stestr run --load-list FILENAME
10:28:58 bauzas gibi: lemme try to do it
10:30:23 bauzas but after this, I'll need to go off until end of the day as I need to help my daughters due to the french strike
10:31:56 gibi bauzas: yepp
10:32:30 gibi bauzas: I tried the list of test from the failed worker from the recent gate failure and also tried the list from my past reproducers in that bug report
10:32:46 gibi neither triggered the problem after couple of hours of continuous run
10:33:02 gibi now simply run the test with --until-failure and --random
10:33:06 gibi but this is a long shot
10:33:44 bauzas gibi: ok, I'll try to run it the same
10:34:01 bauzas gibi: I guess you know about https://stestr.readthedocs.io/en/latest/MANUAL.html#automated-test-isolation-bisection ?
10:34:41 gibi i think that only works after we can reproduce the fauilure
10:35:03 gibi and it needs a stable reproduction not a stohastic one
10:35:35 bauzas gibi: I'll try to look at that
10:35:56 bauzas folks, as I also said downstream, I'll be off until 1630UTC
10:39:33 bauzas gibi: so I loaded the subunit
10:39:36 bauzas into stestr
10:39:49 bauzas (functional-py38) [sbauza@sbauza nova]$ stestr load < /tmp/zuul-logs.bNmdDA/testrepository.subunit
10:40:06 bauzas I should now be able to run the failing test
10:41:34 bauzas anyway, I need to disappear :(
10:55:22 gokhanis hello folks, after rebooting my compute host, I can't attach my cinder volumes to instances. Nova throws "unable to lock /var/lib/nova/mnt/dgf/volume-xx for metadata change: No locks available" Full logs are in https://paste.openstack.org/show/beicZ71J17WeNwLjghKc/ What can be reason of this problem and how can I solve this issue ?
11:06:33 sean-k-mooney gokhanis: what kind fo cinder backedn are you using
11:06:57 sean-k-mooney is this nfs based?
11:08:16 gokhanis sean-k-mooney, yes I am using cinder nfs driver.
11:08:22 sean-k-mooney this does not sound like it actuly a nova issue but rather and issue with how you have deployed it. ie. like you are hitting a ulimit or similar in the env
11:08:34 sean-k-mooney gokhanis: ok is it nfs3 or nfs4
11:08:54 sean-k-mooney nfs3 has locking issues and we dont really support it anymore
11:09:19 sean-k-mooney ideally you should use nfs 4.2 or newer
11:10:42 sean-k-mooney gokhanis: i think your hitiing this https://bugzilla.redhat.com/show_bug.cgi?id=1547095
11:11:47 sean-k-mooney gokhanis: apprently you can "fix" it by using the nolock mount option for nfsv3
11:12:07 sean-k-mooney gokhanis: also relevent https://access.redhat.com/solutions/2780381
11:12:50 gokhanis sean-k-mooney, I deployed my env with Openstack Ansible. I have create nfs share on zfs storage and I used it as cinder nfs backend.
11:12:56 sean-k-mooney gokhanis: form a downstream perspetive redhat discontinued support for nfsv3 in our porudct because of these locking issues
11:14:21 sean-k-mooney gokhanis: yep so on the zfs culster you need to either ensure that the zfs pool is exported as nfsv4, nfsv3 with the nfsv4 lock manager or you have to ensure the nfs share is mounted with the nolock option on the compute hosts

Earlier   Later