Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-18
16:38:26 bauzas anyway, this is a guess
16:38:38 bauzas nothing was changed in this module since Feb 22
16:38:39 dansmith if anything, it makes me think we can make it larger as this might have been the OOM limit I was running into with 2G
16:38:40 sean-k-mooney dansmith: i agree
16:39:03 bauzas https://github.com/openstack/tempest/blob/master/tempest/api/compute/admin/test_volume.py
16:39:20 sean-k-mooney dansmith: the trade of is we only have 80G of disk space in ci
16:39:33 sean-k-mooney so back to 2G perhaps 20G proably not
16:39:34 bauzas oh wait
16:39:36 dansmith I understand, disk space is not the issue though
16:39:40 bauzas maybe we change the image ref
16:40:19 opendevreview Kashyap Chamarthy proposed openstack/nova master: libvirt: At start-up allow skiping compareCPU() with a workaround https://review.opendev.org/c/openstack/nova/+/870794
16:41:00 kashyap Duh, forgot to commit 2 files
16:41:20 opendevreview Kashyap Chamarthy proposed openstack/nova master: libvirt: At start-up allow skiping compareCPU() with a workaround https://review.opendev.org/c/openstack/nova/+/870794
16:45:41 bauzas https://github.com/openstack/devstack/blob/master/lib/tempest#L213-L220
16:45:43 bauzas hmmmm
16:46:14 gibi propsed the skip for this test https://review.opendev.org/c/openstack/tempest/+/870974 I checked no other test using the _create_image_with_custom_property util function
16:46:29 bauzas 2023-01-18 11:54:33.763321 | controller | ++ lib/tempest:get_active_images:155 : '[' cirros-raw = cirros-0.5.2-x86_64-disk ']'
16:46:37 bauzas https://storage.gra.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_362/870924/2/check/nova-ceph-multistore/3626391/job-output.txt
16:46:46 bauzas we only get the cirros image
16:46:53 bauzas shouldn't be that large
16:48:22 dansmith bauzas: you understand that the ceph job uses a 1G cirros image right?
16:49:37 bauzas Jan 18 11:57:14.947333 np0032776548 glance-api[110229]: DEBUG glance.image_cache [None req-76dfbfc9-8d31-4a47-a529-e95d8077cfc0 tempest-AttachSCSIVolumeTestJSON-1393201534 tempest-AttachSCSIVolumeTestJSON-1393201534-project-admin] Tee'ing image '0bc12eec-2802-48e8-bedf-0931be582d19' into cache {{(pid=110229) get_caching_iter /opt/stack/glance/glance/image_cache/__init__.py:343}}
16:49:49 bauzas dansmith: oh sorry no, wasn't knowing
16:49:58 gibi 950MB image is downloaded in my case
16:50:10 dansmith bauzas: [08:31:27] <dansmith> bauzas: I recently increased the size of the image used on the ceph job from 16MB to 1G
16:50:17 bauzas missed that line
16:50:23 dansmith :)
16:50:53 bauzas ok, so we know that we cache 1GB in memory
16:50:57 gibi te test is nice as it get the image metadata first so in the log there is a "size": 996147200 before the data is downloaded
16:51:24 bauzas gibi: I'm waiting for my test job to return but I guess we'll see a size of 1GB in memory for that variable
16:52:29 bauzas dansmith: about your question (why do we trigger now the kill and not earlier), my guess is that we were just below the line
16:52:34 opendevreview Balazs Gibizer proposed openstack/nova master: DNM: Test that OOM triggering test is skipped https://review.opendev.org/c/openstack/nova/+/870950
16:53:34 dansmith bauzas: yeah, like I said, we're probably just swapping it all and never touching it again, so pressure is high and we're close to the edge :)
16:53:45 bauzas one way to alleviate this issue would be to make sure we run that greedy test into a specific test runner worker
16:54:06 opendevreview Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324
16:54:10 bauzas https://stestr.readthedocs.io/en/latest/MANUAL.html#test-scheduling
16:54:29 gibi bauzas: we just shouldn't load the whole image data in memory at once
16:54:29 bauzas tempest exposes the worker configs from stestr
16:54:57 gibi as in general image size can be way bigger than memory size
16:55:10 bauzas gibi: that's true
16:55:20 bauzas caching the metadata seems ok to me
16:55:34 bauzas caching the data itself seems unnecessary
16:55:44 bauzas unless you want to compare bytes per bytes
16:55:46 gibi yeah meatadata is bound by glance API
16:55:59 gibi image size is unbound
16:56:05 bauzas but agreed you could and should compare streams and not objects
16:56:15 bauzas for the dataz
16:56:28 dansmith it's just a naive test not thinking about the world outside a 16MB image
16:56:30 bauzas glance team is added on the bug report
16:56:38 dansmith anyone using this for verification of a real cloud likely has an even larger image,
16:56:41 dansmith so it's clearly not okay to do this
16:56:45 bauzas agreed
16:56:55 bauzas whoami-rajat: hey, happy new year :)
16:57:03 opendevreview Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324
16:58:53 opendevreview Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324
16:59:49 opendevreview Merged openstack/nova master: Strictly follow placement allocation during PCI claim https://review.opendev.org/c/openstack/nova/+/855650
17:01:08 bauzas hmmm, stackalytics.io is fone
17:01:10 bauzas gone*
17:01:57 gibi what do you mean by gone?
17:02:00 gibi it loads for me with
17:02:01 gibi Last updated on 18 Jan 2023 12:44:21 UTC
17:02:06 bauzas I got a timeout
17:02:16 bauzas oh this works now
17:04:20 gibi gmann, dansmith: was there some policy change recently that can cause that the nova functional test locally gets 403 from placement?
17:04:33 gibi {'errors': [{'status': 403, 'title': 'Forbidden', 'detail': 'Access was denied to this resource.\n\n placement:allocations:list ', 'request_id': 'req-c5c03029-ba79-4e4a-8ec8-03deadb24ded'}]}
17:04:43 dansmith gibi: reliably?
17:04:46 gibi yes
17:04:51 dansmith oof, then probably
17:04:51 gibi this is pure nova master functional test
17:04:53 gmann gibi: yeah we changed the default and placement fixture needed fix whihc is merged i think
17:05:00 bauzas gibi: yeah
17:05:09 gmann gibi: can you rebase the placement repo?
17:05:12 bauzas gibi: we switched to new RBAC policies
17:05:28 gibi it is just my nova repo locally
17:05:32 gmann this one https://review.opendev.org/c/openstack/placement/+/869525
17:05:58 gmann gibi: because nova functional test use placement fixture from placement repo
17:06:17 bauzas which we pull as a dependency, right?
17:06:31 gmann yes
17:06:57 gibi openstack-placement>=1.0.0
17:06:57 gibi so nova's tox.ini has
17:07:06 gibi that should pull in the latest openstack-placement
17:07:12 gibi in the functional venv
17:07:35 bauzas tox -r ?
17:07:42 gibi I have openstack-placement==8.0.0
17:07:44 gibi in the venv
17:07:56 gibi I guess we merged the fix in placement but we haven't released it
17:08:04 gmann i do not think we released placement with that
17:08:17 gmann yeah not released yet, we should do
17:08:17 gibi so nova's tox.ini pulls placement from pypi
17:08:25 gibi ^^ yepp
17:08:43 gibi or change nova's tox.ini to pull placement from github
17:08:55 gmann yeah for now this can be workaround
17:08:57 gmann let me push it today unless bauzas you want to do?
17:10:15 gmann I feel placement fixture import from placement in functional test should be changed like we do for cinder/glance fixture otherwise we need new placement release for any change in there
17:11:11 bauzas gmann: do the push and I'll +1
17:11:30 gmann bauzas: ok
17:12:22 gibi gmann: based on the constraint in tox.ini openstack-placement>=1.0.0 this is the first time we need such a release due to the fixture
17:13:37 bauzas gibi: gmann: shall we consider to pull from gh ?
17:13:57 gibi I would keep pypi
17:14:07 gmann gibi: yeah because this actually change the things like default policy. but this can occur if change in default for policy or config unless we change nova functional test to move to those new defaults. for example

Earlier   Later