Earlier  
Posted Nick Remark
#openstack-nova - 2020-07-24
15:14:57 sean-k-mooney so it should be 8 https://github.com/openstack/devstack-plugin-ceph/blob/master/devstack/settings#L17
15:15:01 sean-k-mooney by default
15:15:12 dansmith it's overridden in our job somewhere
15:15:16 dansmith you can see in the devstacklog
15:15:24 sean-k-mooney CEPH_LOOPBACK_DISK_SIZE is
15:15:31 sean-k-mooney is VOLUME_BACKING_FILE_SIZE
15:15:34 dansmith ues
15:15:36 dansmith both are
15:15:47 sean-k-mooney ok cool
15:16:12 dansmith VOLUME_BACKING_FILE_SIZE=24G
15:16:18 dansmith and the df shows 24G on /var/lib/ceph
15:17:48 sean-k-mooney ah yes it does
15:17:55 dansmith we run "ceph df" to get the DISK_GB we report,
15:17:56 dansmith and don't really do much to it,
15:18:08 dansmith so it really seems like we're being told 10G
15:19:52 dansmith lyarwood: do you know anything about what ceph df may be telling us about total pool size that differs from the backing store's size?
15:20:34 lyarwood dansmith: nope, AFAIK it just reports the size of the images_rbd_pool
15:21:16 dansmith seems straightforward :)
15:21:38 lyarwood https://github.com/openstack/nova/blob/master/nova/virt/libvirt/storage/rbd_utils.py#L374-L382 - ah well melwitt has a handy comment here that might help
15:21:50 lyarwood gah highlights are a new lines off but you get the point
15:21:58 dansmith oh I read that,
15:22:01 dansmith but didn't grok until now
15:22:10 dansmith so replication makes the thing looks smaller I guess?
15:22:37 dansmith seems weird to go from 24G to 10G, as that's not an even factor
15:23:15 dansmith er, no wait
15:23:24 dansmith that's for max_avail, which is "free" not total right?
15:24:30 lyarwood right sorry and you're seeing 10 reported as the total capacity right?
15:24:34 dansmith correct
15:24:40 lyarwood kk sorry then that isn't it
15:26:43 dansmith I guess one thing we could do is increase the ceph backing size to 36G and see if DISK_GB goes up
15:26:45 bauzas can someone tell me what the fuck is ? http://paste.openstack.org/show/796292/
15:26:57 bauzas tl;dr: ssh: connect to host review.openstack.org port 29418: Network is unreachable
15:27:11 bauzas have I missed a memo ?
15:27:25 melwitt there's a openstackstatus above ^ said there will be a short outage
15:29:00 openstackgerrit Sylvain Bauza proposed openstack/nova-specs master: WIP: Offline Reshape tool spec https://review.opendev.org/742908
15:29:05 bauzas yay, it worked
15:29:09 bauzas melwitt: thanks
15:29:46 bauzas calling it a day
15:32:15 melwitt dansmith: MAX_AVAIL should be total actually, just taking number of replicas into account. if you only have 1 replica (default NUM_REPLICAS=1) then MAX_AVAIL should match whatever total says in 'ceph df'
15:32:49 dansmith melwitt: you're reporting free as max_avail though in that thing aren't you?
15:32:59 dansmith or does MAX_AVAIL != max_avail ?
15:32:59 melwitt but if you've set NUM_REPLICAS=2 when you deployed a devstack, then since the devstack ceph plugin creates 2 OSDs on the same HDD in that case, it would be 2x the real disk
15:33:07 melwitt no MAX_AVAIL is a ceph thing
15:33:30 melwitt (if you're referring to what is written about ceph df in rbd_utils.py)
15:33:32 dansmith you mean half the disk I assume
15:33:34 dansmith yeah
15:33:57 melwitt no like the old behavior used to report 20G if you had a 10G disk, of you had created 2 OSDs that point at the same HDD
15:34:02 dansmith so maybe (24 - overhead) / 2 == 10 or something
15:34:22 melwitt you're using NUM_REPLICAS=1 right? you didn't set it in the job
15:34:31 melwitt if so, there shouldn't be a difference
15:34:38 dansmith I'm not setting it, but let me look if it's getting set
15:35:06 melwitt I doubt it, I've never seen it set in CI before. I had to set it locally to do the testing for that MAX_AVAIL change
15:35:26 dansmith yeah I don't even see that variable anywhere
15:35:41 dansmith is that a devstack-plugin-ceph thing?
15:35:45 melwitt yeah sec
15:36:02 lyarwood dansmith: https://docs.ceph.com/docs/jewel/rados/operations/pools/#create-a-pool ; sudo ceph -c /etc/ceph/ceph.conf osd pool create vms 8 8 ; that doesn't mean create a 8GB pool
15:36:09 melwitt bah sorry it's CEPH_REPLICAS https://github.com/openstack/devstack-plugin-ceph/blob/master/devstack/lib/ceph#L109
15:36:15 dansmith lyarwood: yeah we established that :)
15:36:22 lyarwood ah sorry wasn't watching irc
15:36:30 dansmith lyarwood: somewhere in the plugin I saw a comment that made it sound like that was size
15:37:16 melwitt 10G honestly I would have thought is just the cloud image's disk size, no?
15:37:20 dansmith melwitt: yeah 1
15:37:23 melwitt or do we probably use something larger in CI
15:37:26 dansmith melwitt: no, said above, it's 24G
15:37:37 dansmith melwitt: https://zuul.opendev.org/t/openstack/build/13d8a055ff1b4be0b627205f4d51d50f/log/controller/logs/df.txt
15:37:46 dansmith and it's overridden to 24G in the devstack log
15:37:50 lyarwood that's total for the three different pools
15:37:59 lyarwood vms images and volumes?
15:38:04 melwitt oh I see
15:38:14 dansmith lyarwood: and are the pools set to something specific for size? that's what we're trying to find and can't :)
15:38:38 dansmith lyarwood: the way it looks now I'd assume it just reports that they're all 24G in size, with various amounts free like zfs does for filesystems on a pool,
15:38:42 dansmith but I'm just guessing
15:38:52 dansmith I'm stacking a ceph devstack so I can poke but right now all I have is logs
15:39:28 dansmith if total decreases as we use space, then we're not really reporting the right thing to placement
15:39:32 dansmith which could be part of the problem of coruse
15:40:02 lyarwood dansmith: yup true and that's also going to bounce around alot during a tempest run
15:40:07 dansmith yep
15:40:34 dansmith I'm pretty sure this is not a consequence of my job, by the way, I think mine is just a little slower because we have some glance features turned on, so we probably have a little more of a logjam than normal
15:41:06 dansmith oh jeez, you know what I just realized?
15:41:25 dansmith we might be snapshotting to the file store and not the ceph store in some cases, actually
15:41:27 dansmith hmm
15:41:55 dansmith nova does the snapshots itself so maybe not, but if we ever do a raw image upload.. the default store is the file store
15:42:04 dansmith not that that would cause this, but it might be changing the timing characteristics
15:42:32 dansmith I'll have to think on that a bit
15:42:33 melwitt well, this doesn't look promising for MAX_AVAIL, it sounds like it would decrease with use and is not a total https://access.redhat.com/solutions/3537961
15:43:08 dansmith ah yeah
15:43:41 dansmith melwitt: did you read this? https://access.redhat.com/solutions/2273951
15:44:03 dansmith we're not replicated I guess so maybe that doesn't affect us in CI, but probably has some impact for real users of this
15:44:12 melwitt no
15:45:47 melwitt so there are multiple reasons MAX_AVAIL shouldn't be used :(
15:46:31 dansmith not it!
15:47:42 dansmith the other problem I'm guessing,
15:47:44 melwitt yeah... I'm thinking whether to revert that or tweak it to take total and divide by pool size, the latter would do what was actually desired and report total with replication considered
15:48:11 dansmith is that if we report the real actual total (even minus replication overhead), but other pools can consume space from the same store,
15:48:17 dansmith we will tell placement we have more room than it can allocate
15:48:50 dansmith so really we need to sum up all the pools on the same store, and then set reserved= for any space they use I guess, but then we race with those other uses in our reporting
15:48:55 dansmith and could go negative
15:49:19 melwitt yeah, I'm trying to remember, I could have sworn this get_pool_info was only used to report free space, not total space, but I could be totally making that up
15:49:45 melwitt or that that's what it's used for ultimately in higher layers
15:50:23 melwitt let me look up what "total" used to be, maybe it meant "total available"

Earlier   Later