Earlier  
Posted Nick Remark
#openstack-nova - 2020-07-24
15:36:22 lyarwood ah sorry wasn't watching irc
15:36:30 dansmith lyarwood: somewhere in the plugin I saw a comment that made it sound like that was size
15:37:16 melwitt 10G honestly I would have thought is just the cloud image's disk size, no?
15:37:20 dansmith melwitt: yeah 1
15:37:23 melwitt or do we probably use something larger in CI
15:37:26 dansmith melwitt: no, said above, it's 24G
15:37:37 dansmith melwitt: https://zuul.opendev.org/t/openstack/build/13d8a055ff1b4be0b627205f4d51d50f/log/controller/logs/df.txt
15:37:46 dansmith and it's overridden to 24G in the devstack log
15:37:50 lyarwood that's total for the three different pools
15:37:59 lyarwood vms images and volumes?
15:38:04 melwitt oh I see
15:38:14 dansmith lyarwood: and are the pools set to something specific for size? that's what we're trying to find and can't :)
15:38:38 dansmith lyarwood: the way it looks now I'd assume it just reports that they're all 24G in size, with various amounts free like zfs does for filesystems on a pool,
15:38:42 dansmith but I'm just guessing
15:38:52 dansmith I'm stacking a ceph devstack so I can poke but right now all I have is logs
15:39:28 dansmith if total decreases as we use space, then we're not really reporting the right thing to placement
15:39:32 dansmith which could be part of the problem of coruse
15:40:02 lyarwood dansmith: yup true and that's also going to bounce around alot during a tempest run
15:40:07 dansmith yep
15:40:34 dansmith I'm pretty sure this is not a consequence of my job, by the way, I think mine is just a little slower because we have some glance features turned on, so we probably have a little more of a logjam than normal
15:41:06 dansmith oh jeez, you know what I just realized?
15:41:25 dansmith we might be snapshotting to the file store and not the ceph store in some cases, actually
15:41:27 dansmith hmm
15:41:55 dansmith nova does the snapshots itself so maybe not, but if we ever do a raw image upload.. the default store is the file store
15:42:04 dansmith not that that would cause this, but it might be changing the timing characteristics
15:42:32 dansmith I'll have to think on that a bit
15:42:33 melwitt well, this doesn't look promising for MAX_AVAIL, it sounds like it would decrease with use and is not a total https://access.redhat.com/solutions/3537961
15:43:08 dansmith ah yeah
15:43:41 dansmith melwitt: did you read this? https://access.redhat.com/solutions/2273951
15:44:03 dansmith we're not replicated I guess so maybe that doesn't affect us in CI, but probably has some impact for real users of this
15:44:12 melwitt no
15:45:47 melwitt so there are multiple reasons MAX_AVAIL shouldn't be used :(
15:46:31 dansmith not it!
15:47:42 dansmith the other problem I'm guessing,
15:47:44 melwitt yeah... I'm thinking whether to revert that or tweak it to take total and divide by pool size, the latter would do what was actually desired and report total with replication considered
15:48:11 dansmith is that if we report the real actual total (even minus replication overhead), but other pools can consume space from the same store,
15:48:17 dansmith we will tell placement we have more room than it can allocate
15:48:50 dansmith so really we need to sum up all the pools on the same store, and then set reserved= for any space they use I guess, but then we race with those other uses in our reporting
15:48:55 dansmith and could go negative
15:49:19 melwitt yeah, I'm trying to remember, I could have sworn this get_pool_info was only used to report free space, not total space, but I could be totally making that up
15:49:45 melwitt or that that's what it's used for ultimately in higher layers
15:50:23 melwitt let me look up what "total" used to be, maybe it meant "total available"
15:51:48 melwitt no, looks like it was total. had total, total used, and total available
15:52:47 sean-k-mooney dansmith: one thing that i just tought of
15:53:02 sean-k-mooney by default replicate pools have a replciation factor of 3
15:53:17 sean-k-mooney so if we have 24G of space we would only have 8 useable
15:53:31 melwitt but looking at the clip again https://github.com/openstack/nova/blob/master/nova/virt/libvirt/storage/rbd_utils.py#L374-L382 I did parse out total_bytes to go with 'total', max_avail to go with 'free', and bytes_used to go with 'used'. so this should be fine....
15:53:32 dansmith I'm confused about whether we're replicating or not
15:53:36 dansmith sean-k-mooney: ^
15:53:51 sean-k-mooney that is the default unless we create a erasure encoded pool
15:53:54 dansmith and even still, 24/3==10 only for very small values of 3 :P
15:54:13 dansmith hmm, okay what is CEPH_REPLICAS then?
15:54:31 melwitt that's the number of replicas for when it creates the pools
15:54:41 sean-k-mooney well we have 24G for ceph but we have multiple pools right?
15:54:51 sean-k-mooney the images pool will also be using that
15:55:01 dansmith right, vms and images
15:56:55 sean-k-mooney https://github.com/openstack/devstack-plugin-ceph/blob/master/devstack/lib/ceph#L109
15:56:57 sean-k-mooney its 1
15:57:05 sean-k-mooney CEPH_REPLICAS
15:57:22 sean-k-mooney wich for ci makes sense
15:57:23 dansmith right, I think we established that earlier :)
15:57:31 melwitt yeah, I was saying earlier I've never seen CI use anything other than the default of 1
15:58:09 sean-k-mooney well we dont need to test anything else in our ci since we are not really testing ceph
15:58:18 dansmith https://pastebin.com/6gcGhTHQ
15:58:21 sean-k-mooney just ceph integration with other thngs
15:58:26 dansmith this is what my ceph df shows on a clean devstack
15:58:38 dansmith interestingly I didn't update my backing size from 8 to 24, but still got 24
15:58:50 gibi_pto so I'm going away for a week. I will be back on 3rd of Aug
15:59:01 dansmith gibi_pto: p/
15:59:06 gibi_pto o/
15:59:10 lyarwood \o
15:59:46 sean-k-mooney dansmith: i think VOLUME_BACKING_FILE_SIZE is a devstack setting
16:00:05 dansmith oh, I see, and ceph plugin uses that, gotcha
16:00:11 sean-k-mooney yes https://github.com/openstack/devstack/blob/e0d06adffcf4c8da1aefebc66f2de9a440badbf6/stackrc#L766
16:00:21 sean-k-mooney and devstack defaults it to 24
16:00:39 sean-k-mooney so that is where that is comming form
16:01:02 sean-k-mooney that was orginically for cinder
16:03:08 sean-k-mooney oh
16:03:09 sean-k-mooney https://zuul.opendev.org/t/openstack/build/13d8a055ff1b4be0b627205f4d51d50f/log/controller/logs/ceph/ceph_log.txt
16:03:19 sean-k-mooney pgmap v5: 0 pgs: ; 0 B data, 704 KiB used, 9.0 GiB / 10 GiB avail
16:03:38 sean-k-mooney so ceph does think it has only 10G
16:03:53 dansmith oh nice, but from where?
16:04:08 sean-k-mooney there is a ceph follder at the root of the contoler logs
16:04:22 dansmith from my devstack: 2020-07-24 08:25:52.400945 mgr.x client.14099 192.168.201.41:0/3299763660 2 : cluster [DBG] pgmap v5: 0 pgs: ; 0B data, 188MiB used, 23.8GiB / 24.0GiB avail
16:04:30 dansmith no I mean where is it getting the 10G
16:04:57 sean-k-mooney im wondering if we are using the filestore backend and didnt resize the filesystem or something?
16:05:09 sean-k-mooney although DF on the host shose 24G right
16:05:15 dansmith it does,
16:05:19 sean-k-mooney is that the block device size or filesystem
16:05:22 dansmith and my local devstack shows the 24G
16:05:27 dansmith filesystem
16:06:16 dansmith doesn't look like we grab the ceph configs
16:06:19 sean-k-mooney https://zuul.opendev.org/t/openstack/build/13d8a055ff1b4be0b627205f4d51d50f/log/controller/logs/ceph/ceph-osd.0_log.txt#10
16:06:33 sean-k-mooney so its using filestore
16:06:42 sean-k-mooney but wy is that 10G
16:06:56 sean-k-mooney oh its using bluestore not file store
16:06:59 sean-k-mooney but same question
16:07:01 dansmith you mean files in /var/lib/ceph right?
16:07:17 sean-k-mooney bluestore(/var/lib/ceph/osd/ceph-0) _setup_block_symlink_or_file resized block file to 10 GiB
16:07:33 sean-k-mooney its not using the mount

Earlier   Later