| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-07-24 | |||
| 14:34:28 | dansmith | because we're not seeing problems, just placement is refusing to find space | |
| 14:36:59 | dansmith | Jul 24 12:44:22.575632 ubuntu-bionic-ovh-bhs1-0018770257 devstack@placement-api.service[50512]: DEBUG placement.wsgi_wrapper [req-eeb6d563-2483-4e4f-91e8-2dc3a694ade4 req-c57d5bd6-fc4e-469d-9784-cdfe1652d653 service placement] Placement API returning an error response: Unable to allocate inventory: Unable to create allocation for 'DISK_GB' on resource provider 'e786426a-5ae2-4732-8cf6-16325fd2bf2a'. The requested amount would exceed | |
| 14:37:00 | dansmith | the capacity. {{(pid=50513) call_func /opt/stack/placement/placement/wsgi_wrapper.py:31}} | |
| 14:37:16 | dansmith | Over capacity for DISK_GB on resource provider e786426a-5ae2-4732-8cf6-16325fd2bf2a. Needed: 1, Used: 10, Capacity: 10.0 | |
| 14:37:22 | dansmith | 10G doesn't sound right | |
| 14:40:45 | mriedem | random drive by comment but https://review.opendev.org/#/c/586363/ | |
| 14:41:06 | mriedem | anyway related to ceph ci jobs? | |
| 14:41:30 | dansmith | I'm trying to figure out, but we are running a ceph df right before we report inventory | |
| 14:41:40 | openstackgerrit | Alex Deiter proposed openstack/nova master: Detach is broken for multi-attached fs-based volumes https://review.opendev.org/741712 | |
| 14:54:48 | sean-k-mooney | dansmith: by the way if the traceback is unrelated then we likely have another silent bug as we are not catching the excpetion in the missing image case | |
| 14:54:56 | dansmith | sean-k-mooney: yep | |
| 14:55:06 | dansmith | so we're calling ceph df to get the total size of the pool and reporting that | |
| 14:55:07 | sean-k-mooney | dansmith: i think you are right that its unrelated | |
| 14:55:16 | dansmith | as best I can tell, the ceph is backed by a 24G partition | |
| 14:55:27 | dansmith | so I dunno where the 10G is coming from | |
| 14:56:04 | sean-k-mooney | this is using the ceph image backend in nova so the local_GB should be the ceph pool size right | |
| 14:56:25 | dansmith | well, it should be yes | |
| 14:56:43 | dansmith | ceph has 24G, so I'm trying to find where our images pool would be limited to 10G but not seeing it | |
| 14:56:47 | dansmith | one thing that might explain this, | |
| 14:57:09 | dansmith | is that our normal ceph job was using qcow on rbd, which is not what you're supposed to do, | |
| 14:57:24 | sean-k-mooney | oh ya because we have to flatten it | |
| 14:57:31 | sean-k-mooney | it should be raw | |
| 14:57:36 | sean-k-mooney | to get the cow optimization | |
| 14:57:40 | dansmith | and so we convert the image to raw, which is 44M per image instead of 12 or something.. although we shouldn't really be using that much space, so... hmm | |
| 14:57:53 | dansmith | and this is just placement saying we're out of space, not ceph | |
| 14:58:17 | dansmith | I wonder if glance is incorrectly determining the size of the new image after it flattens or something | |
| 14:58:19 | sean-k-mooney | well with after teh first image import is all cow clones in ceph right | |
| 14:58:23 | dansmith | and telling us we need a lot more than we do or something | |
| 14:58:30 | dansmith | right | |
| 15:01:20 | dansmith | check this out: Jul 24 12:44:22.270293 ubuntu-bionic-ovh-bhs1-0018770257 nova-scheduler[55176]: WARNING nova.scheduler.host_manager [None req-eeb6d563-2483-4e4f-91e8-2dc3a694ade4 tempest-MultipleCreateTestJSON-968818181 tempest-MultipleCreateTestJSON-968818181] Host ubuntu-bionic-ovh-bhs1-0018770257 has more disk space than database expected (8 GB > 1 GB) | |
| 15:01:37 | sean-k-mooney | reserved_host_disk_mb IS 0 TOO | |
| 15:02:46 | sean-k-mooney | that is strange do we have the hoststate update enabled | |
| 15:03:09 | sean-k-mooney | im pretty sure we do | |
| 15:03:44 | sean-k-mooney | ya we do | |
| 15:04:23 | sean-k-mooney | disk_allocation_ratio=1.0,disk_available_least=8,free_disk_gb=10,f | |
| 15:04:33 | dansmith | we're only asking placement for DISK_GB=1 allocation so I don't think we're getting a bad number from glance or anything | |
| 15:06:15 | sean-k-mooney | what do our flavor look like | |
| 15:06:27 | sean-k-mooney | actully no never mind | |
| 15:06:30 | sean-k-mooney | this is not bfv | |
| 15:06:51 | sean-k-mooney | the flavor should be either 1 or 2GB per instance i think | |
| 15:07:20 | dansmith | and it seems like 1 since we're asking for that size allocation | |
| 15:09:44 | sean-k-mooney | ya its based on teh image size https://github.com/openstack/devstack/blob/2ecd1823850ae0e00ad0ecebbbceb312be60ccf4/lib/tempest#L204-L206 | |
| 15:09:53 | sean-k-mooney | so for cirros image it will be 1g | |
| 15:10:04 | dansmith | sudo ceph -c /etc/ceph/ceph.conf osd pool create vms 8 8 | |
| 15:10:10 | dansmith | that's 8G for the vms pool | |
| 15:10:18 | dansmith | I dunno where we're getting 10G | |
| 15:10:26 | sean-k-mooney | i dont think that is the size | |
| 15:10:43 | sean-k-mooney | i think that is the buckest to share it in | |
| 15:10:45 | sean-k-mooney | let me check | |
| 15:11:15 | dansmith | hmm, okay it seems like size | |
| 15:11:35 | sean-k-mooney | i think its the placment groups but its been a while | |
| 15:11:41 | dansmith | okay yeah, maybe you're right | |
| 15:12:39 | sean-k-mooney | ceph osd pool create <pool-name> <pg-num> <pgp-num> [replicated] \ | |
| 15:12:41 | sean-k-mooney | [crush-ruleset-name] [expected-num-objects] | |
| 15:12:52 | sean-k-mooney | so ya its not the size | |
| 15:13:04 | dansmith | yeah | |
| 15:13:22 | dansmith | I still dunno where we're getting 10G, | |
| 15:13:34 | sean-k-mooney | same | |
| 15:14:27 | dansmith | because CEPH_LOOPBACK_DISK_SIZE=24G | |
| 15:14:57 | sean-k-mooney | so it should be 8 https://github.com/openstack/devstack-plugin-ceph/blob/master/devstack/settings#L17 | |
| 15:15:01 | sean-k-mooney | by default | |
| 15:15:12 | dansmith | it's overridden in our job somewhere | |
| 15:15:16 | dansmith | you can see in the devstacklog | |
| 15:15:24 | sean-k-mooney | CEPH_LOOPBACK_DISK_SIZE is | |
| 15:15:31 | sean-k-mooney | is VOLUME_BACKING_FILE_SIZE | |
| 15:15:34 | dansmith | ues | |
| 15:15:36 | dansmith | both are | |
| 15:15:47 | sean-k-mooney | ok cool | |
| 15:16:12 | dansmith | VOLUME_BACKING_FILE_SIZE=24G | |
| 15:16:18 | dansmith | and the df shows 24G on /var/lib/ceph | |
| 15:17:48 | sean-k-mooney | ah yes it does | |
| 15:17:55 | dansmith | we run "ceph df" to get the DISK_GB we report, | |
| 15:17:56 | dansmith | and don't really do much to it, | |
| 15:18:08 | dansmith | so it really seems like we're being told 10G | |
| 15:19:52 | dansmith | lyarwood: do you know anything about what ceph df may be telling us about total pool size that differs from the backing store's size? | |
| 15:20:34 | lyarwood | dansmith: nope, AFAIK it just reports the size of the images_rbd_pool | |
| 15:21:16 | dansmith | seems straightforward :) | |
| 15:21:38 | lyarwood | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/storage/rbd_utils.py#L374-L382 - ah well melwitt has a handy comment here that might help | |
| 15:21:50 | lyarwood | gah highlights are a new lines off but you get the point | |
| 15:21:58 | dansmith | oh I read that, | |
| 15:22:01 | dansmith | but didn't grok until now | |
| 15:22:10 | dansmith | so replication makes the thing looks smaller I guess? | |
| 15:22:37 | dansmith | seems weird to go from 24G to 10G, as that's not an even factor | |
| 15:23:15 | dansmith | er, no wait | |
| 15:23:24 | dansmith | that's for max_avail, which is "free" not total right? | |
| 15:24:30 | lyarwood | right sorry and you're seeing 10 reported as the total capacity right? | |
| 15:24:34 | dansmith | correct | |
| 15:24:40 | lyarwood | kk sorry then that isn't it | |
| 15:26:43 | dansmith | I guess one thing we could do is increase the ceph backing size to 36G and see if DISK_GB goes up | |
| 15:26:45 | bauzas | can someone tell me what the fuck is ? http://paste.openstack.org/show/796292/ | |
| 15:26:57 | bauzas | tl;dr: ssh: connect to host review.openstack.org port 29418: Network is unreachable | |
| 15:27:11 | bauzas | have I missed a memo ? | |
| 15:27:25 | melwitt | there's a openstackstatus above ^ said there will be a short outage | |
| 15:29:00 | openstackgerrit | Sylvain Bauza proposed openstack/nova-specs master: WIP: Offline Reshape tool spec https://review.opendev.org/742908 | |
| 15:29:05 | bauzas | yay, it worked | |
| 15:29:09 | bauzas | melwitt: thanks | |
| 15:29:46 | bauzas | calling it a day | |
| 15:32:15 | melwitt | dansmith: MAX_AVAIL should be total actually, just taking number of replicas into account. if you only have 1 replica (default NUM_REPLICAS=1) then MAX_AVAIL should match whatever total says in 'ceph df' | |
| 15:32:49 | dansmith | melwitt: you're reporting free as max_avail though in that thing aren't you? | |
| 15:32:59 | dansmith | or does MAX_AVAIL != max_avail ? | |