| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-09-05 | |||
| 21:10:39 | dansmith | but technically thin provisioning can blow up in your face if you overcommit | |
| 21:10:47 | cdent | right, that's exactly the issue | |
| 21:10:50 | dansmith | I do know that. | |
| 21:11:22 | dansmith | IMHO, allocation_ratio can only be over 1.0 for disk if you like to gamble | |
| 21:11:34 | dansmith | and shouldn't really be related to use of thin provisioning, | |
| 21:11:42 | dansmith | except in that it may be the mechanism by which you gamble | |
| 21:12:00 | cdent | the nearby concrete problem here is: datastore with 50TB total, allocated to 50TB, but in reality has 41TB free because of "thing provisioning" | |
| 21:12:19 | dansmith | but yep | |
| 21:12:19 | cdent | so the short term fix is bump allocation_ratio and watch real usage "real close like" | |
| 21:12:27 | dansmith | yeah, i.e. gambling | |
| 21:12:33 | cdent | that was not mockery, I just can't type | |
| 21:12:37 | cdent | which you also know | |
| 21:12:50 | dansmith | so, every night before the op goes to sleep, | |
| 21:13:07 | dansmith | he decides how far he thinks he can get before 8am the next morning and sets the allocation_ratio accordingly | |
| 21:13:13 | cdent | pretty much | |
| 21:13:32 | dansmith | solid plan | |
| 21:14:40 | cdent | so the questioning is: if you know what placement thinks about usage, and you know what the datastore thinks of its free space, you ought to be able to calculate a dynamic allocation ratio every now and again, and save that op some sleep | |
| 21:15:12 | dansmith | if your usage is very consistent | |
| 21:15:23 | dansmith | the other thing to think about/remember is: | |
| 21:15:52 | dansmith | thin provisioning is somewhat of a space-saving thing, but it's also just to speed up the actual provisioning step | |
| 21:16:06 | dansmith | most filesystems will end up filling out the block device over time, | |
| 21:16:29 | cdent | yeah, which is why it would need to be a failure regular thing | |
| 21:16:33 | dansmith | unless they intentionally compact and fstrim() now and then to let go of things back to the underlying storage | |
| 21:16:36 | cdent | like maybe every update_provider_tree | |
| 21:17:00 | cdent | that was an awesome freudian slide on my part: meant fairly, but failure works too | |
| 21:17:16 | dansmith | but if you go to sleep with it set at 1.1, | |
| 21:17:29 | dansmith | and something happens where a big tenant has a script that runs amok and fills the disk with logs, | |
| 21:17:37 | dansmith | you could end up getting a page | |
| 21:17:55 | dansmith | so it's still a ticking time bomb, even if you calculate a conservative ratio based on recent history | |
| 21:18:24 | cdent | I'm not sure I'm following that logic: at time 1 we set it to 1.1. a few minutes later the actual usage on the store goes up, so at time 2 we set it to 1.0 | |
| 21:18:48 | cdent | if 1 and 2 are close together the risk is lessed (but not entirely removed) | |
| 21:19:00 | dansmith | you mean some automated thing that looks for the crash on the horizon and adjusts the ratio before it becomes a problem? | |
| 21:21:34 | dansmith | I'm trying to think about how that helps | |
| 21:21:36 | cdent | in this vmware case the datastore knows how many bytes it says it has "free". this is different from what placement says. The ratio of that difference can be a multiplier for setting allocation_ratio. If you're saying "the datastores sense of 'free' is not real" then yeah, sure, we got a problem. | |
| 21:21:46 | dansmith | until you get to 1.0, it might as well be 1.0, and once you're past 1.0 you can't go back | |
| 21:22:10 | dansmith | it's free to the datastore, but it's not "uncommitted" | |
| 21:22:58 | cdent | I think that may be where the vmware stuff is actually doing something "amazing" that is a bit different from sparse files | |
| 21:23:37 | dansmith | the other two things that it can be doing are compression and dedup | |
| 21:23:48 | dansmith | both of which have an optimal ratio, but that aren't consistent | |
| 21:24:23 | dansmith | like, if your users are into storing usenet archives, you might get 1.5x from compression, but if they're archiving "those" usenet articles, it'll be much closer to 1.0x | |
| 21:24:50 | dansmith | and if multiple tenants are storing the same things, then you might get 2x from dedup, but not if they're using different block sizes or encruption | |
| 21:24:55 | dansmith | *encryption | |
| 21:24:57 | dansmith | dang | |
| 21:24:58 | dansmith | you know. | |
| 21:25:10 | cdent | yah, all that makes sense | |
| 21:25:30 | cdent | but the fundamental issue here is 41TB free, but placement says no bytes free | |
| 21:25:33 | cdent | *issue for the customer | |
| 21:25:57 | dansmith | right, I get that | |
| 21:26:06 | dansmith | 41TB is free because 41TB hasn't yet been used, | |
| 21:26:28 | dansmith | but it _is_ committed to some instances | |
| 21:26:29 | dansmith | according to nova and placement | |
| 21:26:48 | cdent | and for many of these flavors, probably never will, because somebody made big flavors, and people like big things even when they don't need them | |
| 21:27:02 | cdent | yes to all that | |
| 21:27:05 | dansmith | you can tell placement to oversubscribe that to converge the two numbers to zero, but the end result will be someone writes a byte to disk and gets an EIO because you lost the bet | |
| 21:27:25 | dansmith | yeah, but again, with disk, it all ends up filling out | |
| 21:27:38 | dansmith | if you have a linux guest with a 1TB disk and 500MB used, | |
| 21:27:45 | dansmith | and you write a 10MB file and then delete it, in a loop, | |
| 21:27:53 | dansmith | you will end up touching all sectors on the disk | |
| 21:28:18 | dansmith | if you're fstrim()'ing then that will tell the underlying disk that you're only using some of those at a time, | |
| 21:28:21 | dansmith | but if you're not, it doesn't know | |
| 21:28:32 | dansmith | now, I'll grant you this: | |
| 21:28:35 | cdent | as I understand things, in the vc situation, the datastore will reject a workload if _it_ thinks it doesn't have enough space to cope | |
| 21:28:49 | cdent | so it's not going to lead to that level of disaster, at least not as quickly | |
| 21:28:59 | dansmith | if the expectation or requirement is that the guest agent is running in all these instances and can ensure that it's trim'ing then you won't converge on full in the same way | |
| 21:29:23 | cdent | the split brain nature of the vc situation, sets up a lot of weird | |
| 21:29:25 | dansmith | based on what logic? | |
| 21:29:36 | cdent | vmware magic? | |
| 21:29:49 | cdent | it's what that stuff does. who knows how it works. | |
| 21:30:07 | dansmith | well, unless you can explain some magic, I don't believe :) | |
| 21:30:16 | cdent | let me see if I can find something | |
| 21:30:49 | mriedem | https://www.youtube.com/watch?v=1-bt81UBDLI ? | |
| 21:34:08 | cdent | hrmm, too many layers of indirection | |
| 21:34:50 | cdent | so: in the general case a fix for this kind of thing is: train people to not use big disk using flavors when they don't need that. use a volume. | |
| 21:35:40 | dansmith | I'm not saying you shouldn't gamble, | |
| 21:35:52 | dansmith | I'm just saying if you do, be sure to let me know so I can move my instances elsewhere :) | |
| 21:36:37 | dansmith | I think most probably do because they're economizing, and they take that risk, mitigated by willingness to throw disks at a growing problem | |
| 21:36:51 | dansmith | if you're going for risk-and-reward, you gamble | |
| 21:37:06 | dansmith | if you're going for "make sure the nukes don't EIO at a bad time" then maybe you don' | |
| 21:37:07 | dansmith | t | |
| 21:38:06 | cdent | aye aye | |
| 21:38:48 | cdent | still querying around for an answer on the "is there magic here" and so far the answer is "it's complicated" | |
| 21:38:57 | cdent | and that, kids, is why grampy chris hates everything | |
| 21:40:15 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove redundant image GET call in _do_rebuild_instance https://review.openstack.org/600260 | |
| 21:40:19 | mriedem | cdent: i imagine the VC resource tracker is infinitely worse than nova's | |
| 21:40:24 | mriedem | which gives me some comfort | |
| 21:40:38 | mriedem | by worse, i mean complicated | |
| 21:41:08 | cdent | thankfully I get to stay far away from that stuff | |
| 22:13:05 | mriedem | fun, just saw a gate failure where rebuild creates the libvirt guest then while polling the domain to see if it's running, libvirt says it's not found | |
| 22:13:08 | mriedem | and we blow up | |
| 22:24:40 | openstackgerrit | Elancheran S proposed openstack/nova stable/pike: Add exact match aggregate image properties matcher/filter https://review.openstack.org/599870 | |
| 22:25:32 | openstackgerrit | Elancheran S proposed openstack/nova master: Add exact match aggregate image properties matcher/filter https://review.openstack.org/593167 | |
| 22:31:15 | openstackgerrit | Jonte Watford proposed openstack/nova master: Modified version of 0027-Numa-object-string-representations.patch with some updates from the current numa files for nova: numa.py instance_numa_topology.py https://review.openstack.org/600269 | |
| 23:31:02 | Sundar | melwitt: Please ping me when you can | |
| #openstack-nova - 2018-09-06 | |||
| 00:42:36 | openstackgerrit | fupingxie proposed openstack/nova master: Add an example to add more pci devices in nova.conf https://review.openstack.org/592243 | |
| 00:49:52 | openstackgerrit | Merged openstack/nova master: Move str to six.string_types https://review.openstack.org/599493 | |
| 00:49:59 | openstackgerrit | Merged openstack/nova master: Fix a failure to format config sample https://review.openstack.org/597986 | |
| 01:02:19 | naichuans | mriedem: Got it, Matt. I can't got PTG this time, maybe bauzas: and efried: could take care about it with you? | |
| 01:22:23 | melwitt | Sundar: I got your email, I'll add info to our etherpad about cyborg/nova on monday from 2-3pm | |
| 01:23:46 | openstackgerrit | fupingxie proposed openstack/nova master: Delete allocations for instances that have been moved to another node https://review.openstack.org/582899 | |
| 01:33:02 | Sundar | melwitt: Thank you | |