Earlier  
Posted Nick Remark
#openstack-sdks - 2018-07-24
19:44:50 mordred in both cases the data is uploaded in a set of 'segment' objects. in static large objects there is then a manifest object that lists the segment objects and the order in which they should be served
19:45:33 mordred in dynamic the segment objects follow a naming scheme with an empty manifest object (so you don't have to keep track of the list and then upload the manifest object)
19:45:52 mordred each of the segment objects is a completely independent swift object, so I'd expect uploading to them with deflate would work
19:46:40 mordred oh - then when you fetch the objects, as a client, you just fetch the manifest object, and swift knows that what you want to do is stream the concatenation of the segment objects
19:47:11 mordred it's the download that I'm not sure what happens
19:47:42 corvus the download?
19:47:49 mordred the fetch
19:48:32 corvus i'm confused, i thought you just described what happens in that case
19:48:58 mordred yes - in the standard case - I do not know how it interacts with streaming compression
19:49:12 mordred largely becaues I've never tested it - I'd imagine it would DTRT
19:49:17 mordred but don't actually know
19:49:26 corvus oh i see what you're saying
19:50:54 mordred corvus: out of curiosity - what is the iterator that has the data in it?
19:51:28 mordred oh. the deflatefilter
19:51:35 mordred *duh* it's in the traceback
19:52:29 corvus yeah. it's a thing i wrote that compresses 16k at a time and produces an unknown amount of data each chunk
19:55:14 mordred corvus: well - here's another question then ...
19:57:26 mordred corvus: does max_file_size apply to the compresesd or uncompresed value
19:57:41 corvus excellent question! it's all very vague
19:58:21 corvus i'm assuming, for the moment, that swift doesn't attempt to do any decoding. so probably it applies to the compressed value.
19:58:29 corvus (i'm assuming it is stored in the way i send it)
19:59:04 corvus i have little factual basis for making that assumption
20:00:39 mordred yeah
20:00:52 mordred corvus: for now, if you make a len method on your object that just returns 0
20:00:57 mordred if should get you past your issue
20:01:30 mordred file_size is only used a little later to determine if file_size < max_segment_size ...
20:01:43 mordred so if you return 0, the code pass data on through directly
20:02:33 openstackgerrit Merged openstack/os-service-types master: Updated from OpenStack Service Type Authority https://review.openstack.org/585062
20:02:33 openstackgerrit Merged openstack/os-service-types master: Switch to stestr https://review.openstack.org/585349
20:03:03 openstackgerrit Monty Taylor proposed openstack/openstacksdk master: Update create_object to handled chunked data https://review.openstack.org/585532
20:03:14 mordred corvus: ^^ also, I think that should at least do a workaround
20:03:42 corvus mordred: though does it get passed through to the underlying session method? i need that not to have a len method, or it'll avoid chunked uploads.
20:04:54 corvus mordred: i'll try your patch out in a little bit
20:05:13 mordred corvus: ah.
20:05:48 mordred corvus: oh for the love of ...
20:06:34 mordred corvus: we don't use file_size in the data path at all
20:07:58 openstackgerrit Monty Taylor proposed openstack/openstacksdk master: Update create_object to handled chunked data https://review.openstack.org/585532
20:08:18 mordred corvus: ^^ that should fix it more appropriately
20:09:08 corvus mordred: ack. i'll give it a spin when i finish dealing with trailing slashes :)
20:09:19 mordred corvus: trailing slashes are the worst
20:09:29 mordred corvus: they're almost as terrible as INCORRECT WHITESPACE
20:09:38 notmyname mordred: hello
20:09:45 mordred yay it's notmyname !
20:10:30 mordred notmyname: we have several questions - I will try to ask them in some semblance of order
20:11:41 mordred notmyname: first of all, if you are uploading a large object and you want to upload compressed with a deflate header - does the compressed or uncompressed size count towards max_file_size?
20:12:36 corvus specifically, "content-encoding: deflate" is what's happening here
20:12:46 mordred yeah. corvus is likely to say smarter words than me
20:14:55 notmyname that is likely to be dependent on something between the client and swift itself
20:15:15 notmyname there's not anything in swift that will accept compressed data and store in uncompressed
20:15:39 notmyname however, if you've got some caching thing (CDN or otherwise) that understands those headers, that work can be done there
20:16:02 notmyname swift will happily store the content-encoding header, if you send it, and return it on a read request.
20:17:49 notmyname lol, rackspace took my name off the author byline ;-) https://blog.rackspace.com/cloud-files-cdn-compresses-at-the-edge
20:17:57 mordred ok. so - what about SLO/DLO objects and concatenation? if the segments are uploaded compressed and swift doesnt' natively do any uncompression, I'm guessing that could get weird for the read?
20:18:08 mordred notmyname: haha
20:18:45 mordred and by 'get weird' I mean 'not work"
20:19:31 notmyname let me try something...
20:21:17 timburke i'd expect you'd want to open the large object, stream it through a compressor, and break out segments from that compressed stream. that way when you go to download the large object, you'll get a singular large, compressed stream
20:23:05 timburke breaking the large object into segments first then compressing is unlikely to end well
20:23:11 notmyname oh hi timburke
20:23:21 corvus makes sense
20:23:39 corvus unfortunately, we have to know ahead of time if we're going to upload a large object or a normal one yeah?
20:24:10 notmyname yeah, what he said. "breaking the large object into segments first then compressing is unlikely to end well". and I just confirmed I wasn't forgetting something about this against a dev box
20:25:13 mordred then I guess you could put content-encoding: deflate header on the manifest object - and a browser would theoretically dtrt?
20:25:31 notmyname corvus: you can use an SLO even if the total object size is much less than a single "normal" object limit
20:25:50 notmyname mordred: ya
20:26:11 mordred notmyname: oh. well that's certainly an interesting thought ...
20:26:45 mordred notmyname, timburke: the overall problem we're trying to solve is what to do with the intersection of openstacksdk transparently creating large objects for you and a user of openstacksdk wanting to pass in an interable that is a compressed stream
20:27:19 mordred it seems like one way to deal with it might be to just always create a SLO if someone passes in an interable instead of a bytes or a filename
20:27:20 timburke corvus: depends on how many api requests you're willing to make :-) one option would be to always upload as a large object (like notmyname said) or upload one segment's worth to the base name, then do a server-side copy to the segment location once you realize you need a large object
20:27:46 mordred ooh. that second one sounds reasonable too
20:27:54 corvus if it's not crazy to create a SLO when not strictly necessary, maybe we could make the decision based on the size of the uncompressed data. so if it's > max size, go ahead and SLO even if it's not strictly necessary.... we'd still only do it for "big" files :)
20:27:57 corvus or that second one. :)
20:28:19 notmyname is this for log files?
20:28:33 notmyname for the zuul jobs?
20:28:44 mordred notmyname: yup
20:28:52 corvus notmyname: for starters (so unlikely to hit it) but container/machine images probably aren't far behind.
20:29:04 timburke if you've got enough memory, you could buffer the first MB or so, if it all fits do it as a normal object; otherwise fall back to SLO
20:29:12 notmyname then in that case, I'd optimize for simpler client write path instead of optimal read latency
20:29:23 notmyname since these will be frequently written and rarely read
20:29:39 notmyname timburke has the right idea
20:30:16 notmyname .read(1024*1024) on the input, if you get the full MB, then do a SLO. if not, write a normal object
20:31:19 corvus memory is actually an issue; we could end up attempting a lot (hundreds? many many hundreds?) of these simultaneously on a 8g vm
20:36:39 timburke no upload pooling? i feel like you'd probably be able to saturate your i/o with tens of workers rather than hundreds... but maybe this is getting into the need to have a simple client
20:37:24 timburke server-side copy (or always SLO, all the time) may work out best
20:38:11 notmyname corvus: mordred: so the general answer is that swift will store the bytestream you send it and also headers that may have some definition for clients (eg content-encoding). SLOs aren't special in that the segments are simply slices of the resulting range. swift doesn't do any interpretation of the contents of objects
20:39:05 mordred timburke: yah - server-side copy or always SLO both sound like good general options
20:39:45 mordred there's definitely a balancing act we're trying to do here with wanting SDK to DTRT and yet also providing enough knobs so that we can do the zuul log upload thing efficiently
20:41:45 mordred notmyname, timburke: thanks both of you - this has been super helpful
20:42:02 corvus ++
20:43:46 timburke fwiw, swiftclient opts for the buffering thing when uploading from stdin -- i think we go up to 16MB (or something like that?) then start uploading 16MB segments. since its stdin, there's only one upload per-process, so we don't feel too bad about the memory
20:49:33 mordred timburke: yah - the fun part of this story is that once sdk switches to "oh, you wanted an SLO" - it does so with a pool of threads (similar to swiftuploader in swiftclient)
20:49:56 mordred of course, actually ... now that I think about it - that won't work for iterators that don't have seek anyway
20:50:51 mordred since it does it in parallel for files by opening multiple handles and seeking on them ... so to support SLO from an input stream we'll need to reengineer what we're doing anyway
20:50:53 mordred corvus: ^^
20:51:38 corvus this may be a limited use case. we won't want to use the deflatefilter for, say, already compressed images. i've only got it set up to engage for text/ types with no encoding right now.
20:51:57 mordred corvus: ah - cool.
20:52:15 mordred corvus: I'm almost starting to feel like we should add your compressiong streaming code into sdk itself so that we can put it further down the stack
20:52:45 mordred corvus: like, put it around the file reads after the seek in the SLO segment uploads
20:53:07 corvus mordred: i think it may be pretty domain-specific; i don't think it's appropriate for everything
20:53:07 mordred and make an option to create_object "compress=False" or something like that (just thinking out loud)

Earlier   Later