Repository navigation
Use streams more efficiently #802
Description
Activity
stephenplusplus commented
on Aug 16, 2015 ContributorAuthorMore actionsHere's a test implementation: https://gist.github.com/stephenplusplus/b67bcaf570f33181d297
stephenplusplus commented
on Aug 21, 2015 ContributorAuthorMore actionsMoved to a repo for further testing: https://github.com/stephenplusplus/gcloud-streams-test
stephenplusplus commented
on Aug 31, 2015 ContributorAuthorMore actionsI think after all of the profiling over in the gcloud-streams-test repo, it's safe to conclude conforming to a more stream-y approach would be nice, but it comes at quite a cost; both hackiness and memory usage. The suggested approach from this issue underperforms our approach in all aspects. The repo will stay open for discussion (see stephenplusplus/gcloud-streams-test#1), but for us right now, I think we should stick with what we're doing.
I'm confused how a "give me everything, then I'll stream it to you" would be more memory efficient than "when I get a page of items, I'll stream that, and then continue fetching only if you ask me for more"...
I'd expect the "full buffer" implementation to consume more memory, and more CPU intensive. Can you explain why it's the opposite?
stephenplusplus commented
on Aug 31, 2015 ContributorAuthorMore actionsI didn't explain the pagination part that we use with our current solution. I'll update the first post.
We get all of the results back that the API gives us, drain it, then if there's another page, we repeat. I haven't seen an API response hike up into the multi-MB payload size, but if it did, it would be quickly drained. Compare that to the all-stream solution, where we spend multi-MB in memory just to create multiple Stream objects. With our current solution, we only need to create one.
Related: https://twitter.com/stephenplusplus/status/634791341172637697
stephenplusplus commented
on Aug 31, 2015 ContributorAuthorMore actions"when I get a page of items, I'll stream that, and then continue fetching only if you ask me for more"...
Sorry, just caught that. They both would behave the same here. If a user manually ends the stream (
this.end()), the stream closes, and no more API requests are made.The solution we're using now might not be right for every case, but I think given our smaller-sized API responses, it's more practical. It would be great to get more eyes on this question though (stephenplusplus/gcloud-streams-test#1) -- I think everyone is just busy, equally stumped, or indifferent :)
2 remaining items
- added 7 commits that reference this issue
on Jan 27, 2026 - added 2 commits that reference this issue
on Feb 25, 2026 - added a commit that references this issue
on Mar 12, 2026 - added a commit that references this issue
on Mar 18, 2026 - added a commit that references this issue
on Mar 27, 2026 - added a commit that references this issue
on May 5, 2026
Our support for readable streams in our library is now almost everywhere possible, which is awesome. But, the way we use them is not in the most efficient way.
How we do it, using bucket.getFiles for example:
How we could do it:
How we do it currently (simplified):
How it would look (simplified):
With the second example, the object is only in memory until it is flushed to the next handler in the user's own pipeline. At that point, the user is free to let it buffer up until they're all in or write them to a destination in chunks, not letting any extra memory pile up.
* JSONStream: https://github.com/dominictarr/JSONStream