Skip to content

bucket api using a lot of memory due to retry-request #1392

Description

@richtera

Since the bucket api is using the retry-request module it's growing my node heap to around 7G due to large files being up and downloaded. Node is not willing to give up all that memory right away. Since I really don't need retry for such large requests and would rather know when an error happens due to rate restrictions is there a way to turn of the usage of retry-request?
Thanks
Andy

Activity

  1. stephenplusplus commented on Jun 24, 2016

    @stephenplusplus
    Contributor

    You can disable the retries by providing { maxRetries: 0 } to Storage:

    var gcs = gcloud.storage({ maxRetries: 0 });

    Let me know if that helps.

  2. richtera commented on Jun 24, 2016

    @richtera
    Author

    That will prevent a retry, but I need it to revert to using a plain request rather than a retry-request. The retry-request will pipe all of the request and response data into a streamcache which is what's taking up all the memory. I need to be able to disable caching of the stream data in memory.
    Thanks
    Andy

  3. stephenplusplus commented on Jun 24, 2016

    @stephenplusplus
    Contributor

    Ah, I see. I didn't realize stream-cache was keeping that in memory indefinitely. I'll see if I can think of something to release those after we don't need the cache stream anymore.

  4. richtera commented on Jun 24, 2016

    @richtera
    Author

    It's not indefinitely and does get eventually released but node does not free up the underlying OS memory for quite a while and maybe never. And for large files I would rather have it fail than clobbering memory to this extent.
    More detail... I think the file object will keep the streamCache in scope. I had to refactor to have the file object go out of scope before the heapdump showed the memory released.
    Thanks
    Andy

  5. stephenplusplus commented on Jun 24, 2016

    @stephenplusplus
    Contributor

    I would guess the buffers a stream-cache instance retains go out of scope when the stream itself goes out of scope. But I think we could try only retaining that data until we pipe a single stream to it.

    In other words, I believe stream-cache now is holding internally all of the data it receives, because it thinks we want to replay them to possibly multiple streams. I think we should modify it to stop holding data internally after it has already had a stream connected to it. I'll put this into code and see if this works. Hope you won't mind if I ping you to try it out!

  6. stephenplusplus commented on Jun 24, 2016

    @stephenplusplus
    Contributor

    @richtera I made a branch of retry-request that removes stream-cache completely, instead just using a normal passthrough stream to hold data. Can you see if it improves the memory usage in your project?

    $ npm install --save stephenplusplus/retry-request#rm-stream-cache
  7. richtera commented on Jun 27, 2016

    @richtera
    Author

    Works great. My memory footprint went from 7G to 700M :). Sorry it took a couple of days to validate.
    Thanks.
    Andy

  8. stephenplusplus commented on Jun 27, 2016

    @stephenplusplus
    Contributor

    Great! Thanks for catching this, it's going to be very appreciated by our users :) I'm going to run it through a few more tests, then merge and release 👍

  9. stephenplusplus commented on Jun 27, 2016

    @stephenplusplus
    Contributor

    Released the fix in retry-request under 1.3.1. gcloud will automatically pick it up in new installs. Thanks again!

  10. 22 remaining items

  11. added 2 commits that reference this issue on Feb 26, 2026
    50e7d64
    180abce
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

coretype: questionRequest for information or clarification. Not an issue.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions