Skip to content

Remove resumable option from file.createWritableStream #399

Description

@ryanseys

If everyone is for it, I will remove the resumable option from file.createWritableStream in GCS. It will default to false, because you can't really resume a stream reliably. This is how the gsutil behaves. If you pipe in data to gsutil, it will not support resumable. This will help with #397.

Looking for a 👍 and I'll go ahead and rip it out.

Activity

  1. added
    type: questionRequest for information or clarification. Not an issue.
    api: storageIssues related to the Cloud Storage API.
    on Feb 18, 2015
  2. self-assigned this
    on Feb 18, 2015
  3. added this to the Storage Stable milestone on Feb 18, 2015
  4. stephenplusplus commented on Feb 19, 2015

    @stephenplusplus
    Contributor

    Why can't we resume a stream reliably? .upload is just a wrapper around .createWriteStream, so any security felt from using .upload will be superficial.

    I think it being on by default is okay. If you're in the situation like #397, turning it off may be best for your use case. Don't we only enable resumable if the file is over 5 MB? Maybe it's worth considering pushing that limit higher before shutting off by default?

  5. ryanseys commented on Feb 24, 2015

    @ryanseys
    ContributorAuthor

    But if the data is coming from a stream, we don't know if it's coming from a file somewhere, so how can we detect when a file is reused? Similarily, if the data is coming from a stream, how can we know what the size will be to determine if it should be resumed or not resumed? In the case of upload, we're given a filename so can stat the file to determine its size and store the filename somewhere to determine later if the file is attempting to be re-uploaded.

  6. stephenplusplus commented on Feb 24, 2015

    @stephenplusplus
    Contributor

    so how can we detect when a file is reused?

    The way we are doing it now. We have all the pieces we need: the target file name and the data being sent. If a user starts another upload attempt, and we see that we've recorded a failed upload to that target, we detect that the incoming data is the same based on the first chunk of data sent in.

    The command line tool can do resumable streams, I think we should too.

    Just caught the part in your opening message about gsutil not supporting piped resumable uploads. If those engineers decided it's not reliable, then I guess I can't suggest that we outsmarted them. But really, I think our implementation is good.

    Similarily, if the data is coming from a stream, how can we know what the size will be to determine if it should be resumed or not resumed?

    Forgot about that.

  7. ryanseys commented on Feb 24, 2015

    @ryanseys
    ContributorAuthor

    I'm reading their docs now and I think you may be correct and I'm wrong. See here: https://cloud.google.com/storage/docs/gsutil/commands/cp#streaming-transfers

    From the link:

    Streaming transfers (other than uploads using the JSON API) do not support resumable uploads/downloads. If you have a large amount of data to upload (say, more than 100 MiB) it is recommended to write the data to a local file and then copy that file to the cloud rather than streaming it (and similarly for large downloads).

    By default gsutil uses the the JSON API so I guess they do support resumable?

  8. ryanseys commented on Feb 24, 2015

    @ryanseys
    ContributorAuthor

    Well if it works and actually helps performance when the file is huge, then cool! 👍 I just would like a reasonable default value for performance gainz (TM). When a file is small though, in the case of the benchmark test, using resumable seems like too much overhead. How can we deal with that?

  9. stephenplusplus commented on Feb 24, 2015

    @stephenplusplus
    Contributor

    I think we just have to make a choice:

    • resumable on by default for safety, but allow turning it off
    • resumable off by default for perf, but allow turning it on

    My vote still goes to letting a developer shut it off for small uploads. The overhead from one upload operation of a 1MB file isn't going to be noticeable. And if your app disagrees, then just turn it off. The perf gainz (TM) noticed in #397 are because it is uploading 750 files at once. If that's a use case for your app, I would say again, you might benefit from disabling it.

  10. ryanseys commented on Feb 24, 2015

    @ryanseys
    ContributorAuthor

    Closing this, as we discussed "offline", we'll keep this on by default and delegate some responsibility to the developer in choosing the appropriate time to turn it off (e.g. when they are uploading 750 50kb files)

  11. added a commit that references this issue on Aug 22, 2022
  12. 23 remaining items

  13. added a commit that references this issue on Feb 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

api: storageIssues related to the Cloud Storage API.type: questionRequest for information or clarification. Not an issue.

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions