Skip to content
This repository was archived by the owner on Mar 31, 2026. It is now read-only.
This repository was archived by the owner on Mar 31, 2026. It is now read-only.

Performance gap between GCS Python Client and boto3 for object downloads #1458

Description

@dreamtalen

Environment details

  • GCP VM instance: c4a-highcpu-72
  • OS type and version: Linux, 6.1.0-31-cloud-arm64
  • Python version: Python 3.11.2
  • pip version: pip 23.0.1
  • google-cloud-storage version: 2.19.0 & 3.1.0

Steps to reproduce

When comparing download performance between Google Cloud Storage Python Client and boto3 (AWS SDK), we observed that GCS client is significantly slower (about 50% slower) than boto3 for downloading the same objects stored in a GCS bucket.

GCS Client Implementation (Two methods tested)

  1. Using blob.download_as_bytes():
blob = bucket.blob(key)
data = blob.download_as_bytes()
  1. Using blob.open() (10%-30% faster than method 1, but still 50% slower than boto3):
with blob.open("rb") as f:
    data = f.read()

boto3 Implementation

response = s3_client.get_object(Bucket=bucket_name, Key=key)
data = response['Body'].read()

Performance Results

With ThreadPoolExecutor(max_workers=16), I got following average throughput downloading 64MB x 1000 objects from GCS bucket to memory:

  • boto3 get_object : 12 Gbps
  • GCS download_as_bytes(): 3.2 Gbps in 2.19.0 & 4.2 Gbps in 3.1.0
  • GCS blob.open(): 4.5 Gbps in both 2.19.0 & 3.1.0

Questions

  1. Is this performance gap expected?
  2. Are there any recommended optimizations or best practices for improving download performance with the GCS Python client?
  3. Are there any internal differences in how GCS supports S3-compatible APIs handling downloads that might explain the performance gap?

Additional Context

Benchmarking scripts are available at https://github.com/dreamtalen/boto3-benchmark/tree/main/google-cloud-storage

Thanks!

Activity

  1. added
    type: questionRequest for information or clarification. Not an issue.
    on Apr 9, 2025
  2. cojenco commented on Apr 11, 2025

    @cojenco
    Contributor

    Thank you for the detailed report. To better understand this, our team will set up benchmarks in a similar environment to the one you've described to try and reproduce your findings. This will help us understand why you're seeing this so we can investigate further.

    In the meantime, could you try using Transfer Manager that utilizes multiple workers, in threads or processes, to maximize throughput. You can find more information in our docs and blog post.
    Below is a code sample on how to use transfer_manager.download_many()

    def download_many_blobs_with_transfer_manager(
    bucket_name, blob_names, destination_directory="", workers=8
    ):
    """Download blobs in a list by name, concurrently in a process pool.
    The filename of each blob once downloaded is derived from the blob name and
    the `destination_directory `parameter. For complete control of the filename
    of each blob, use transfer_manager.download_many() instead.
    Directories will be created automatically as needed to accommodate blob
    names that include slashes.
    """
    # The ID of your GCS bucket
    # bucket_name = "your-bucket-name"
    # The list of blob names to download. The names of each blobs will also
    # be the name of each destination file (use transfer_manager.download_many()
    # instead to control each destination file name). If there is a "/" in the
    # blob name, then corresponding directories will be created on download.
    # blob_names = ["myblob", "myblob2"]
    # The directory on your computer to which to download all of the files. This
    # string is prepended (with os.path.join()) to the name of each blob to form
    # the full path. Relative paths and absolute paths are both accepted. An
    # empty string means "the current working directory". Note that this
    # parameter allows accepts directory traversal ("../" etc.) and is not
    # intended for unsanitized end user input.
    # destination_directory = ""
    # The maximum number of processes to use for the operation. The performance
    # impact of this value depends on the use case, but smaller files usually
    # benefit from a higher number of processes. Each additional process occupies
    # some CPU and memory resources until finished. Threads can be used instead
    # of processes by passing `worker_type=transfer_manager.THREAD`.
    # workers=8
    from google.cloud.storage import Client, transfer_manager
    storage_client = Client()
    bucket = storage_client.bucket(bucket_name)
    results = transfer_manager.download_many_to_path(
    bucket, blob_names, destination_directory=destination_directory, max_workers=workers
    )

    Additionally, could you also try enabling client-side tracing? This will provide us with more detailed insight that can help us pinpoint any potential bottlenecks on the client side. Here are the instructions on how to enable tracing in the python storage client: https://cloud.google.com/storage/docs/traces#python

  3. dreamtalen commented on Apr 11, 2025

    @dreamtalen
    Author

    Thanks for your reply, @cojenco !

    I tried download_many, although it differs a bit from my original plan since I was benchmarking downloads directly to memory. download_many appears to only support downloads to disk.

    To minimize the impact of disk performance, I used a RAM disk as the destination. With 16 thread workers, I observed a throughput of around 2.7 Gbps, which didn’t show improvement compared to download_as_bytes() or blob.open() then read().

    transfer_manager.download_many_to_path(
                    bucket,
                    blob_names,
                    destination_directory="/mnt/ramdisk",
                    max_workers=16,
                    worker_type=transfer_manager.THREAD
                )

    As for process workers, I ran into issues with daemonic processes are not allowed to have children on my current script. Since my focus is on comparing in downloading to memory and multi-threaded setups, I decided to defer testing with process workers for now.

    Another question is, looks like blob.open() then read() ultimately calls download_as_bytes() internally (source), but it consistently performs about 30% faster. I’m curious why there’s such a noticeable difference despite the shared code?

    I’ll try client-side tracing later, thanks for your suggestion. (Update: I tried to enable Cloud Trace but ran into the permission error: Error while writing to Cloud Trace: 403 The caller does not have permission. I have granted the Cloud Trace Agent role to my service account.. will keep investigating)

  4. dreamtalen commented on May 9, 2025

    @dreamtalen
    Author

    Hi @cojenco, just checking in to see if there are any updates or if you need any help reproducing my results. I’ve uploaded the benchmarking script here:https://github.com/dreamtalen/boto3-benchmark/tree/main/google-cloud-storage. Thanks!

  5. dreamtalen commented on May 23, 2025

    @dreamtalen
    Author

    FYI, running the same download test with gcloud storage cp can get 3.5 GiB/s (28 Gbps) throughput. However, gcloud test isn't a direct apples-to-apples comparison, since I can't choose threads/processes to configure the concurrency the same as my benchmark. But I hope this baseline is a helpful diagnostic step.

    I also did benchmarks across various concurrency settings show that the google-cloud-storage Python client currently falls behind boto3 across different multithreading and multiprocessing levels. Please find the image below:
    Image

  6. shubham-up-47 commented on May 26, 2025

    @shubham-up-47
    Contributor

    Hi @dreamtalen I am trying to replicate the issue, using your scripts.

  7. shubham-up-47 commented on May 28, 2025

    @shubham-up-47
    Contributor

    Hi @dreamtalen I am trying to replicate the issue, using your scripts.

    We were able to replicate the issue, currently investigating the root cause and possible fixes.

  8. shubham-up-47 commented on Jun 18, 2025

    @shubham-up-47
    Contributor

    Hi @dreamtalen I am trying to replicate the issue, using your scripts.

    We were able to replicate the issue, currently investigating the root cause and possible fixes.

    Hi @dreamtalen, We were able to find out the root cause of the issue. The primary reason for the performance disparity is found as the default data transfer strategy,

    Apart from this, we found some little overhead due to the checksum matching logic,

    We will be introducing the single shot read/download soon. While we recommend to keep the checksum matching enabled, since it is important for the data integrity. But soon, we will give an option to disable the checksum matching through a boolean, to give a better control.

  9. dreamtalen commented on Jun 18, 2025

    @dreamtalen
    Author

    Hi @dreamtalen I am trying to replicate the issue, using your scripts.

    We were able to replicate the issue, currently investigating the root cause and possible fixes.

    Hi @dreamtalen, We were able to find out the root cause of the issue. The primary reason for the performance disparity is found as the default data transfer strategy,

    Apart from this, we found some little overhead due to the checksum matching logic,

    We will be introducing the single shot read/download soon. While we recommend to keep the checksum matching enabled, since it is important for the data integrity. But soon, we will give an option to disable the checksum matching through a boolean, to give a better control.

    Hi @shubham-up-47, that makes sense, and thanks for the detailed explanation!. Looking forward to trying it out once the single-shot version is out.

  10. shubham-up-47 commented on Jun 24, 2025

    @shubham-up-47
    Contributor

    Hi @dreamtalen, Some updates over the issue,

  11. chandra-siri commented on Jun 24, 2025

    @chandra-siri
    Collaborator

    Hi @dreamtalen, Some updates over the issue,

    Just a caution, if checksum verification is disabled, corrupted data may go undetected.

    @dreamtalen I see you've done benchmarking for google-cloud-storage version: 2.19.0 & 3.1.0

    There's a difference between both of these versions as far as the checksum verification is concerned. A snippet from official documentation.

    In Python Storage 3.0, uploads and downloads now have a default of “auto” where applicable. “Auto” will use crc32c checksums, except for unusual cases where the fast (C extension) crc32c implementation is not available, in which case it will use md5 instead. Before Python Storage 3.0, the default was md5 for most downloads and None for most uploads

    you can know more by going through Checksum Defaults

  12. shubham-up-47 commented on Jul 7, 2025

    @shubham-up-47
    Contributor

    Hi @dreamtalen,

    Now, you should be able to get high speed by enabling flags single_shot_download=True and raw_download=True in the API call download_as_bytes

    def download_as_bytes(
    in the latest clone (you can mark checksum=None if you want to disable the checksumming and make it little more faster).

    Let us know, whether this fix works well or not. Feel free to close the issue if its giving the expected speed.

  13. dreamtalen commented on Jul 7, 2025

    @dreamtalen
    Author

    Hi @shubham-up-47,

    Thanks! I’ve verified that setting single_shot_download=True delivers similar performance to boto3.

    Just curious - what exactly does raw_download=True do? I tried toggling it on and off, but didn’t notice any significant difference. Is it necessary in this case?

  14. shubham-up-47 commented on Jul 8, 2025

    @shubham-up-47
    Contributor

    Glad to hear that setting single_shot_download=True worked!

    You didn't see a difference because raw_download=True only matters in case of compressed objects. Imagine you have a file named my_data.txt.gz in your GCS bucket and it's stored in GCS with Content-Encoding: gzip.

    • raw_download=False (default): Using the default download_as_bytes(), the SDK will automatically decompress the my_data.txt.gz data for you, so you will receive the decompressed content of the file my_data.txt.
    • raw_download=True: When you use download_as_bytes(raw_download=True), you'll get the raw bytes of the my_data.txt.gz file, which are the compressed bytes. You'll need to use gzip.decompress to get the original content.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

api: storageIssues related to the googleapis/python-storage API.type: questionRequest for information or clarification. Not an issue.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions