Repository navigation
Performance gap between GCS Python Client and boto3 for object downloads #1458
Description
Activity
- addedapi: storageIssues related to the googleapis/python-storage API.Issues related to the googleapis/python-storage API.
on Apr 7, 2025 - addedtype: questionRequest for information or clarification. Not an issue.Request for information or clarification. Not an issue.
on Apr 9, 2025 Thank you for the detailed report. To better understand this, our team will set up benchmarks in a similar environment to the one you've described to try and reproduce your findings. This will help us understand why you're seeing this so we can investigate further.
In the meantime, could you try using
Transfer Managerthat utilizes multiple workers, in threads or processes, to maximize throughput. You can find more information in our docs and blog post.
Below is a code sample on how to usetransfer_manager.download_many()def download_many_blobs_with_transfer_manager( bucket_name, blob_names, destination_directory="", workers=8 ): """Download blobs in a list by name, concurrently in a process pool. The filename of each blob once downloaded is derived from the blob name and the `destination_directory `parameter. For complete control of the filename of each blob, use transfer_manager.download_many() instead. Directories will be created automatically as needed to accommodate blob names that include slashes. """ # The ID of your GCS bucket # bucket_name = "your-bucket-name" # The list of blob names to download. The names of each blobs will also # be the name of each destination file (use transfer_manager.download_many() # instead to control each destination file name). If there is a "/" in the # blob name, then corresponding directories will be created on download. # blob_names = ["myblob", "myblob2"] # The directory on your computer to which to download all of the files. This # string is prepended (with os.path.join()) to the name of each blob to form # the full path. Relative paths and absolute paths are both accepted. An # empty string means "the current working directory". Note that this # parameter allows accepts directory traversal ("../" etc.) and is not # intended for unsanitized end user input. # destination_directory = "" # The maximum number of processes to use for the operation. The performance # impact of this value depends on the use case, but smaller files usually # benefit from a higher number of processes. Each additional process occupies # some CPU and memory resources until finished. Threads can be used instead # of processes by passing `worker_type=transfer_manager.THREAD`. # workers=8 from google.cloud.storage import Client, transfer_manager storage_client = Client() bucket = storage_client.bucket(bucket_name) results = transfer_manager.download_many_to_path( bucket, blob_names, destination_directory=destination_directory, max_workers=workers ) Additionally, could you also try enabling client-side tracing? This will provide us with more detailed insight that can help us pinpoint any potential bottlenecks on the client side. Here are the instructions on how to enable tracing in the python storage client: https://cloud.google.com/storage/docs/traces#python
Thanks for your reply, @cojenco !
I tried
download_many, although it differs a bit from my original plan since I was benchmarking downloads directly to memory.download_manyappears to only support downloads to disk.To minimize the impact of disk performance, I used a RAM disk as the destination. With 16 thread workers, I observed a throughput of around 2.7 Gbps, which didn’t show improvement compared to
download_as_bytes()orblob.open()thenread().transfer_manager.download_many_to_path( bucket, blob_names, destination_directory="/mnt/ramdisk", max_workers=16, worker_type=transfer_manager.THREAD )
As for process workers, I ran into issues with
daemonic processes are not allowed to have childrenon my current script. Since my focus is on comparing in downloading to memory and multi-threaded setups, I decided to defer testing with process workers for now.Another question is, looks like
blob.open()thenread()ultimately callsdownload_as_bytes()internally (source), but it consistently performs about 30% faster. I’m curious why there’s such a noticeable difference despite the shared code?I’ll try client-side tracing later, thanks for your suggestion. (Update: I tried to enable Cloud Trace but ran into the permission error:
Error while writing to Cloud Trace: 403 The caller does not have permission.I have granted the Cloud Trace Agent role to my service account.. will keep investigating)Hi @cojenco, just checking in to see if there are any updates or if you need any help reproducing my results. I’ve uploaded the benchmarking script here:https://github.com/dreamtalen/boto3-benchmark/tree/main/google-cloud-storage. Thanks!
FYI, running the same download test with
gcloud storage cpcan get 3.5 GiB/s (28 Gbps) throughput. However, gcloud test isn't a direct apples-to-apples comparison, since I can't choose threads/processes to configure the concurrency the same as my benchmark. But I hope this baseline is a helpful diagnostic step.I also did benchmarks across various concurrency settings show that the google-cloud-storage Python client currently falls behind boto3 across different multithreading and multiprocessing levels. Please find the image below:
Hi @dreamtalen I am trying to replicate the issue, using your scripts.
Hi @dreamtalen I am trying to replicate the issue, using your scripts.
We were able to replicate the issue, currently investigating the root cause and possible fixes.
Hi @dreamtalen I am trying to replicate the issue, using your scripts.
We were able to replicate the issue, currently investigating the root cause and possible fixes.
Hi @dreamtalen, We were able to find out the root cause of the issue. The primary reason for the performance disparity is found as the default data transfer strategy,
-
By default, GCS Python SDK iterates over small chunks (default size as 8 KB) of the response object, which introduces significant overhead, generating lower speed
-
By default, Boto3 Python SDK utilizes a single-shot read (response.raw.read()) of the response object, which downloads the entire 64 MB object at once causing better speed (while boto3 too supports chunk download, but it is not their default behavior)
Apart from this, we found some little overhead due to the checksum matching logic,
-
By default, GCS Python SDK does checksum matching
-
While Boto3 Python SDK by default don’t do checksum matching here, as shown in the customer script
We will be introducing the single shot read/download soon. While we recommend to keep the checksum matching enabled, since it is important for the data integrity. But soon, we will give an option to disable the checksum matching through a boolean, to give a better control.
Reacted by Yongming Ding and Joshua SleeperReacted by Yongming Ding and Joshua SleeperReacted by Joshua Sleeper-
Hi @dreamtalen I am trying to replicate the issue, using your scripts.
We were able to replicate the issue, currently investigating the root cause and possible fixes.
Hi @dreamtalen, We were able to find out the root cause of the issue. The primary reason for the performance disparity is found as the default data transfer strategy,
- By default, GCS Python SDK iterates over small chunks (default size as 8 KB) of the response object, which introduces significant overhead, generating lower speed
- By default, Boto3 Python SDK utilizes a single-shot read (response.raw.read()) of the response object, which downloads the entire 64 MB object at once causing better speed (while boto3 too supports chunk download, but it is not their default behavior)
Apart from this, we found some little overhead due to the checksum matching logic,
- By default, GCS Python SDK does checksum matching
- While Boto3 Python SDK by default don’t do checksum matching here, as shown in the customer script
We will be introducing the single shot read/download soon. While we recommend to keep the checksum matching enabled, since it is important for the data integrity. But soon, we will give an option to disable the checksum matching through a boolean, to give a better control.
Hi @shubham-up-47, that makes sense, and thanks for the detailed explanation!. Looking forward to trying it out once the single-shot version is out.
Reacted by shubham-up-47, Chandra Shekhar Sirimala and Joshua SleeperHi @dreamtalen, Some updates over the issue,
- We have raised one PR feat: Adding support of single shot download #1493 to support the single_shot_download in download_as_bytes
- You can pass checksum=None in it to disable checksumming, you will get additional speed by this (blob.open is faster because checksum=None is passed in it by default)
Reacted by Chandra Shekhar SirimalaHi @dreamtalen, Some updates over the issue,
- We have raised one PR feat: Adding support of single shot download #1493 to support the single_shot_download in download_as_bytes
- You can pass checksum=None in it to disable checksumming, you will get additional speed by this (blob.open is faster because checksum=None is passed in it by default)
Just a caution, if checksum verification is disabled, corrupted data may go undetected.
@dreamtalen I see you've done benchmarking for
google-cloud-storageversion: 2.19.0 & 3.1.0There's a difference between both of these versions as far as the checksum verification is concerned. A snippet from official documentation.
In Python Storage 3.0, uploads and downloads now have a default of “auto” where applicable. “Auto” will use crc32c checksums, except for unusual cases where the fast (C extension) crc32c implementation is not available, in which case it will use md5 instead. Before Python Storage 3.0, the default was md5 for most downloads and None for most uploads
you can know more by going through Checksum Defaults
Hi @dreamtalen,
- the feature is merged now feat: Adding support of single shot download #1493
- and we made a new release chore(main): release 3.2.0 #1496
Now, you should be able to get high speed by enabling flags
single_shot_download=Trueandraw_download=Truein the API call download_as_bytesin the latest clone (you can markpython-storage/google/cloud/storage/blob.py
Line 1393 in 7e0412a
def download_as_bytes( checksum=Noneif you want to disable the checksumming and make it little more faster).Let us know, whether this fix works well or not. Feel free to close the issue if its giving the expected speed.
Reacted by Yongming DingHi @shubham-up-47,
Thanks! I’ve verified that setting
single_shot_download=Truedelivers similar performance to boto3.Just curious - what exactly does
raw_download=Truedo? I tried toggling it on and off, but didn’t notice any significant difference. Is it necessary in this case?Reacted by shubham-up-47 and Chandra Shekhar SirimalaGlad to hear that setting
single_shot_download=Trueworked!You didn't see a difference because
raw_download=Trueonly matters in case of compressed objects. Imagine you have a file namedmy_data.txt.gzin your GCS bucket and it's stored in GCS withContent-Encoding: gzip.- raw_download=False (default): Using the default download_as_bytes(), the SDK will automatically decompress the my_data.txt.gz data for you, so you will receive the decompressed content of the file
my_data.txt. - raw_download=True: When you use download_as_bytes(raw_download=True), you'll get the raw bytes of the
my_data.txt.gzfile, which are the compressed bytes. You'll need to use gzip.decompress to get the original content.
Reacted by Yongming Ding- raw_download=False (default): Using the default download_as_bytes(), the SDK will automatically decompress the my_data.txt.gz data for you, so you will receive the decompressed content of the file
Environment details
google-cloud-storageversion: 2.19.0 & 3.1.0Steps to reproduce
When comparing download performance between Google Cloud Storage Python Client and boto3 (AWS SDK), we observed that GCS client is significantly slower (about 50% slower) than boto3 for downloading the same objects stored in a GCS bucket.
GCS Client Implementation (Two methods tested)
blob.download_as_bytes():blob.open()(10%-30% faster than method 1, but still 50% slower than boto3):boto3 Implementation
Performance Results
With
ThreadPoolExecutor(max_workers=16), I got following average throughput downloading 64MB x 1000 objects from GCS bucket to memory:get_object: 12 Gbpsdownload_as_bytes(): 3.2 Gbps in 2.19.0 & 4.2 Gbps in 3.1.0blob.open(): 4.5 Gbps in both 2.19.0 & 3.1.0Questions
Additional Context
raw_download=Trueblob.open("rb", chunk_size=xxx)Benchmarking scripts are available at https://github.com/dreamtalen/boto3-benchmark/tree/main/google-cloud-storage
Thanks!