Repository navigation
Performance gap between OCI Python SDK and boto3 for object downloads #755
Description
Activity
- addedSDKIssue pertains to the SDK itself and not specific to any serviceIssue pertains to the SDK itself and not specific to any service
on Apr 28, 2025 Hi @richachugh11 , thanks for taking a look and helping apply the correct label!
Just checking in to see if there’s any update on this issue or if I can provide any additional details to help move it forward. Thanks!@dreamtalen
There could be multiple factors why the performance is relatively slower. Depending upon the network performance, chunk-size, the download speed could vary. For an efficient download, we suggest users to use the Python SDK Download manager. Can you please check the below example and use the download manager to compare the results :https://github.com/oracle/oci-python-sdk/blob/master/examples/download_manager_example.py
Hi @y-chandra,
Thanks for the reply! I’ll give the Download Manager a try.
In the meantime, I have reported a similar issue to the GCS Python SDK. The root cause, as noted here, is that the GCS SDK reads the response in small chunks (default 8 KB), which introduces significant overhead and results in lower download speeds. In contrast, boto3 can download an entire 64 MB object in a single shot, leading to much better performance.
I’m curious how the OCI Python SDK handles this. Does it support single-shot downloads like boto3? If not, I suspect the same performance limitation may apply.
I've given it a try on Download Manager, and I'm seeing the same performance results as with
get_object. It appears the Download Manager, withallow_multipart=Falsefor smaller objects (like my 64MB test file, mirroring my boto3 comparison), acts as a simple wrapper aroundget_object(https://github.com/oracle/oci-python-sdk/blob/master/src/oci/object_storage/transfer/internal/download/DownloadManager.py#L193), leading to expectedly similar performance.Here's the code snippet I used:
download_configuration = DownloadConfiguration(allow_multipart=False, max_retries=1, backoff_type=BACKOFF_FULL_JITTER_VALUE) download_manager = DownloadManager(download_configuration, oci_client, DownloadState.DOWNLOADING) response = download_manager.get_object(namespace, self.bucket_name, key) return response.data.contentOn a separate note, I also encountered an
AttributeError: 'tuple' object has no attribute 'data'when trying to useget_object_to_byteswith the Download Manager. I've included the traceback below, but we can discuss this in a separate issue if that's easier:File "/lustre/fs12/portfolios/general/projects/general_cs_infra/users/yongmingd/benchmark-vs-boto3/oci_benchmark.py", line 213, in task self.download_object(oci_client, namespace, object_key, metrics) File "/lustre/fs12/portfolios/general/projects/general_cs_infra/users/yongmingd/benchmark-vs-boto3/oci_benchmark.py", line 183, in download_object download_manager.get_object_to_bytes(namespace, self.bucket_name, key) File "/home/yongmingd/.local/lib/python3.9/site-packages/oci/object_storage/transfer/internal/download/DownloadManager.py", line 266, in get_object_to_bytes response.data = stream.getvalue() AttributeError: 'tuple' object has no attribute 'data'Given these results, I still suspect the core issue lies with the chunk-size/single-shot download mechanism I mentioned in the previous comment.
@y-chandra gentle reminder - can you please to the above comment?
We are currently investigating this issue. We will update the issue when we have a resolution
Reacted by Yongming Ding and Paulius ŠarkaReacted by Yongming DingHi @y-chandra, do we have an update on this issue?
Environment details
ociversion: 2.111.0Issue
We are comparing the download performance of the OCI Python SDK and boto3 (AWS SDK). For the same objects stored in an OCI bucket, we’ve observed that the OCI SDK is approximately 20% to 50% slower than boto3 when downloading to memory.
Methods Tested with OCI SDK
response.data.content:Get idea from this issue, this method is ~60% faster than method 1 but still ~20% slower than boto3:
Note: We tested various chunk sizes, but they did not yield further improvements.
boto3 Baseline Implementation
Performance Results
With
ThreadPoolExecutor(max_workers=16), I got following average throughput downloading 64MB x 1000 objects from the same OCI bucket to memory:get_object: 9.8 Gbpsresponse.data.content: 4.1 Gbpsresponse.data.raw.stream: 6.8 GbpsThe gap remains consistent across multiple test runs, including various multithreaded and multiprocessed setups.
Questions
Thanks!