Skip to content

Performance gap between OCI Python SDK and boto3 for object downloads #755

Description

@dreamtalen

Environment details

  • Python version: 3.9.18
  • pip version: 23.2.1
  • oci version: 2.111.0

Issue

We are comparing the download performance of the OCI Python SDK and boto3 (AWS SDK). For the same objects stored in an OCI bucket, we’ve observed that the OCI SDK is approximately 20% to 50% slower than boto3 when downloading to memory.

Methods Tested with OCI SDK

  1. Using response.data.content :
response = self._oci_client.get_object(
    namespace_name=self._namespace, bucket_name=bucket, object_name=key, range=bytes_range
)
return response.data.content 
  1. Using response.data.raw.stream
    Get idea from this issue, this method is ~60% faster than method 1 but still ~20% slower than boto3:
response = self._oci_client.get_object(
    namespace_name=self._namespace, bucket_name=bucket, object_name=key, range=bytes_range
)
content = bytearray()
for chunk in response.data.raw.stream(1024 * 1024, decode_content=False):  # 1MB chunks
    content.extend(chunk)
return bytes(content)

Note: We tested various chunk sizes, but they did not yield further improvements.

boto3 Baseline Implementation

response = s3_client.get_object(Bucket=bucket_name, Key=key)
return response['Body'].read()

Performance Results

With ThreadPoolExecutor(max_workers=16), I got following average throughput downloading 64MB x 1000 objects from the same OCI bucket to memory:

  • boto3 get_object: 9.8 Gbps
  • OCI SDK response.data.content: 4.1 Gbps
  • OCI SDK response.data.raw.stream: 6.8 Gbps

The gap remains consistent across multiple test runs, including various multithreaded and multiprocessed setups.

Questions

  1. Is this performance gap expected?
  2. Are there any recommended optimizations or best practices for improving download performance with the OCI Python SDK?
  3. Are there any internal differences in how OCI supports S3-compatible APIs handling downloads that might explain the performance gap?

Thanks!

Activity

  1. added
    SDKIssue pertains to the SDK itself and not specific to any service
    on Apr 28, 2025
  2. dreamtalen commented on May 5, 2025

    @dreamtalen
    Author

    Hi @richachugh11 , thanks for taking a look and helping apply the correct label!
    Just checking in to see if there’s any update on this issue or if I can provide any additional details to help move it forward. Thanks!

  3. y-chandra commented on Jul 17, 2025

    @y-chandra
    Member

    @dreamtalen
    There could be multiple factors why the performance is relatively slower. Depending upon the network performance, chunk-size, the download speed could vary. For an efficient download, we suggest users to use the Python SDK Download manager. Can you please check the below example and use the download manager to compare the results :

    https://github.com/oracle/oci-python-sdk/blob/master/examples/download_manager_example.py

  4. dreamtalen commented on Jul 17, 2025

    @dreamtalen
    Author

    Hi @y-chandra,

    Thanks for the reply! I’ll give the Download Manager a try.

    In the meantime, I have reported a similar issue to the GCS Python SDK. The root cause, as noted here, is that the GCS SDK reads the response in small chunks (default 8 KB), which introduces significant overhead and results in lower download speeds. In contrast, boto3 can download an entire 64 MB object in a single shot, leading to much better performance.

    I’m curious how the OCI Python SDK handles this. Does it support single-shot downloads like boto3? If not, I suspect the same performance limitation may apply.

  5. dreamtalen commented on Jul 18, 2025

    @dreamtalen
    Author

    I've given it a try on Download Manager, and I'm seeing the same performance results as with get_object. It appears the Download Manager, with allow_multipart=False for smaller objects (like my 64MB test file, mirroring my boto3 comparison), acts as a simple wrapper around get_object(https://github.com/oracle/oci-python-sdk/blob/master/src/oci/object_storage/transfer/internal/download/DownloadManager.py#L193), leading to expectedly similar performance.

    Here's the code snippet I used:

    download_configuration = DownloadConfiguration(allow_multipart=False, max_retries=1, backoff_type=BACKOFF_FULL_JITTER_VALUE)
    download_manager = DownloadManager(download_configuration, oci_client, DownloadState.DOWNLOADING)
    response = download_manager.get_object(namespace, self.bucket_name, key)
    return response.data.content
    

    On a separate note, I also encountered an AttributeError: 'tuple' object has no attribute 'data' when trying to use get_object_to_bytes with the Download Manager. I've included the traceback below, but we can discuss this in a separate issue if that's easier:

    File "/lustre/fs12/portfolios/general/projects/general_cs_infra/users/yongmingd/benchmark-vs-boto3/oci_benchmark.py", line 213, in task
        self.download_object(oci_client, namespace, object_key, metrics)
    File "/lustre/fs12/portfolios/general/projects/general_cs_infra/users/yongmingd/benchmark-vs-boto3/oci_benchmark.py", line 183, in download_object
        download_manager.get_object_to_bytes(namespace, self.bucket_name, key)
    File "/home/yongmingd/.local/lib/python3.9/site-packages/oci/object_storage/transfer/internal/download/DownloadManager.py", line 266, in get_object_to_bytes
        response.data = stream.getvalue()
    AttributeError: 'tuple' object has no attribute 'data'
    

    Given these results, I still suspect the core issue lies with the chunk-size/single-shot download mechanism I mentioned in the previous comment.

  6. pradeep-mj commented on Jul 29, 2025

    @pradeep-mj

    @y-chandra gentle reminder - can you please to the above comment?

  7. y-chandra commented on Aug 5, 2025

    @y-chandra
    Member

    We are currently investigating this issue. We will update the issue when we have a resolution

  8. dreamtalen commented on Feb 2, 2026

    @dreamtalen
    Author

    Hi @y-chandra, do we have an update on this issue?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    SDKIssue pertains to the SDK itself and not specific to any service

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions