Skip to content

How to speed up reading of a large file from object store? #72

Description

@johnnydepup

I have a file of about 160 MB on Object store and it is taking almost 8 minutes to download that. What can I do to speed that up?
This is what I'm doing right now.

client = oci.object_storage.ObjectStorageClient(..)
obj = client.get_object(..., ...., ...)

for chunk in obj.data.raw.stream(2048 ** 2, decode_content=False):
    file.write(chunk)

I tried playing the buffer size but that didn't seem to make any difference. What else can I do to download that file faster?

Activity

  1. arthall commented on Sep 28, 2018

    @arthall
    Contributor

    @durgaswaroop I don't know of a way to speed up the process of downloading a file by changing settings in the SDK. I not aware of any cases where the SDK was the bottleneck when downloading a file. What region are you using and have you tried other regions?

  2. itemir commented on Oct 2, 2018

    @itemir

    Hi @durgaswaroop, your code should work. When in doubt, one of the best and easiest ways is to check what OCI CLI does (it is written in Python). As you can see here, it uses a chunk size of 1,048,576.

    I ran a quick test with the following code:

    #!/usr/bin/python
    import oci
    
    config = oci.config.from_file()
    
    onfig = oci.config.from_file()
    compartment_id = config["tenancy"]
    object_storage = oci.object_storage.ObjectStorageClient(config)
    
    namespace = object_storage.get_namespace().data
    
    bucket_name="MyBucket"
    object_name="160MB.img"
    
    obj = same_obj = object_storage.get_object(namespace,
                                               bucket_name,
                                               object_name)
    
    file=open("downloaded.img","w+")
    for chunk in obj.data.raw.stream(2048 ** 2, decode_content=False):
        file.write(chunk)
    file.close()
    

    Running this on a compute node in the same region, I could download a 160 MB image in less than 6 seconds:

    $ time ./get.py
    
    real	0m5.874s
    user	0m0.813s
    sys	0m0.573s
    $ ls -l downloaded.img
    -rw-rw-r-- 1 itemir itemir 167772160 Oct  2 02:47 downloaded.img
    $
    

    So your download time should primarily depend on your available bandwidth. How much bandwidth do you have and what region do you try to connect to? One useful test would be to perform the same operation with OCI CLI and see if you see any significant difference between download times. You can find more information about OCI CLI at this link.

    If you come to the conclusion that the issue is with the code, one way of speeding it up is to split the download to multiple parts. Our APIs (and SDKs) support ranged GETs. For example, you can get the first 1MB of an object with a code like this in Python:

        obj = object_storage.get_object(namespace,
                                                   bucket_name,
                                                   object_name,
    	    				       range='bytes=0-1048575')
    

    Then you have the option to download multiple parts in multiple threads and combine them as a single file. Object Storage is a highly distributed system, so you can distribute it as much as you want. This is a much more advanced use-case though, so before going there, please verify your connection and available throughput to target region.

  3. arthall commented on Oct 18, 2018

    @arthall
    Contributor

    @durgaswaroop I'm closing this issue based on the response from @itemir. Please reopen the issue if you are still experiencing a problem.

  4. bassg0navy commented on Feb 24, 2022

    @bassg0navy

    Hi @durgaswaroop, your code should work. When in doubt, one of the best and easiest ways is to check what OCI CLI does (it is written in Python). As you can see here, it uses a chunk size of 1,048,576.

    I ran a quick test with the following code:

    #!/usr/bin/python
    import oci
    
    config = oci.config.from_file()
    
    onfig = oci.config.from_file()
    compartment_id = config["tenancy"]
    object_storage = oci.object_storage.ObjectStorageClient(config)
    
    namespace = object_storage.get_namespace().data
    
    bucket_name="MyBucket"
    object_name="160MB.img"
    
    obj = same_obj = object_storage.get_object(namespace,
                                               bucket_name,
                                               object_name)
    
    file=open("downloaded.img","w+")
    for chunk in obj.data.raw.stream(2048 ** 2, decode_content=False):
        file.write(chunk)
    file.close()
    

    Running this on a compute node in the same region, I could download a 160 MB image in less than 6 seconds:

    $ time ./get.py
    
    real	0m5.874s
    user	0m0.813s
    sys	0m0.573s
    $ ls -l downloaded.img
    -rw-rw-r-- 1 itemir itemir 167772160 Oct  2 02:47 downloaded.img
    $
    

    So your download time should primarily depend on your available bandwidth. How much bandwidth do you have and what region do you try to connect to? One useful test would be to perform the same operation with OCI CLI and see if you see any significant difference between download times. You can find more information about OCI CLI at this link.

    If you come to the conclusion that the issue is with the code, one way of speeding it up is to split the download to multiple parts. Our APIs (and SDKs) support ranged GETs. For example, you can get the first 1MB of an object with a code like this in Python:

        obj = object_storage.get_object(namespace,
                                                   bucket_name,
                                                   object_name,
    	    				       range='bytes=0-1048575')
    

    Then you have the option to download multiple parts in multiple threads and combine them as a single file. Object Storage is a highly distributed system, so you can distribute it as much as you want. This is a much more advanced use-case though, so before going there, please verify your connection and available throughput to target region.

    I'd like to do something similar, can you explain the " file=open("downloaded.img","w+")" portion of code a bit more? Are we downloading the get_object request to this file on server?

  5. bassg0navy commented on Feb 24, 2022

    @bassg0navy

    Sorry, for the format (didn't mean to quote reply in my comment). New to github.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions