Repository navigation
How to speed up reading of a large file from object store? #72
Description
Activity
@durgaswaroop I don't know of a way to speed up the process of downloading a file by changing settings in the SDK. I not aware of any cases where the SDK was the bottleneck when downloading a file. What region are you using and have you tried other regions?
Hi @durgaswaroop, your code should work. When in doubt, one of the best and easiest ways is to check what OCI CLI does (it is written in Python). As you can see here, it uses a chunk size of 1,048,576.
I ran a quick test with the following code:
#!/usr/bin/python import oci config = oci.config.from_file() onfig = oci.config.from_file() compartment_id = config["tenancy"] object_storage = oci.object_storage.ObjectStorageClient(config) namespace = object_storage.get_namespace().data bucket_name="MyBucket" object_name="160MB.img" obj = same_obj = object_storage.get_object(namespace, bucket_name, object_name) file=open("downloaded.img","w+") for chunk in obj.data.raw.stream(2048 ** 2, decode_content=False): file.write(chunk) file.close()Running this on a compute node in the same region, I could download a 160 MB image in less than 6 seconds:
$ time ./get.py real 0m5.874s user 0m0.813s sys 0m0.573s $ ls -l downloaded.img -rw-rw-r-- 1 itemir itemir 167772160 Oct 2 02:47 downloaded.img $So your download time should primarily depend on your available bandwidth. How much bandwidth do you have and what region do you try to connect to? One useful test would be to perform the same operation with OCI CLI and see if you see any significant difference between download times. You can find more information about OCI CLI at this link.
If you come to the conclusion that the issue is with the code, one way of speeding it up is to split the download to multiple parts. Our APIs (and SDKs) support ranged GETs. For example, you can get the first 1MB of an object with a code like this in Python:
obj = object_storage.get_object(namespace, bucket_name, object_name, range='bytes=0-1048575')Then you have the option to download multiple parts in multiple threads and combine them as a single file. Object Storage is a highly distributed system, so you can distribute it as much as you want. This is a much more advanced use-case though, so before going there, please verify your connection and available throughput to target region.
@durgaswaroop I'm closing this issue based on the response from @itemir. Please reopen the issue if you are still experiencing a problem.
Hi @durgaswaroop, your code should work. When in doubt, one of the best and easiest ways is to check what OCI CLI does (it is written in Python). As you can see here, it uses a chunk size of 1,048,576.
I ran a quick test with the following code:
#!/usr/bin/python import oci config = oci.config.from_file() onfig = oci.config.from_file() compartment_id = config["tenancy"] object_storage = oci.object_storage.ObjectStorageClient(config) namespace = object_storage.get_namespace().data bucket_name="MyBucket" object_name="160MB.img" obj = same_obj = object_storage.get_object(namespace, bucket_name, object_name) file=open("downloaded.img","w+") for chunk in obj.data.raw.stream(2048 ** 2, decode_content=False): file.write(chunk) file.close()Running this on a compute node in the same region, I could download a 160 MB image in less than 6 seconds:
$ time ./get.py real 0m5.874s user 0m0.813s sys 0m0.573s $ ls -l downloaded.img -rw-rw-r-- 1 itemir itemir 167772160 Oct 2 02:47 downloaded.img $So your download time should primarily depend on your available bandwidth. How much bandwidth do you have and what region do you try to connect to? One useful test would be to perform the same operation with OCI CLI and see if you see any significant difference between download times. You can find more information about OCI CLI at this link.
If you come to the conclusion that the issue is with the code, one way of speeding it up is to split the download to multiple parts. Our APIs (and SDKs) support ranged GETs. For example, you can get the first 1MB of an object with a code like this in Python:
obj = object_storage.get_object(namespace, bucket_name, object_name, range='bytes=0-1048575')Then you have the option to download multiple parts in multiple threads and combine them as a single file. Object Storage is a highly distributed system, so you can distribute it as much as you want. This is a much more advanced use-case though, so before going there, please verify your connection and available throughput to target region.
I'd like to do something similar, can you explain the " file=open("downloaded.img","w+")" portion of code a bit more? Are we downloading the get_object request to this file on server?
Sorry, for the format (didn't mean to quote reply in my comment). New to github.
I have a file of about 160 MB on Object store and it is taking almost 8 minutes to download that. What can I do to speed that up?
This is what I'm doing right now.
I tried playing the buffer size but that didn't seem to make any difference. What else can I do to download that file faster?