In order to produce object metadata and its overall size the underlying S3 client will issue a HEAD request:
|
val metadata: ObjectMetadata = |
|
client.getObjectMetadata(request.getBucketName, request.getKey) |
|
|
|
val totalLength: Long = metadata.getContentLength |
This is problematic for two reasons:
- Sometimes the object ACT will not allow
HEAD requests, resulting in error 400
- In most use cases
totalLength is not used and is not important.
To expand on the second point, the GeoTrellis GeoTiffReader will read from the start of the file until it consumes the full header. At that point it will use the offset information from the GeoTiff header to read the data segments.
Note that HttpRangeReader already has logic for this case:
|
val totalLength: Long = { |
|
val headers = if(useHeadRequest) { |
|
request.method("HEAD").asString |
|
} else { |
|
request.method("GET").execute { is => "" } |
|
} |
|
val contentLength = headers |
|
.header("Content-Length") |
|
.flatMap({ cl => Try(cl.toLong).toOption }) match { |
|
case Some(num) => num |
|
case None => -1L |
|
} |
|
|
|
/** |
|
* "The Accept-Ranges response HTTP header is a marker used by the server |
|
* to advertise its support of partial requests. The value of this field |
|
* indicates the unit that can be used to define a range." |
|
* https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Accept-Ranges |
|
*/ |
|
require(headers.header("Accept-Ranges") == Some("bytes"), |
|
"Server doesn't support ranged byte reads") |
|
|
|
require(contentLength > 0, |
|
"Server didn't provide (required) \"Content-Length\" headers, unable to do range-based read") |
|
|
|
contentLength |
|
} |
At a minimum the totalSize should be turned into lazy val which would avoid the issue in most cases. S3RangeReader should also have a constructor flag to avoid head requests. This should cause re-evaluation of the RangeReader interface, maybe we don't need that method at all?
In order to produce object metadata and its overall size the underlying S3 client will issue a
HEADrequest:geotrellis/s3/src/main/scala/geotrellis/spark/io/s3/util/S3RangeReader.scala
Lines 39 to 42 in f86ef9e
This is problematic for two reasons:
HEADrequests, resulting in error400totalLengthis not used and is not important.To expand on the second point, the GeoTrellis GeoTiffReader will read from the start of the file until it consumes the full header. At that point it will use the offset information from the GeoTiff header to read the data segments.
Note that
HttpRangeReaderalready has logic for this case:geotrellis/spark/src/main/scala/geotrellis/spark/io/http/util/HttpRangeReader.scala
Lines 38 to 64 in f86ef9e
At a minimum the
totalSizeshould be turned intolazy valwhich would avoid the issue in most cases. S3RangeReader should also have a constructor flag to avoid head requests. This should cause re-evaluation of theRangeReaderinterface, maybe we don't need that method at all?