Repository navigation
Better HDFS Support - #1556
Merged
echeipesh merged 21 commits intoJul 5, 2016
Merged
Better HDFS Support#1556
Conversation
…ly, reducing memory pressure
| MapFile.Writer.valueClass(classOf[BytesWritable]), | ||
| MapFile.Writer.compression(SequenceFile.CompressionType.NONE)) | ||
| writer.setIndexInterval(1) | ||
| writer.setIndexInterval(32) |
Member
There was a problem hiding this comment.
how it would effect query time? (just curious)
Contributor
Author
There was a problem hiding this comment.
Layer query time would not be effected at all by this. When we're reading off ranges we're already seeking through the file, so not having as many points would have minimal impact if any. This is going to have more of an impact on random access through value reader. Spinning up a cluster to figure that part out. For reference the default interval value is 128
…s on attributeStore type
…ks in cache lookups
| /** | ||
| * When record being written would exceed the block size of the current MapFile | ||
| * opens a new file to continue writing. This allows to split partition into block-sized | ||
| * chunks without foreknowledge of how bit it is. |
Member
|
+1 after comments addressed |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR improves HDFS layer writing and random value reading support.
HadoopValuesReadernow opens upMapFile.Readersfor each map file in the layer. These readers cache the index of available keys so they are able to provide quick lookups. This replaces and improves previous method of usingFileInputFormatto query for a single record.HadoopRDDWriterhas multiple improvements:Current testing shows:
GroupedCumulativeIteratorapproach are able to complete in settings where .groupBy/sort/write jobs are killed for memory violations.Future improvements that on the mind:
HadoopValueReaderand to produce quicker lookup in most cases.HadoopValueReaderwhen fetching a record, cache all the tiles that share that index before filtering down to a single tile. These records are very likely to be asked for next and should be stored in LRU cache.