Repository navigation
Vector Tiles - #1563
Vector Tiles#1563
Conversation
…types into the feature category
… future, added better filter funcitonality, added more test data and test data generating examples
…elevation into vector-tiles-decoder
Vector tiles decoder
…is into feature/vector-tiles
| case LineTo(ds) => unparseCmd(2, ds.length) +: params(ds) // (+:) is bad! | ||
| case ClosePath => Array(unparseCmd(7, 1)) | ||
| }) | ||
| } |
There was a problem hiding this comment.
@lossyrob flatMapping here is the clean thing, but very likely not the fast thing.
- This is to make for easier addition of backends
- These traits don't assume a backend, and should first be extended by classes like `ProtobufTile`, etc.
|
@lossyrob Something I need to get back into on Monday: We'll be storing RDDs of high-level |
|
Unless it's okay to Avro serialize: case class ProtobufLayer(
name: String,
extent: Int,
rawFeatures: Seq[vt.Tile.Feature] // from the source `vt.Tile.Layer`
) extends Layer { ... } // `Layer` is a trait from the parent packageand have these within RDDs. Then any geometry manipulation on them will be lazy and internal as usual, as if they had come straight from protobuf bytes. |
|
It turns out the |
- Using lazy Streams here allows us to avoid strictly holding the Stream head,
meaning no Features of a geomtype we don't care about will be parsed,
unless we ask.
- This is advantangeous for queries as well. If you are looking for a Feature
match on some metadata point, only Features will be parsed until you find
what you're looking for.
- Potential disadvantage being intermittent instances of the opposite
Single/Multi you're looking for will fully parse as well. The alternative
is the ad-hoc reimplementation of laziness with internal mutable data
structures.
Consider the following scenarios, where P and MP are Polygon and
MultiPolygon respectively:
[P MP P MP P MP] -- A list of alternating raw Polygon features.
(1) The user wants to find a particular Polygon, which unknown to them
is the second one in the list. They have to parse the first P,
do *something* to the first MP, parse the second P and match on it,
then stop.
With Streams, the original list now looks like: [ MP P MP ]
With custom laziness, it looks like: [ MP MP P MP ]
The custom laziness wins for speed here, since we were able to
cancel the parsing of the first MP early, and the Streams
approach fully parsed the first MP.
(2) The user wants to perform another operation, this time across all
Ps. Both approaches must thus map over the entire list.
With Streams, the original list is now empty: []
With custom laziness, it looks like: [ MP MP MP ]
It's harder to tell who wins here, since while the Streams
had to waste time fully parsing each MP, the custom laziness
had to reparse old MPs it had already looked through.
(3) The user wants to perform another operation on all Ps. The streams
can go ahead since everything has been parsed. The custom approach
must reparse all the MPs to check for Ps, since it wouldn't know
there weren't any left.
(4) The user wants to perform an operation on MPs this time. The streams
can go ahead since the MPs are already parsed. The custom approach
has to reparse the MPs for the fourth time.
My takeaway: the custom approach is better for one-off operations on a
particular geometry type. The stream approach quickly overtakes the other
if you plan multiple operations over the same geometries.
|
@lossyrob I did some more thinking on Streams vs custom laziness: Consider the following scenarios, where P and MP are Polygon and MultiPolygon respectively: [P MP P MP P MP] -- A list of alternating raw Polygon features. (1) The user wants to find a particular Polygon, which unknown to them (2) The user wants to perform another operation, this time across all (3) The user wants to perform another operation on all Ps. The streams (4) The user wants to perform an operation on MPs this time. The streams My takeaway: the custom approach is better for one-off operations on a |
|
A good shower made me realize your three-stage idea wins out, since we'd never have to reparse MPs until Step 4 above. @echeipesh warnings of over-engineering echo in my ear, mind you. |
| def fromPBTile( | ||
| tile: vt.Tile, | ||
| key: SpatialKey, | ||
| layout: LayoutDefinition |
There was a problem hiding this comment.
We actually just want to pass in the Extent. SpatialKey/LayoutDefinition are Sparky things, and vector tiles shouldn't depend on sparky things (but be used by sparky things in the layer versions of them)
There was a problem hiding this comment.
I've already made this change in the next PR.
- And so the only thing require we during IO is that Extent. This keeps the API simple.
- Renamed to avoid confusion with GeoTrellis `Extent`.
- This shaves off a bit more time from the vanilla decoding process
|
@lossyrob this is good to go. |
|
@lossyrob |
|
@fosskers just needs an update and we're good to go. |
|
@lossyrob updated to master. |
|
💯 |
TODO
scalapbGeometrytypesCommandconversiontraitsGeometryStreamsFeatureconstructiontoCommandsinstances)LayoutDefinitions on decodeLayoutDefinitions on encode.mvtfilesInvalidCommand, etc)Motivation
Invented by Mapbox, they are a combination of the ideas of finite-sized tiles and vector geometries. Mapbox maintains the official implementation spec for VectorTile codecs.
VectorTiles are advantageous over raster tiles in that:
Raw VectorTile data is stored in the protobuf format. Any codec implementing
the spec must decode and encode data according to this .proto schema.
Post-PR Next Steps
case classesLayoutDefinitionto not assume Rasters. (inner TileLayout does)VectorTileRDD