Skip to content

Repository files navigation

GeoTessera

GeoTessera is a Python library for reading, exporting, and displaying Tessera embeddings.

🚀 TESSERA v2 is here

The TESSERA v2 code is now live — model weights and inference code are available in the ucam-eo/tessera repository. v2 is our next-generation pixel-wise Earth foundation model; see the preprint, TESSERA v2: Scaling Pixel-wise Earth Foundation Models.

Want to try v2 embeddings early? You can pre-request v2 embeddings for your region and become an early tester:

⚠️ Heads-up: we are still ramping up the compute, storage, and release infrastructure for v2, so v2 embeddings will be produced slowly at first and there is no guaranteed turnaround time. If you need embeddings now, use v1.1 instead — wall-to-wall global coverage for 2017–2025 is available today.

Overview

GeoTessera provides access to geospatial embeddings from the Tessera foundation model, which processes Sentinel-1 and Sentinel-2 satellite imagery to generate 128-channel representation maps at 10m resolution. These embeddings compress a full year of temporal-spectral features into dense representations optimized for downstream geospatial analysis tasks. Read more details about the model.

The Zarr API reads selected points and regions from the cloud store and returns dequantized values on their native UTM grid. The tile API downloads individual NPY or GeoTIFF files for offline use; NPY tiles are deprecated and will be removed.

Stream regions to GeoTIFF or a web map

download reads from Zarr by default and writes one GeoTIFF per UTM zone. webmap reads three embedding bands and creates an RGB web map.

geotessera download --bbox '-3.0,53.4,-2.9,53.5' --year 2024 --output region/
geotessera webmap --bbox '-3.0,53.4,-2.9,53.5' --year 2024 --output map/ --serve

GeoTIFFs contain float32 embeddings on the native grid, with NaN nodata. Use --bands to select zero-based bands, --depth to select a published embedding prefix, or --store-url to read another Zarr store. Country and vector regions select their bounding boxes.

Rerunning a Zarr export replaces its completed files. Use --source tiles for individual tiles that skip existing files on rerun. --format npy also selects individual tiles.

Every download records its dataset version and variant in tessera_metadata.json in the output directory, and GeoTIFFs record them in their tags. Downloading into a directory that holds another dataset, or merging GeoTIFFs of different datasets, fails.

Web maps reuse matching completed mosaics and tiles. Keep the output directory and its JSON completion files. Use --force to refresh a map when the source changes at the same URL, or geotessera serve map/ to view an existing map without regenerating it.

See the CLI reference for options, output files, and restart behavior.

Global coverage

Since GeoTessera 0.11, the default v1.1-dclimate dataset provides wall-to-wall global v1.1 embeddings for every year from 2017 to 2025. No request is needed; GeoTesseraZarr() and geotessera download read any land area directly.

We recommend v1.1 for all users. Earlier versions remain selectable to reproduce existing results.

If you find missing or defective embeddings, please open an issue with the bounding box (lon,lat, 4 decimals) and year(s).

Table of Contents

Installation

GeoTessera requires Python 3.12 or later.

pip install geotessera

For development:

git clone https://github.com/ucam-eo/geotessera
cd geotessera
uv sync --all-extras --dev   # or: pip install -e ".[docs]"

Sphinx lives in the docs extra and the test harness in the dev group, so a plain pip install geotessera pulls in neither.

Cloud-Native Zarr Access

The Zarr API reads selected pixels from the public store and returns dequantized float32 embeddings. Point and region reads use the native UTM grid. Patches can combine pixels across zone boundaries.

from geotessera import GeoTesseraZarr

gt = GeoTesseraZarr()  # v1.1 dClimate Icechunk store
print(gt.years)  # [2017, ..., 2025]

# One embedding, with a status explaining any missing value
vec, status = gt.probe(0.12, 52.20, year=2024)
print(status)  # 'valid', 'water', 'nodata', or 'outside'

# Many points, one bulk read per UTM zone
X = gt.sample_points([(0.12, 52.20), (-2.97, 53.44)], year=2024)  # (2, 128)

# A lon/lat bounding box as a mosaic on the native UTM grid
bbox = (0.05, 52.15, 0.20, 52.25)
mosaic, transform, crs = gt.read_region(bbox, year=2024)

# A fixed-size patch centred on a point, merged across UTM zones when needed
patch, transform, crs = gt.read_patch(0.12, 52.20, year=2024, size_px=256)

# Stream a large region in row strips rather than holding it in memory
for block, transform, crs in gt.iter_region(bbox, year=2024, strip_rows=512):
    predictions = model.predict(block.reshape(-1, 128))

Export a region without holding the full array in memory:

files = gt.export_geotiffs(bbox, 2024, "region/", bands=[0, 1, 2])

gt.dataset names the published dataset being read, such as 1.1-dclimate. Exports record it in their tags and in the directory's tessera_metadata.json, and exporting into a directory that holds another dataset fails.

Other dataset versions are selected by store URL. v2 stores also publish matryoshka prefixes. depth=16 reads the first 16 dimensions; the bytes transferred depend on the store's chunk layout:

from geotessera.registry import zarr_store_url

gt = GeoTesseraZarr(zarr_store_url("v2"))
X16 = gt.sample_points([(0.12, 52.20)], year=2024, depth=16)  # (N, 16)

A location ending in .icechunk opens an Icechunk repository. Its UTM zone and hemisphere groups read as one utmNN zone on the northern CRS, with negative northings south of the equator; NaN scales mark unembedded pixels, reported as nodata.

Set cache_dir to persist Zarr metadata between runs. Byte-range reads of sharded embeddings are cached within the process. Icechunk stores, including the default, ignore it.

gt = GeoTesseraZarr(zarr_store_url("v2"), cache_dir="tessera-cache")

For direct access to one UTM zone, gt.open_zone(lon=0.15) returns an xarray dataset with a .tessera accessor that works in that zone's own eastings and northings. The store layout follows the geoemb: convention for geospatial embedding data; the zarr quickstart and the examples repository walk through complete workflows.

Agent Skill

The repository ships an Agent Skill that teaches coding agents the zarr interface. Claude Code users can install it as a plugin:

/plugin marketplace add ucam-eo/geotessera
/plugin install geotessera@ucam-eo

The skill itself is skills/geotessera/SKILL.md, in the open Agent Skills format, so it also works with other SKILL.md-compatible agents when copied into their skills directory.

Architecture

Core Concepts

This section describes the tile download interface; the zarr backend above streams the same embeddings without downloading files. The tile workflow has two steps:

  1. Retrieve embeddings: Fetch raw numpy arrays for a geographic bounding box
  2. Export to desired format: Save as raw numpy arrays or convert to georeferenced GeoTIFF files

Coordinate System and Tile Grid

The Tessera embeddings use a 0.1-degree grid system:

  • Tile size: Each tile covers 0.1° × 0.1° (approximately 11km × 11km at the equator)
  • Tile naming: Tiles are named by their center coordinates (e.g., grid_0.15_52.05)
  • Tile bounds: A tile at center (lon, lat) covers:
    • Longitude: [lon - 0.05°, lon + 0.05°]
    • Latitude: [lat - 0.05°, lat + 0.05°]
  • Resolution: 10m per pixel (variable number of pixels per tile depending on latitude)

File Structure and Downloads

When you request embeddings, GeoTessera downloads files over HTTPS from the public Source Cooperative repository into the output directory you specify, where they persist for re-use:

Embedding Files (via fetch_embedding)

  1. Quantized embeddings (grid_X.XX_Y.YY.npy):

    • Shape: (height, width, 128)
    • Data type: int8 (quantized for storage efficiency)
    • Contains the compressed embedding values
  2. Scale files (grid_X.XX_Y.YY_scales.npy):

    • Shape: (height, width) or (height, width, 128)
    • Data type: float32
    • Contains scale factors for dequantization
  3. Use geotessera.dequantize_embedding(quantized_embedding, scales) to dequantize arrays. The helper handles scale broadcasting and missing values.

  4. Persistent Storage: Files are downloaded into your chosen output directory and skipped on rerun, so interrupted downloads resume cleanly

Landmask Files (for GeoTIFF export)

When exporting to GeoTIFF, additional landmask files are fetched:

  • Landmask tiles (grid_X.XX_Y.YY.tiff):
    • Provide UTM projection information
    • Define precise geospatial transforms
    • Contain land/water masks
    • Cached alongside the embedding tiles for re-use

Data Flow

User Request (lat/lon bbox)
    ↓
Parquet Registry Lookup (find available tiles from manifest.parquet)
    ↓
HTTPS Downloads from Source Cooperative to Output Directory (integrity verified)
    ├── embedding.npy (quantized) → output dir
    └── embedding_scales.npy → output dir
    ↓
Dequantization (multiply arrays)
    ↓
Output Format
    ├── NumPy arrays → Direct analysis
    └── GeoTIFF → GIS integration

Storage Note: Only the Parquet manifests (tens to a couple of hundred MB per dataset) are cached under ~/.cache/geotessera. Embedding tiles are downloaded on demand into the output directory you specify and persist there for re-use across runs.

Quick Start

Check Available Data

Before downloading, check what data is available:

# Map the default v1.1 dclimate dataset from its tile registry
geotessera coverage --output coverage_map.png

# Map it for the UK, or for one year
geotessera coverage --country uk
geotessera coverage --year 2024 --output coverage_2024.png

# Map NPY tile coverage, with coverage.json and globe.html
geotessera coverage --dataset-variant cambridge --output tiles_coverage.png

Download Embeddings

Download individual tiles as NumPy arrays or GeoTIFF files. Tiles are deprecated and will be removed. v1.1 tiles exist only for the cambridge variant, which these commands use with a warning; add --dataset-variant cambridge to select it explicitly.

# Download as GeoTIFF (default, with georeferencing)
geotessera download --source tiles \
  --bbox "-0.2,51.4,0.1,51.6" \
  --year 2024 \
  --output ./london_tiffs

# Download quantized arrays with scales and landmasks.
geotessera download --source tiles \
  --bbox "-0.2,51.4,0.1,51.6" \
  --format npy \
  --year 2024 \
  --output ./london_arrays

# Download using a GeoJSON/Shapefile region
geotessera download --source tiles \
  --region-file cambridge.geojson \
  --format tiff \
  --year 2024 \
  --output ./cambridge_tiles

# Download specific bands only
geotessera download --source tiles \
  --bbox "-0.2,51.4,0.1,51.6" \
  --bands "0,1,2" \
  --year 2024 \
  --output ./london_rgb

Create Visualizations

Generate PCA visualizations and web maps from downloaded GeoTIFFs:

# Create a PCA mosaic from downloaded tiles
geotessera visualize ./london_tiffs pca_mosaic.tif

# Use histogram equalization for maximum contrast
geotessera visualize ./london_tiffs pca_balanced.tif --balance histogram

# Create web tiles and serve interactively
geotessera webmap pca_mosaic.tif --serve

# Serve existing web visualizations locally
geotessera serve ./london_web --open

Python API

Core Methods

The GeoTessera class downloads embedding tiles as files; for streaming access use the zarr backend. NPY tiles are deprecated and will be removed. v1.1 tiles exist only for the cambridge variant, which GeoTessera() uses with a warning; pass dataset_variant="cambridge" to select it explicitly. The tile interface provides two main methods for retrieving embeddings:

from geotessera import GeoTessera

# Initialize the client
gt = GeoTessera()

# Method 1: Fetch a single tile
embedding, crs, transform = gt.fetch_embedding(lon=0.15, lat=52.05, year=2024)
print(f"Shape: {embedding.shape}")  # e.g., (1200, 1200, 128)
print(f"CRS: {crs}")  # Coordinate reference system from landmask

# Method 2: Fetch all tiles in a bounding box
bbox = (-0.2, 51.4, 0.1, 51.6)  # (min_lon, min_lat, max_lon, max_lat)
tiles_to_fetch = gt.registry.load_blocks_for_region(bounds=bbox, year=2024)
embeddings = gt.fetch_embeddings(tiles_to_fetch)

for year, tile_lon, tile_lat, embedding_array, crs, transform in embeddings:
    print(f"Tile ({tile_lat}, {tile_lon}): {embedding_array.shape}")

Export Formats

Export as GeoTIFF

# Export embeddings for a region as individual GeoTIFF files
# Step 1: Get the tiles for the region
bbox = (-0.2, 51.4, 0.1, 51.6)
tiles_to_fetch = gt.registry.load_blocks_for_region(bounds=bbox, year=2024)

# Step 2: Export those tiles as GeoTIFFs
files = gt.export_embedding_geotiffs(
    tiles_to_fetch=tiles_to_fetch,
    output_dir="./output",
    bands=None,  # Export all 128 bands (default)
    compress="lzw"  # Compression method
)

print(f"Created {len(files)} GeoTIFF files")

# Export specific bands only (e.g., first 3 for RGB visualization)
files = gt.export_embedding_geotiffs(
    tiles_to_fetch=tiles_to_fetch,
    output_dir="./rgb_output",
    bands=[0, 1, 2]  # Only export first 3 bands
)

Work with NumPy Arrays

# Fetch and process embeddings directly
tiles_to_fetch = gt.registry.load_blocks_for_region(bounds=bbox, year=2024)
embeddings = gt.fetch_embeddings(tiles_to_fetch)

for year, tile_lon, tile_lat, embedding, crs, transform in embeddings:
    # Compute statistics
    mean_values = np.mean(embedding, axis=(0, 1))  # Mean per channel
    std_values = np.std(embedding, axis=(0, 1))    # Std per channel

    # Extract specific pixels
    center_pixel = embedding[embedding.shape[0]//2, embedding.shape[1]//2, :]

    # Apply custom processing
    processed = your_analysis_function(embedding)

Visualization Functions

from geotessera.visualization import (
    create_rgb_mosaic,
    visualize_global_coverage
)
from geotessera.web import (
    create_coverage_summary_map,
    geotiff_to_web_tiles
)

# Create an RGB mosaic from multiple GeoTIFF files
create_rgb_mosaic(
    geotiff_paths=["tile1.tif", "tile2.tif"],
    output_path="mosaic.tif",
    bands=(0, 1, 2)  # RGB bands
)

# Generate web tiles for interactive maps
geotiff_to_web_tiles(
    geotiff_path="mosaic.tif",
    output_dir="./web_tiles",
    zoom_levels=(8, 15)
)

# Create a global coverage visualization
visualize_global_coverage(
    tessera_client=gt,
    output_path="global_coverage.png",
    year=2024,  # Or None for all years
    width_pixels=2000,
    tile_color="red",
    tile_alpha=0.6
)

CLI Reference

Use geotessera COMMAND --help for command options. The command reference describes output formats, region selection, caching, and repeated runs.

download

Export a region from Zarr as GeoTIFFs. Add --source tiles for individual GeoTIFF tiles, or --format npy for quantized arrays with scales and landmasks.

geotessera download --bbox '-3.0,53.4,-2.9,53.5' --bands 0,1,2 --output region/

visualize

Create a PCA mosaic from a GeoTIFF file or a directory of GeoTIFF or NPY tiles. One sampled PCA model and color scale are applied across the inputs. The output contains display-scaled uint8 components; use the original embeddings for analysis.

geotessera visualize region/ pca.tif

webmap

Create web tiles and a viewer from an RGB GeoTIFF or a Zarr region. Tile generation requires the GDAL command-line tools.

geotessera webmap pca.tif --output map/ --serve
geotessera webmap --bbox '-3.0,53.4,-2.9,53.5' --output region_map/

coverage

Show data availability as a PNG map and HTML globe. Icechunk datasets, including the default, are drawn from their tile registry as a PNG map alone. Use --by-source to compare the datasets with NPY tiles.

geotessera coverage --country 'United Kingdom' --year 2024

serve

Serve an existing map directory over HTTP. An occupied port causes an error.

geotessera serve map/ --port 8001 --html viewer.html

info

List every dataset, its formats and store URLs, and summarise the selected dataset; or inspect local GeoTIFF and NPY files.

geotessera info
geotessera info --tiles region/

Registry System

Overview

GeoTessera uses a Parquet-based registry system to efficiently manage and access the large Tessera dataset:

  • Per-version manifests: Each dataset version has its own manifest.parquet listing every (year, lon, lat) tile available for that version's variants
  • Fast queries: Uses pandas DataFrames for efficient spatial and temporal filtering
  • Block-based organization: Internal 5×5 degree geographic blocks for efficient queries
  • Minimal storage: Only manifest files (tens to a couple of hundred MB per dataset) are cached locally
  • Integrity checking: Every download is verified against the response Content-Length, and against an MD5 computed over the streamed body whenever the server's ETag is a content MD5 (single-part uploads)
    • A mismatch rejects the download and triggers a retry, so corrupt or truncated files never reach the cache

Dataset Versions and Variants

Tessera embeddings are published as dataset versions (e.g. v1, v1.1, v2) and, within a version, as variants, each a separate inference run. Embeddings from different variants do not interoperate, even within one version: train and predict on the same (version, variant). Each variant is published in one or more formats:

Version Variant NPY tiles (npy/) Zarr (zarr/) Icechunk
1.0 vultr (default) v1/ v1/ —
1.1 dclimate (default) — v1.1-dclimate/ s3://tessera-embeddings/v1.1/dclimate.icechunk
1.1 cambridge v1.1-cam/ v1.1/ —
2.0 2B-L~beta1 (default) v2-2B-L~beta1/ v2-2B-L~beta1/ —
2.0 2B-L~beta2 v2-2B-L~beta2/ v2-2B-L~beta2/ —

The default version is v1.1. Streamed reads (GeoTesseraZarr(), download, webmap) use the version's default variant, dclimate, read from its Zarr store on Source Cooperative; the same run is also published as an Icechunk store. NPY tiles of v1.1 exist only for cambridge, so GeoTessera and download --format npy fall back to it with a warning; pass --dataset-variant cambridge to select it explicitly. coverage and info read the Icechunk store's Parquet tile registry in place of an NPY manifest. The table is DATASETS in geotessera/registry.py.

NPY tiles are deprecated and will be removed; use Zarr or Icechunk. List the datasets at any time with geotessera info, and select them on the CLI with --dataset-version and --dataset-variant, or in Python:

gt = GeoTessera(dataset_version="v1.1", dataset_variant="cambridge")

Use geotessera coverage --by-source to render each (version, variant) source in a distinct colour on the coverage map and globe viewer.

Registry Sources

The registry can be loaded from multiple sources (in priority order):

  1. Local file (via registry_path parameter)
  2. Local directory (via --registry-dir or registry_dir parameter, looks for manifest.parquet, falling back to the legacy registry.parquet)
  3. Remote URL (via registry_url parameter)
  4. Default remote (from https://data.source.coop/tessera/tessera/npy/{dataset}/manifest.parquet, where {dataset} encodes the (version, variant) pair: v1, v1.1-cam, v2-2B-L~beta1)
# Use local manifest file
gt = GeoTessera(registry_path="/path/to/manifest.parquet")

# Use local registry directory
gt = GeoTessera(registry_dir="/path/to/registry-dir")

# Use default remote manifest (downloads and caches automatically)
gt = GeoTessera()  # Default behavior

Registry Structure

The Parquet manifest contains columns for:

  • Coordinates: lon, lat (tile center coordinates)
  • Year: year (data year, 2017-2025)
  • Size: file_size (file size in bytes for download planning)
# Example manifest query
import pandas as pd
manifest = pd.read_parquet("manifest.parquet")
print(manifest.head())

Regenerating Manifests (Maintainers)

The per-version manifests can be rebuilt at any time by scanning the public Source Cooperative repository itself — no local copy of the data is needed:

# Rescan every dataset and write one manifest per npy/ directory
# (./manifests/{v1,v1.1-cam,v2-2B-L~beta1}/manifest.parquet) plus a
# per-version landmasks.parquet (./manifests/{v1,v1.1,v2}/landmasks.parquet)
geotessera-registry s3scan s3://tessera/tessera/npy/ \
    --landmasks-uri s3://tessera/tessera/landmasks/ \
    --output ./manifests

# Or scope to a single dataset — the directory name encodes the
# (version, variant) pair, so no --variant flag is needed
geotessera-registry s3scan "s3://tessera/tessera/npy/v1.1-cam/" \
    --landmasks-uri s3://tessera/tessera/landmasks/v1.1/ \
    --output ./manifests

The scan uses anonymous S3 ListObjectsV2 calls against https://data.source.coop, so no credentials are required to regenerate. Uploading the results does require source.coop write access, and the two parquet files go to different trees (matching where clients fetch them — note the npy/ tree is keyed by dataset directory, the landmasks/ tree by plain version):

aws s3 cp manifests/v1.1-cam/manifest.parquet \
    s3://tessera/tessera/npy/v1.1-cam/manifest.parquet \
    --endpoint-url https://data.source.coop
aws s3 cp manifests/v1.1/landmasks.parquet \
    s3://tessera/tessera/landmasks/v1.1/landmasks.parquet \
    --endpoint-url https://data.source.coop

The s3scan summary panel prints these per-file upload commands for you. Transient listing failures (Cloudflare 503s, timeouts) are retried with exponential backoff; if a shard still fails after retries, the affected manifest is not written and the command exits non-zero so an incomplete manifest can never be uploaded. See the maintenance guide for the full workflow, including caching caveats.

How Registry Loading Works

  1. Load Parquet manifest → Download and cache the dataset's manifest (if not local)
  2. Request tiles for bbox → Query DataFrame for tiles in region
  3. Filter by year and variant → Select tiles matching the requested year/variant
  4. Find available tiles → Return list of matching tiles
  5. HTTPS download → Fetch tiles on demand from the Source Cooperative mirror into the output directory, with integrity checks
  6. Persist → Downloaded tiles stay in the output directory and are skipped on rerun

Data Organization

Tessera Data Structure

Remote Server (https://data.source.coop/tessera/tessera)
├── npy/                                       # NPY embeddings + scales
│   │                                          # (one dir per (version, variant) dataset)
│   ├── v1/                                    # 1.0 — all variants share this dir
│   │   ├── manifest.parquet                   # Per-dataset tile manifest
│   │   └── 2024/grid_0.15_52.05/grid_0.15_52.05{,_scales}.npy
│   ├── v1.1-cam/                              # 1.1 / cambridge
│   │   ├── manifest.parquet
│   │   └── 2024/grid_0.15_52.05/grid_0.15_52.05{,_scales}.npy
│   └── v2-2B-L~beta1/                         # 2.0 / 2B-L~beta1 (beta)
│       └── 2024/grid_0.15_52.05/grid_0.15_52.05{,_scales}.npy
├── landmasks/                                 # Landmask TIFFs (per version)
│   ├── v1/
│   │   ├── landmasks.parquet                  # Landmask manifest
│   │   └── grid_0.15_52.05.tiff               # Landmask with projection info
│   ├── v1.1/
│   │   ├── landmasks.parquet
│   │   └── grid_0.15_52.05.tiff
│   └── v2/
│       ├── landmasks.parquet
│       └── grid_0.15_52.05.tiff
└── zarr/                                      # Cloud-native zarr stores
    ├── v1/                                    # 1.0 / vultr: 60 UTM zone groups + RGB pyramid
    ├── v1.1/                                  # 1.1 / cambridge
    ├── v1.1-dclimate/                         # 1.1 / dclimate, the default
    ├── v2-2B-L~beta1/                         # v2 stores add matryoshka
    └── v2-2B-L~beta2/                         # prefix arrays (d4, d16)

Icechunk (s3://tessera-embeddings/v1.1)
├── dclimate.icechunk/                         # 1.1 / dclimate, one group per zone and hemisphere
└── dclimate.registry/parts/                   # Parquet tile registry

Local Cache Structure

~/.cache/geotessera/                 # Default cache location (manifests only)
├── v1/                              # 1.0 dataset dir + v1 landmask registry
│   ├── manifest.parquet
│   └── landmasks.parquet
├── v1.1-cam/                        # 1.1/cambridge dataset dir
│   └── manifest.parquet
├── v1.1/                            # v1.1 landmask registry
│   └── landmasks.parquet
├── v2-2B-L~beta1/                   # 2.0 beta dataset dir
│   └── manifest.parquet
└── v2/                              # v2 landmask registry
    └── landmasks.parquet

# Note: Embedding and landmask tiles are NOT stored here. They are downloaded
# into the output directory you specify and persist there for re-use.

Coordinate Reference Systems

  • Embeddings: Stored in simple arrays, referenced by center coordinates
  • GeoTIFF exports: Use UTM projection from corresponding landmask tiles
  • Web visualizations: Reprojected to Web Mercator (EPSG:3857)

Cache Configuration

Tile workflows cache manifests and write embedding and landmask files to the output directory for reuse. Zarr reads persist metadata in cache_dir and cache byte ranges within the process. Keep exported files or completed web map directories to reuse their data between runs.

Python API

from geotessera import GeoTessera

# Use custom cache directory for registry
gt = GeoTessera(cache_dir="/path/to/cache")

# Use default cache location (recommended)
gt = GeoTessera()

CLI

# Specify custom cache directory
geotessera download --source tiles --cache-dir /path/to/cache ...

# Use default cache location
geotessera download --source tiles ...

Default Cache Locations

When cache_dir is not specified, the registry is cached in platform-appropriate locations:

  • Linux/macOS: $XDG_CACHE_HOME/geotessera or ~/.cache/geotessera
  • Windows: %LOCALAPPDATA%/geotessera

Hash Verification

GeoTessera verifies every downloaded file (embeddings, scales, and landmasks) against the response Content-Length, and additionally against an MD5 computed over the streamed body whenever the server's ETag is a content MD5 (a single-part upload; this covers landmask TIFFs and scales files). Large multipart-uploaded embedding tiles carry a composite ETag that is not a content hash, so they are length-checked only. A mismatch rejects the download and triggers a retry with backoff, so corrupt or truncated files never reach the cache.

Contributing

Contributions are welcome! Please see our Contributing Guide for details. This project is licensed under the MIT License - see the LICENSE file for details.

Citation

If you use Tessera in your research, please cite the arXiv paper:

@misc{feng2025tesseratemporalembeddingssurface,
      title={TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis}, 
      author={Zhengpeng Feng and Clement Atzberger and Sadiq Jaffer and Jovana Knezevic and Silja Sormunen and Robin Young and Madeline C Lisaius and Markus Immitzer and David A. Coomes and Anil Madhavapeddy and Andrew Blake and Srinivasan Keshav},
      year={2025},
      eprint={2506.20380},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2506.20380}, 
}

Links

Star History

Star History Chart

About

Python library for the Tessera embeddings

Topics

Resources

Stars

354 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages