Skip to content

Repository files navigation

DisSCube

CI Docs License: MIT

Status: Alpha — stable APIs for the core pipeline; declarative models still evolving.

DisSCube is the spatial data cube engine of the DisSModel ecosystem. It converts raw geospatial sources (rasters, vectors) into derived variables aligned to LUCC (Land Use and Cover Change) modeling grids, ready for Cellular Automata models and spatio-temporal analysis.

📖 Documentation: https://dissmodel.github.io/disscube/

Core concept

SpatialSource  →  Derivation  →  Variable  →  DerivedVariable (Zarr)

A source (SpatialSource) goes through a derivation (SpatialDerivation or Derivation) that applies an operator on a grid (GridSpec), producing a derived variable registered in the SQLite catalog and stored in Zarr.

Installation

git clone https://github.com/DisSModel/disscube.git
cd disscube
python -m venv .venv && source .venv/bin/activate
pip install -e .

Basic usage

1. Initialize the catalog and register a grid

from disscube.client import CubeClient
from disscube.utils.grids import register_local_grid

cube = CubeClient(catalog="catalog.db", store="./data/")

grid = register_local_grid(
    cube,
    name="AC",
    bbox_geo=(-73.99, -11.15, -66.62, -7.11),
    resolution=5_000.0,
)

2. Register a source

from disscube.models import SpatialSource

cube.register_spatial_source(SpatialSource(
    id="mapbiomas_2020",
    name="MapBiomas Acre 2020",
    format="raster",
    asset_url="data/raw/mapbiomas_2020.tif",
    crs="EPSG:4326",
    time=2020,
))

3. Derive — declarative mode (recommended)

from disscube.derivation import Derivation

d = Derivation(
    target="forest_pct",
    source_id="mapbiomas_2020",
    operator="percentage",
    class_code=3,
    role="driver",
    valid_from="2020",
    valid_until="2020",
)

cube.derive_declarative(d, grid_id="AC/5km")

4. Derive — direct mode

from disscube.models import SpatialDerivation, Variable

cube.derive(SpatialDerivation(
    source_id="mapbiomas_2020",
    grid_id="AC/5km",
    role="driver",
    variables=[Variable(name="forest_pct", operator="percentage", class_code=3)],
    valid_from="2020",
    valid_until="2020",
))

5. Load the result

da = cube.load("forest_pct", grid_id="AC/5km")
print(da.shape)   # (rows, cols)

6. Get the cube out

ds = cube.to_dataset(["forest_pct", "dist_roads"], grid_id="AC/5km", period=("2015", "2020"))
# xarray.Dataset: (y, x) static and (time, y, x) temporal variables, CRS and transform via ds.rio

cube.export_geotiff(["forest_pct"], "forest.tif", grid_id="AC/5km")   # one band per variable and year
cube.export_netcdf(["forest_pct"], "cube.nc", grid_id="AC/5km")       # CF-1.8; pip install "disscube[netcdf]"

Exports carry their provenance: each GeoTIFF band and netCDF variable records the spec_hash, the content_hash of the stored data and the source_checksum of the input it came from (cube.provenance("forest_pct") lists them per year).

DisSCube does not need DisSModel. To hand a cube to a DisSModel model, install disscube[dissmodel] and use cube.to_raster_backend(...), which returns a RasterBackend.

Pipeline files (TOML)

A whole data preparation — grid, sources, derived variables — can be declared in one TOML file, the counterpart of a TerraME fill script, and run from the command line:

disscube validate examples/pipelines/quickstart.toml   # no downloads
disscube run      examples/pipelines/quickstart.toml --workspace outputs/quickstart
schema = 1
name = "Quickstart"

[grid]
name = "demo/300m"
crs = "EPSG:31983"
bbox = [570000.0, 9708000.0, 582000.0, 9720000.0]
resolution = 300

[[source]]
id = "landuse"
type = "file"
path = "../data/quickstart/landuse.tif"
crs = "EPSG:31983"

[[derive]]
target = "forest_pct"
source = "landuse"
operator = "percentage"
class_code = 3

See docs/guides/pipeline_files.md and examples/pipelines/.

Examples

examples/ has runnable, self-contained examples that run offline in seconds:

  • 01–03, synthetic data — a quickstart with raster operators, vector drivers, and time series handed off to DisSModel.
  • quickstart.toml — declarative pipeline equivalent to example 01.
python examples/01_quickstart.py
disscube run examples/pipelines/quickstart.toml

Real data workflows, TerraME parity benchmarks, and large-scale case studies are maintained in LambdaGeo/disscube-recipes (cases/terrame_fill, cases/ilha_maranhao, cases/prodes_br163, cases/luccme_br).

See examples/README.md for the full list, and docs/terrame_fill_correspondence.md for how DisSCube's operators relate to TerraME's Fill.

Available operators

Operator Type Resampling Requires class_code
mean zonal average no
sum zonal sum no
std zonal nearest no
min zonal min no
max zonal max no
majority zonal nearest¹ no
minority zonal nearest¹ no
percentage zonal nearest¹ yes
attribute zonal nearest no
presence zonal nearest no
distance proximity (exact, cell centre → nearest feature, source not clipped) — no
min_distance proximity (raster approximation, features inside the grid) nearest no
count proximity nearest no
area polygons (exact share of the cell covered) — no

distance takes params = {crs = …} to measure in another CRS (metres on a geographic grid); the ¹ operators take params = {subcells = n} to cap the fine pixels per cell. Any derivation takes fill = "nearest".

¹ These use needs_fine_alignment=True: GridAligner resamples with nearest at high resolution; the actual reduction (per-window counting) is done by the operator.

Pipeline

SpatialSource
    │
    ▼
Normalizer        — validates / loads a GeoDataFrame (vector) or opens the raster
    │
    ▼
GridAligner       — reprojects per variable with the operator's Resampling
    │
    ▼
Aggregator        — delegates to operator.compute() → one xr.DataArray per variable
    │
    ▼
VariableWriter    — writes Zarr + registers the DerivedVariable in the catalog

Storage layout

data/derived/{grid_id}/{partition}/{spec_hash}/{variable_name}.zarr
  • partition = tile_id, or global for untiled derivations.
  • spec_hash = SHA-256 of the derivation (source + grid + variables + time window, plus the source's checksum when it has one — replacing a source file and registering its new checksum recomputes instead of returning a stale product).

Project structure

disscube/
├── client.py         CubeClient — public entry point
├── models/           GridSpec, SpatialSource, SpatialDerivation, Variable, Derivation…
├── operators/        Operators as classes (self-registered via __init_subclass__)
│   ├── base.py       Operator ABC + OPERATOR_REGISTRY
│   ├── zonal.py      mean, sum, majority, percentage, attribute, presence…
│   └── proximity.py  distance, min_distance, count
├── pipeline/         Pipeline execution & planning (schema, runner) + internal stages
├── catalog/          CatalogStore (Protocol) + SQLite and JSON implementations
├── storage.py        AssetStore (fsspec — local and S3)
├── cli.py            `disscube validate` / `disscube run` / `disscube export`
├── sources/          Adapters that bring external data in as SpatialSources,
│   │                 each with a checksum and a provenance.json sidecar
│   ├── _raster.py    Window2D, windowed reads, composites, mosaics, register_raster
│   ├── bdc.py        Brazil Data Cube cubes via STAC (`bdc` extra)
│   ├── _categorical.py  legends (.qml/.json/.csv) and reclassification
│   ├── mapbiomas.py  MapBiomas annual land-cover maps (Collection 11, 10 m series)
│   ├── prodes.py     PRODES deforestation (download + cache, legend from the .qml)
│   └── classified.py any classified map with its legend, e.g. from SITS
└── utils.py          Checksums (sha256_file) and BDC tile importer (import_bdc_grids)

Adding a new operator

Create a subclass of Operator in any module imported at startup:

from rasterio.warp import Resampling
from disscube.operators.base import Operator

class WeightedMeanOperator(Operator):
    name = "weighted_mean"
    _resampling = Resampling.average

    def compute(self, data, var, grid):
        # data is an xr.DataArray (raster) or a GeoDataFrame (vector)
        ...

The operator is registered automatically and accepted by Derivation / SpatialDerivation with no other change.

Known limitations

The limitations below are scope decisions for the current version, not bugs. They are documented so that users and reviewers understand what is implemented versus what is planned.

In-memory, single-tile processing Each call to derive() loads a tile's full data into memory. There is no lazy (Dask) or distributed processing. For continental-scale grids (e.g. BR/1km), use the tile loop — each tile is processed and saved independently.

Vector aggregation by rasterization (not area-weighted) Operators over vector sources (majority, percentage, attribute, presence, minority) convert geometries to raster before aggregating pixels. For the share of each cell covered by polygons, use area, which intersects the polygons with the cells exactly.

Tile disambiguation in load() CubeClient.load(name) without tile_id raises ValueError when multiple tiles of the same variable exist on the same grid. Automatic mosaicking is not implemented. Always pass tile_id in multi-tile workloads.

SpatialRelation does not act in the pipeline The SpatialRelation model is persisted in the catalog, but no pipeline stage uses relations during derivation — which is why they are excluded from spec_hash. Including them would make the cache key sensitive to metadata that does not affect the result, breaking the reproducibility guarantee. Integration with hierarchical grid strategies is reserved for a future version.

purity_threshold reserved The purity_threshold field on Derivation is included in spec_hash but is not applied to the output — purity masking is not implemented. Setting purity_threshold changes the cache key without changing the result.

STAC: reading only disscube.sources.bdc reads Brazil Data Cube cubes through their STAC catalog (search, windowed reads, per-tile composites, mosaics) and writes local GeoTIFFs that are registered as ordinary sources. Derived variables are not published back as STAC, and the valid_from/valid_until and bbox fields on Derivation only follow STAC naming conventions.

Citation

If you use DisSCube in your research, dynamic modeling, or spatial data pipelines, please cite it using the metadata from CITATION.cff or the following BibTeX entry:

@software{costa_disscube_2026,
  author       = {Costa, S{\'e}rgio Souza},
  title        = {{DisSCube: Declarative spatial data cubes}},
  year         = {2026},
  version      = {0.4.0},
  url          = {https://github.com/DisSModel/disscube}
}

License

DisSCube is part of the DisSModel ecosystem and is released under the MIT License. See LICENSE for details.

About

Spatial data catalog and raster derivation pipeline for simulation workflows.

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages