Status: Alpha — stable APIs for the core pipeline; declarative models still evolving.
DisSCube is the spatial data cube engine of the DisSModel ecosystem. It converts raw geospatial sources (rasters, vectors) into derived variables aligned to LUCC (Land Use and Cover Change) modeling grids, ready for Cellular Automata models and spatio-temporal analysis.
📖 Documentation: https://dissmodel.github.io/disscube/
SpatialSource → Derivation → Variable → DerivedVariable (Zarr)
A source (SpatialSource) goes through a derivation (SpatialDerivation or Derivation) that applies an operator on a grid (GridSpec), producing a derived variable registered in the SQLite catalog and stored in Zarr.
git clone https://github.com/DisSModel/disscube.git
cd disscube
python -m venv .venv && source .venv/bin/activate
pip install -e .from disscube.client import CubeClient
from disscube.utils.grids import register_local_grid
cube = CubeClient(catalog="catalog.db", store="./data/")
grid = register_local_grid(
cube,
name="AC",
bbox_geo=(-73.99, -11.15, -66.62, -7.11),
resolution=5_000.0,
)from disscube.models import SpatialSource
cube.register_spatial_source(SpatialSource(
id="mapbiomas_2020",
name="MapBiomas Acre 2020",
format="raster",
asset_url="data/raw/mapbiomas_2020.tif",
crs="EPSG:4326",
time=2020,
))from disscube.derivation import Derivation
d = Derivation(
target="forest_pct",
source_id="mapbiomas_2020",
operator="percentage",
class_code=3,
role="driver",
valid_from="2020",
valid_until="2020",
)
cube.derive_declarative(d, grid_id="AC/5km")from disscube.models import SpatialDerivation, Variable
cube.derive(SpatialDerivation(
source_id="mapbiomas_2020",
grid_id="AC/5km",
role="driver",
variables=[Variable(name="forest_pct", operator="percentage", class_code=3)],
valid_from="2020",
valid_until="2020",
))da = cube.load("forest_pct", grid_id="AC/5km")
print(da.shape) # (rows, cols)ds = cube.to_dataset(["forest_pct", "dist_roads"], grid_id="AC/5km", period=("2015", "2020"))
# xarray.Dataset: (y, x) static and (time, y, x) temporal variables, CRS and transform via ds.rio
cube.export_geotiff(["forest_pct"], "forest.tif", grid_id="AC/5km") # one band per variable and year
cube.export_netcdf(["forest_pct"], "cube.nc", grid_id="AC/5km") # CF-1.8; pip install "disscube[netcdf]"Exports carry their provenance: each GeoTIFF band and netCDF variable records the
spec_hash, the content_hash of the stored data and the source_checksum of the
input it came from (cube.provenance("forest_pct") lists them per year).
DisSCube does not need DisSModel. To hand a cube to a DisSModel model, install
disscube[dissmodel] and use cube.to_raster_backend(...), which returns a RasterBackend.
A whole data preparation — grid, sources, derived variables — can be declared
in one TOML file, the counterpart of a TerraME fill script, and run from the
command line:
disscube validate examples/pipelines/quickstart.toml # no downloads
disscube run examples/pipelines/quickstart.toml --workspace outputs/quickstartschema = 1
name = "Quickstart"
[grid]
name = "demo/300m"
crs = "EPSG:31983"
bbox = [570000.0, 9708000.0, 582000.0, 9720000.0]
resolution = 300
[[source]]
id = "landuse"
type = "file"
path = "../data/quickstart/landuse.tif"
crs = "EPSG:31983"
[[derive]]
target = "forest_pct"
source = "landuse"
operator = "percentage"
class_code = 3See docs/guides/pipeline_files.md and
examples/pipelines/.
examples/ has runnable, self-contained examples that run offline in seconds:
- 01–03, synthetic data — a quickstart with raster operators, vector drivers, and time series handed off to DisSModel.
- quickstart.toml — declarative pipeline equivalent to example 01.
python examples/01_quickstart.py
disscube run examples/pipelines/quickstart.tomlReal data workflows, TerraME parity benchmarks, and large-scale case studies are maintained in
LambdaGeo/disscube-recipes (cases/terrame_fill, cases/ilha_maranhao, cases/prodes_br163, cases/luccme_br).
See examples/README.md for the full list, and
docs/terrame_fill_correspondence.md
for how DisSCube's operators relate to TerraME's Fill.
| Operator | Type | Resampling | Requires class_code |
|---|---|---|---|
mean |
zonal | average | no |
sum |
zonal | sum | no |
std |
zonal | nearest | no |
min |
zonal | min | no |
max |
zonal | max | no |
majority |
zonal | nearest¹ | no |
minority |
zonal | nearest¹ | no |
percentage |
zonal | nearest¹ | yes |
attribute |
zonal | nearest | no |
presence |
zonal | nearest | no |
distance |
proximity (exact, cell centre → nearest feature, source not clipped) | — | no |
min_distance |
proximity (raster approximation, features inside the grid) | nearest | no |
count |
proximity | nearest | no |
area |
polygons (exact share of the cell covered) | — | no |
distance takes params = {crs = …} to measure in another CRS (metres on a
geographic grid); the ¹ operators take params = {subcells = n} to cap the
fine pixels per cell. Any derivation takes fill = "nearest".
¹ These use
needs_fine_alignment=True: GridAligner resamples withnearestat high resolution; the actual reduction (per-window counting) is done by the operator.
SpatialSource
│
▼
Normalizer — validates / loads a GeoDataFrame (vector) or opens the raster
│
▼
GridAligner — reprojects per variable with the operator's Resampling
│
▼
Aggregator — delegates to operator.compute() → one xr.DataArray per variable
│
▼
VariableWriter — writes Zarr + registers the DerivedVariable in the catalog
data/derived/{grid_id}/{partition}/{spec_hash}/{variable_name}.zarr
partition=tile_id, orglobalfor untiled derivations.spec_hash= SHA-256 of the derivation (source + grid + variables + time window, plus the source'schecksumwhen it has one — replacing a source file and registering its new checksum recomputes instead of returning a stale product).
disscube/
├── client.py CubeClient — public entry point
├── models/ GridSpec, SpatialSource, SpatialDerivation, Variable, Derivation…
├── operators/ Operators as classes (self-registered via __init_subclass__)
│ ├── base.py Operator ABC + OPERATOR_REGISTRY
│ ├── zonal.py mean, sum, majority, percentage, attribute, presence…
│ └── proximity.py distance, min_distance, count
├── pipeline/ Pipeline execution & planning (schema, runner) + internal stages
├── catalog/ CatalogStore (Protocol) + SQLite and JSON implementations
├── storage.py AssetStore (fsspec — local and S3)
├── cli.py `disscube validate` / `disscube run` / `disscube export`
├── sources/ Adapters that bring external data in as SpatialSources,
│ │ each with a checksum and a provenance.json sidecar
│ ├── _raster.py Window2D, windowed reads, composites, mosaics, register_raster
│ ├── bdc.py Brazil Data Cube cubes via STAC (`bdc` extra)
│ ├── _categorical.py legends (.qml/.json/.csv) and reclassification
│ ├── mapbiomas.py MapBiomas annual land-cover maps (Collection 11, 10 m series)
│ ├── prodes.py PRODES deforestation (download + cache, legend from the .qml)
│ └── classified.py any classified map with its legend, e.g. from SITS
└── utils.py Checksums (sha256_file) and BDC tile importer (import_bdc_grids)
Create a subclass of Operator in any module imported at startup:
from rasterio.warp import Resampling
from disscube.operators.base import Operator
class WeightedMeanOperator(Operator):
name = "weighted_mean"
_resampling = Resampling.average
def compute(self, data, var, grid):
# data is an xr.DataArray (raster) or a GeoDataFrame (vector)
...The operator is registered automatically and accepted by Derivation / SpatialDerivation with no other change.
The limitations below are scope decisions for the current version, not bugs. They are documented so that users and reviewers understand what is implemented versus what is planned.
In-memory, single-tile processing
Each call to derive() loads a tile's full data into memory. There is no lazy (Dask) or distributed processing. For continental-scale grids (e.g. BR/1km), use the tile loop — each tile is processed and saved independently.
Vector aggregation by rasterization (not area-weighted)
Operators over vector sources (majority, percentage, attribute, presence, minority) convert geometries to raster before aggregating pixels. For the share of each cell covered by polygons, use area, which intersects the polygons with the cells exactly.
Tile disambiguation in load()
CubeClient.load(name) without tile_id raises ValueError when multiple tiles of the same variable exist on the same grid. Automatic mosaicking is not implemented. Always pass tile_id in multi-tile workloads.
SpatialRelation does not act in the pipeline
The SpatialRelation model is persisted in the catalog, but no pipeline stage uses relations during derivation — which is why they are excluded from spec_hash. Including them would make the cache key sensitive to metadata that does not affect the result, breaking the reproducibility guarantee. Integration with hierarchical grid strategies is reserved for a future version.
purity_threshold reserved
The purity_threshold field on Derivation is included in spec_hash but is not applied to the output — purity masking is not implemented. Setting purity_threshold changes the cache key without changing the result.
STAC: reading only
disscube.sources.bdc reads Brazil Data Cube cubes through their STAC catalog (search, windowed reads, per-tile composites, mosaics) and writes local GeoTIFFs that are registered as ordinary sources. Derived variables are not published back as STAC, and the valid_from/valid_until and bbox fields on Derivation only follow STAC naming conventions.
If you use DisSCube in your research, dynamic modeling, or spatial data pipelines, please cite it using the metadata from CITATION.cff or the following BibTeX entry:
@software{costa_disscube_2026,
author = {Costa, S{\'e}rgio Souza},
title = {{DisSCube: Declarative spatial data cubes}},
year = {2026},
version = {0.4.0},
url = {https://github.com/DisSModel/disscube}
}DisSCube is part of the DisSModel ecosystem and is released under the MIT License. See LICENSE for details.