|
28 | 28 | "\n", |
29 | 29 | "*Part V — Domain Modeling: Land Use & Coastal Systems*\n", |
30 | 30 | "\n", |
31 | | - "Implemented by the [`disslucc-continuous`](https://github.com/DisSModel/disslucc-continuous) and [`disslucc-discrete`](https://github.com/DisSModel/disslucc-discrete) packages.\n", |
32 | | - "\n", |
33 | | - "<div class=\"admonition warning\">\n", |
34 | | - "<p class=\"admonition-title\">Watch out</p>\n", |
35 | | - "<p>This chapter is a draft, deliberately written ahead of the underlying DisSLUCC packages settling — student work on both is still in progress and the API is likely to shift. Treat this as a base to revise once that work stabilizes, not a final reference; the CLI commands shown were not re-verified end to end against the currently installed packages.</p>\n", |
36 | | - "</div>" |
| 31 | + "Implemented by the [`disslucc`](https://github.com/DisSModel/disslucc) package." |
37 | 32 | ] |
38 | 33 | }, |
39 | 34 | { |
|
45 | 40 | "\n", |
46 | 41 | "By the end of this chapter you will be able to:\n", |
47 | 42 | "\n", |
48 | | - "- Explain what DisSLUCC means as a family, not a single package\n", |
| 43 | + "- Explain why `disslucc` implements both continuous and discrete allocation in one package, on one raster substrate\n", |
49 | 44 | "- Choose between continuous and discrete allocation for a given research question\n", |
50 | 45 | "- Describe the Demand/Potential/Allocation loop every LUCC model in the ecosystem shares\n", |
51 | | - "- Know how each package validates itself against its TerraME/LUCCME predecessor" |
| 46 | + "- Know how each allocation style validates itself against its TerraME/LUCCME predecessor, and run it yourself via script or CLI" |
52 | 47 | ] |
53 | 48 | }, |
54 | 49 | { |
|
77 | 72 | "metadata": {}, |
78 | 73 | "id": "ddc11f3a", |
79 | 74 | "source": [ |
80 | | - "## What DisSLUCC Is\n", |
| 75 | + "## What `disslucc` Is\n", |
81 | 76 | "\n", |
82 | | - "**DisSLUCC** isn't one package — it's the name for two libraries that implement spatially explicit land-use and cover change modeling on top of `dissmodel`, mirroring two allocation philosophies LUCCME itself historically supported:\n", |
| 77 | + "`disslucc` implements spatially explicit land-use and cover change modeling on top of `dissmodel`, raster-only, covering two allocation philosophies LUCCME itself historically supported — not as two packages, but as two components within one:\n", |
83 | 78 | "\n", |
84 | | - "| Approach | Package | Style | Unit of allocation |\n", |
85 | | - "|---|---|---|---|\n", |
86 | | - "| Continuous | `disslucc-continuous` | LUCCME-like | area/percentage per cell |\n", |
87 | | - "| Discrete | `disslucc-discrete` | CLUE-S-like | one land use per cell |\n", |
| 79 | + "| Approach | Style | Unit of allocation | Potential | Allocation |\n", |
| 80 | + "|---|---|---|---|---|\n", |
| 81 | + "| Continuous | LUCCME-like | area/percentage per cell | `PotentialLinearRegression` | `AllocationClueLike` |\n", |
| 82 | + "| Discrete | CLUE-S-like | one land use per cell | `PotentialDLogisticRegression` | `AllocationDClueSLike` |\n", |
88 | 83 | "\n", |
89 | | - "Both depend on `dissmodel` as an ordinary package dependency, following the same additive philosophy Chapter 33 described for every satellite package in the ecosystem — neither modifies the core." |
| 84 | + "Both import from the same `disslucc` package and depend on `dissmodel` as an ordinary package dependency, following the same additive philosophy Chapter 33 described for every satellite package in the ecosystem — neither modifies the core. There are two entry points into the same underlying components: a **script-first** path (build the models directly in Python — no TOML, no provenance tracking, the fastest way to experiment) and an **Executor** path (`disslucc.executors`, wraps the same math in `ModelExecutor`/`ExperimentRecord` for automatic provenance, and is what gets registered in [`dissmodel-configs`](https://github.com/DisSModel/dissmodel-configs) to run on `dissmodel-platform`). Both produce identical results; which one you reach for depends on whether you need the provenance record." |
90 | 85 | ] |
91 | 86 | }, |
92 | 87 | { |
|
110 | 105 | "metadata": {}, |
111 | 106 | "id": "8435fe0f", |
112 | 107 | "source": [ |
113 | | - "## Continuous Allocation: disslucc-continuous\n", |
| 108 | + "## Continuous Allocation\n", |
| 109 | + "\n", |
| 110 | + "`PotentialLinearRegression` + `AllocationClueLike` answer \"how much does this cell's land use change\" — the right choice whenever a cell can legitimately hold more than one land use at once (a partially-deforested cell, a partially-urbanized one). Potential is one linear regression per land-use type, each with its own intercept and driving-factor coefficients (`RegressionSpec`); allocation distributes demand across cells according to that potential map, subject to per-class minimum/maximum bounds (`AllocationSpec`).\n", |
114 | 111 | "\n", |
115 | | - "`disslucc-continuous` answers \"how much does this cell's land use change\" — the right choice whenever a cell can legitimately hold more than one land use at once (a partially-deforested cell, a partially-urbanized one). Potential comes from `PotentialLinearRegression`, one regression per land-use type, each with its own intercept and driving-factor coefficients; allocation comes from `AllocationClueLike`, distributing demand across cells according to that potential map, subject to per-class minimum/maximum bounds.\n", |
| 112 | + "Script-first — no `ModelExecutor`, no TOML, the fastest way to run this:\n", |
116 | 113 | "\n", |
117 | 114 | "```python\n", |
118 | | - "from disslucc_continuous import (\n", |
119 | | - " DemandPreComputedValues, load_demand_csv,\n", |
120 | | - " PotentialLinearRegression, RegressionSpec,\n", |
121 | | - " AllocationClueLike, AllocationSpec,\n", |
122 | | - ")\n", |
| 115 | + "from dissmodel.core import Environment\n", |
| 116 | + "from disslucc import DemandInline, PotentialLinearRegression, AllocationClueLike\n", |
| 117 | + "from disslucc.schemas import RegressionSpec, AllocationSpec\n", |
123 | 118 | "\n", |
124 | | - "demand = DemandPreComputedValues(\n", |
125 | | - " annual_demand=load_demand_csv(\"demand.csv\", [\"forest\", \"dev\", \"other\"]),\n", |
126 | | - " land_use_types=[\"forest\", \"dev\", \"other\"],\n", |
127 | | - ")\n", |
| 119 | + "env = Environment(end_time=7)\n", |
| 120 | + "\n", |
| 121 | + "demand = DemandInline(values=demand_matrix, land_use_types=LAND_USE_TYPES)\n", |
128 | 122 | "\n", |
129 | 123 | "potential = PotentialLinearRegression(\n", |
130 | | - " gdf=gdf,\n", |
131 | | - " land_use_types=[\"forest\", \"dev\", \"other\"],\n", |
132 | | - " land_use_no_data=\"other\",\n", |
| 124 | + " backend=backend,\n", |
| 125 | + " demand=demand,\n", |
| 126 | + " land_use_types=LAND_USE_TYPES,\n", |
133 | 127 | " potential_data=[[\n", |
134 | | - " RegressionSpec(const=0.74, betas={\"dist_roads\": -0.22, \"protected\": 0.18}),\n", |
135 | | - " RegressionSpec(const=0.27, betas={\"dist_roads\": -9.9e-7}),\n", |
136 | | - " RegressionSpec(const=0.0),\n", |
| 128 | + " RegressionSpec(const=-0.2, betas={\"slope\": 0.4}), # forest\n", |
| 129 | + " RegressionSpec(const=0.4, betas={\"dist_road\": -0.1, \"slope\": -0.5}), # agriculture\n", |
| 130 | + " RegressionSpec(const=0.3, betas={\"dist_road\": -0.6, \"slope\": -0.3}), # urban\n", |
137 | 131 | " ]],\n", |
138 | 132 | ")\n", |
| 133 | + "\n", |
| 134 | + "allocation = AllocationClueLike(\n", |
| 135 | + " backend=backend, demand=demand, potential=potential,\n", |
| 136 | + " land_use_types=LAND_USE_TYPES,\n", |
| 137 | + " static={\"forest\": 0, \"agriculture\": -1, \"urban\": -1},\n", |
| 138 | + " complementar_lu=\"forest\", cell_area=1.0, max_difference=5.0,\n", |
| 139 | + " allocation_data=[AllocationSpec(static=0), AllocationSpec(static=-1), AllocationSpec(static=-1)],\n", |
| 140 | + ")\n", |
| 141 | + "\n", |
| 142 | + "env.run()\n", |
139 | 143 | "```\n", |
140 | 144 | "\n", |
141 | | - "Validation follows the same benchmark pattern Chapter 32 already described for the ecosystem generally: a dedicated benchmark executor runs vector and raster substrates side by side and checks both against a TerraME/LUCCME reference dataset, asserting mean absolute error and root-mean-square error stay under a configurable tolerance — never trusting a production run without that check passing first." |
| 145 | + "**Validated against real TerraME/LUCCME output**, not synthetic data: Lab1 (Amazon deforestation, csAC region, 6,574 cells, 6 steps) reproduces the reference at **MAE = 0.0036** (RMSE 0.0062, max error 0.027), well within the 0.01 tolerance the original benchmark used to call something \"equivalent to TerraME\". The Pontius & Millones decomposition attributes 90% of that residual to quantity disagreement and 10% to allocation disagreement — the spatial pattern matches almost perfectly; the small gap is in total allocated quantity, not position. The residual itself traces to the original LuccME script's own convergence tolerance (`maxDifference = 1643` against a 2014 demand of 21,607 — a 7.6% band the reference itself doesn't close). Runs at 44.0 ms/step." |
142 | 146 | ] |
143 | 147 | }, |
144 | 148 | { |
145 | 149 | "cell_type": "markdown", |
146 | 150 | "metadata": {}, |
147 | 151 | "id": "1aad283a", |
148 | 152 | "source": [ |
149 | | - "## Discrete Allocation: disslucc-discrete\n", |
| 153 | + "## Discrete Allocation\n", |
| 154 | + "\n", |
| 155 | + "`PotentialDLogisticRegression` + `AllocationDClueSLike` answer a different question: \"which single land use dominates this cell\" — a CLUE-S-style allocation, right whenever ground-truth is itself categorical (a classified land-cover map, not a fractional-cover raster). Potential comes from a logistic regression predicting *which class* a cell most likely belongs to, not *how much* of a fractional quantity it holds (`LogisticRegressionSpec`, adding an `elasticity` term over the continuous case); allocation runs through a competition-based CLUE-S loop governed by a transition matrix (`[region][from][to]`, which pairs are even reachable).\n", |
150 | 156 | "\n", |
151 | | - "`disslucc-discrete` answers a different question: \"which single land use dominates this cell\" — a CLUE-S-style allocation, right whenever ground-truth is itself categorical (a classified land-cover map, not a fractional-cover raster). Potential here comes from `PotentialDLogisticRegression` instead of a linear one — a logistic regression predicts *which class* a cell most likely belongs to, not *how much* of a fractional quantity it holds — and allocation runs through a competition-based `AllocationDClueSLike`. Configuration lives entirely in a TOML file rather than inline Python — land-use types, per-class regression coefficients, elasticities, and the allowed transition matrix:\n", |
| 157 | + "The Executor path is what registers with `dissmodel-configs`/`dissmodel-platform`, and — as of `dissmodel` 0.6.4 — also runs locally via its CLI, configuration in a TOML file rather than inline Python:\n", |
152 | 158 | "\n", |
153 | 159 | "```toml\n", |
154 | 160 | "[model]\n", |
155 | | - "land_use_types = [\"forest\", \"dev\", \"other\"]\n", |
156 | | - "region_attr = \"region\"\n", |
| 161 | + "land_use_types = [\"f\", \"d\", \"o\"]\n", |
| 162 | + "transition_matrix = [[[1, 1, 0], [0, 1, 0], [0, 0, 1]]] # irreversible deforestation\n", |
| 163 | + "\n", |
| 164 | + "[model.parameters]\n", |
| 165 | + "n_steps = 6\n", |
157 | 166 | "\n", |
158 | | - "[[model.potential]]\n", |
159 | | - "const = -2.34\n", |
| 167 | + "[[model.potential_data]]\n", |
| 168 | + "lu = \"f\"\n", |
| 169 | + "const = -2.34187976925989\n", |
160 | 170 | "elasticity = 0.0\n", |
161 | | - "[model.potential.betas]\n", |
162 | | - "soil_decl = -0.03\n", |
163 | | - "dist_road = 3.10\n", |
164 | | - "\n", |
165 | | - "[model.allocation]\n", |
166 | | - "max_difference = 10.0\n", |
167 | | - "max_iteration = 1000\n", |
168 | | - "factor_iteration = 0.0001\n", |
| 171 | + " [model.potential_data.betas]\n", |
| 172 | + " dist_br = 3.10319957497883\n", |
169 | 173 | "```\n", |
170 | 174 | "\n", |
171 | | - "`disslucc-discrete`'s validation reaches the strongest bar anywhere in this book: not a tolerance band, but **exact cell-level parity** against a TerraME/LuccME reference dataset — 100% accuracy, Cohen's κ = 1.0, F1 = 1.0, checked automatically in continuous integration. That exactness is possible specifically *because* the output is categorical — a cell either matches its reference class or it doesn't, with no \"close enough\" in between the way a continuous percentage would have." |
| 175 | + "```bash\n", |
| 176 | + "python -m disslucc.executors.discrete run \\\n", |
| 177 | + " --toml examples/dissmodel-configs/lucc_discrete.toml \\\n", |
| 178 | + " --input data/input/cs_moju.zip \\\n", |
| 179 | + " --param demand_csv=data/input/demand_moju.csv \\\n", |
| 180 | + " --output outputs/result.tif\n", |
| 181 | + "```\n", |
| 182 | + "\n", |
| 183 | + "**Validated against real TerraME/LUCCME output**: Lab15 (deforestation, Moju region, 5,914 cells, 6 steps) reaches **exact cell-for-cell agreement** — zero quantity disagreement, zero allocation disagreement (Pontius & Millones decomposition), 100% accuracy, F1 = 1.0. Runs at 10.3 ms/step.\n", |
| 184 | + "\n", |
| 185 | + "**Discriminance caveat, inherited from the original benchmark and still true here:** the Lab15 scenario is nearly non-discriminative — a trivial static ranking by `(prob_d - prob_f)`, with no CLUE-S, no iteration, no time steps, already reproduces the same cell-by-cell output. So this result confirms the logistic regression coefficients were transcribed correctly; it does *not* by itself prove the CLUE-S allocation algorithm (iteration/convergence via `factor_iteration`) is faithful in a scenario where competition between classes actually matters. A shipped discriminance test makes this explicit rather than letting the strong headline number stand unqualified; a dynamic-covariate scenario that would actually exercise the competition logic is planned." |
172 | 186 | ] |
173 | 187 | }, |
174 | 188 | { |
|
178 | 192 | "source": [ |
179 | 193 | "## Choosing Between Continuous and Discrete\n", |
180 | 194 | "\n", |
181 | | - "Both packages implement the identical Demand/Potential/Allocation loop with different algorithms at each step — the decision between them is about your data, not architecture:\n", |
| 195 | + "Both components implement the identical Demand/Potential/Allocation loop with different algorithms at each step — the decision between them is about your data, not architecture:\n", |
182 | 196 | "\n", |
183 | | - "| | `disslucc-continuous` | `disslucc-discrete` |\n", |
| 197 | + "| | Continuous | Discrete |\n", |
184 | 198 | "|---|---|---|\n", |
185 | 199 | "| Potential | linear regression | logistic regression |\n", |
186 | 200 | "| Allocation | CLUE-like, per-class bounds | competition-based CLUE-S |\n", |
187 | 201 | "| Validation | MAE/RMSE tolerance vs. reference | exact cell-level parity vs. reference |\n", |
| 202 | + "| Speed (validated benchmark) | 44.0 ms/step | 10.3 ms/step |\n", |
188 | 203 | "\n", |
189 | | - "The fastest way to decide: look at your calibration and validation data first. If it's expressed as a class label per cell, `disslucc-discrete` is the model that can be checked against it exactly. If it's expressed as an area or percentage per cell, `disslucc-continuous` is the only one that represents it without lossy discretization forced on it first." |
| 204 | + "The fastest way to decide: look at your calibration and validation data first. If it's expressed as a class label per cell, discrete allocation is what can be checked against it exactly. If it's expressed as an area or percentage per cell, continuous allocation is the only one that represents it without lossy discretization forced on it first." |
190 | 205 | ] |
191 | 206 | }, |
192 | 207 | { |
|
208 | 223 | "source": [ |
209 | 224 | "## Exercises\n", |
210 | 225 | "\n", |
211 | | - "1. **Match the package to the question.** For each research question, name which package fits and why: (a) \"how does the percentage of forest cover in each cell change over the next 20 years,\" (b) \"which cells convert from forest to pasture by 2030.\"\n", |
212 | | - "2. **Why two different regressions?** `PotentialLinearRegression` and `PotentialDLogisticRegression` both take driving-factor coefficients (`betas`) per land-use type. From the model names alone, explain why one needs a linear regression and the other a logistic one — what is each one actually trying to predict?\n", |
213 | | - "3. **Tolerance vs. exact parity.** Explain, in your own words, why a tolerance-based check makes sense for `disslucc-continuous`'s validation but not for `disslucc-discrete`'s.\n", |
| 226 | + "1. **Match the component to the question.** For each research question, name which allocation style fits and why: (a) \"how does the percentage of forest cover in each cell change over the next 20 years,\" (b) \"which cells convert from forest to pasture by 2030.\"\n", |
| 227 | + "2. **Why two different regressions?** `PotentialLinearRegression` and `PotentialDLogisticRegression` both take driving-factor coefficients (`betas`) per land-use type. From the class names alone, explain why one needs a linear regression and the other a logistic one — what is each one actually trying to predict?\n", |
| 228 | + "3. **Tolerance vs. exact parity.** Explain, in your own words, why a tolerance-based check makes sense for the continuous allocation's validation but not for the discrete one's.\n", |
214 | 229 | "4. **Multiscale comparison, by hand.** Sketch, conceptually, how you'd compute a 3×3-window agreement score between two categorical land-use grids of the same shape — what would you compare within each window, and how would you turn that into a single agreement number for the whole grid?" |
215 | 230 | ] |
216 | 231 | }, |
|
233 | 248 | "\n", |
234 | 249 | "### Key concepts introduced\n", |
235 | 250 | "\n", |
236 | | - "- DisSLUCC as two packages, not one, sharing the identical Demand/Potential/Allocation loop LUCCME originally established\n", |
237 | | - "- `disslucc-continuous`: fractional, per-cell land-use change, validated by MAE/RMSE tolerance against a TerraME/LUCCME reference\n", |
238 | | - "- `disslucc-discrete`: categorical, one-class-per-cell allocation, validated by exact cell-level parity — the strongest equivalence claim in the ecosystem, possible because the output is categorical\n", |
239 | | - "- Choosing between them by looking at your own calibration/validation data's type first, not by architectural preference\n", |
| 251 | + "- `disslucc`: one raster-only package implementing both continuous and discrete LUCC allocation, sharing the identical Demand/Potential/Allocation loop LUCCME originally established\n", |
| 252 | + "- **Continuous**: fractional, per-cell land-use change, validated by MAE/RMSE tolerance against a real TerraME/LUCCME reference (MAE 0.0036, 44.0 ms/step)\n", |
| 253 | + "- **Discrete**: categorical, one-class-per-cell allocation, validated by exact cell-level parity — the strongest equivalence claim in the ecosystem, possible because the output is categorical (10.3 ms/step) — with a discriminance test making explicit that this validates coefficient transcription, not full allocation-algorithm fidelity\n", |
| 254 | + "- Two entry points into the same components: script-first (fastest to experiment) and `ModelExecutor`-based (automatic provenance, TOML registration, CLI)\n", |
| 255 | + "- Choosing between allocation styles by looking at your own calibration/validation data's type first, not by architectural preference\n", |
240 | 256 | "- Engineering validation (matching a TerraME reference) versus scientific validation (calibration/validation split against real-world data), and multiscale comparison as an alternative to an overly strict pixel-for-pixel match\n", |
241 | 257 | "\n", |
242 | 258 | "Chapter 27 stays in domain-application territory but moves from land far inland to the coastline itself — a coupled flood and mangrove-migration model, run on both substrates and validated two different ways at once." |
|
252 | 268 | "- Verburg, P. H. et al. (2002). \"Modeling the spatial dynamics of regional land use: the CLUE-S model.\" *Environmental Management*, 30(3), 391-405\n", |
253 | 269 | "- Verburg, P. H. et al. (2006). \"Downscaling of land use change scenarios to assess the dynamics of European landscapes.\" *Agriculture, Ecosystems & Environment*, 114(1), 39-56\n", |
254 | 270 | "- Costanza, R. (1989). \"Model goodness of fit: a multiple resolution procedure.\" *Ecological Modelling*, 47(3-4), 199-215 — the multiscale comparison method this chapter's validation section draws on\n", |
| 271 | + "- Pontius Jr., R. G. & Millones, M. (2011). \"Death to Kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment.\" *International Journal of Remote Sensing*, 32(15), 4407-4429 — the decomposition `disslucc` uses in place of Cohen's kappa\n", |
255 | 272 | "- LuccME documentation (INPE): <http://www.dpi.inpe.br/luccme/>\n", |
256 | | - "- disslucc-continuous and disslucc-discrete on GitHub: <https://github.com/DisSModel>" |
| 273 | + "- `disslucc` on GitHub: <https://github.com/DisSModel/disslucc>" |
257 | 274 | ] |
258 | 275 | } |
259 | 276 | ] |
|
0 commit comments