Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions PRD.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,13 +3,15 @@
| | |
|---|---|
| **Document** | PRD — Laige 2.5D Multi-OS Game Engine (C++) |
| **Version** | 0.3 (draft) |
| **Version** | 0.4 (draft) |
| **Status** | Proposed — pending review |
| **Owner** | Engine team |
| **Last updated** | 2026-09-10 |
| **Last updated** | 2026-10-04 |

> Name: **Laige** — an acronym for *Legendary AI Game Engine*. Confirmed as the final name on 2026-09-10 (decision D-NAME, ADR 0001).
>
> **v0.4 change:** §8.1: the isometric depth-key rebuild budget is re-baselined from ≤ 0.2 ms to **≤ 0.3 ms** (mean, 10k dirty cells after a terrain edit). The absolute-target gate runs on the CI reference machine (ubuntu-24.04, Clang 18.1.3, CMake Debug — methodology §5: "CI is the gate"), where the engine's unchanged `setTile` path measures 0.194–0.267 ms across runner variance — the 0.2 ms bar (calibrated from local-machine runs of 0.081 ms) had zero margin there and flapped on master and PR lanes with identical engine code (no regression). The workload, the engine path, and the measurement method are unchanged. Evidence: `docs/benchmarks/baselines/m2-iso-depth-table-budget-rebaseline.md`.
>
> **v0.3 change:** engine name confirmed as "Laige" (acronym for *Legendary AI Game Engine*); §18 item 1 resolved (ADR 0001).
>
> **v0.2 change:** isometric is designated the **primary projection** — most Laige games will be isometric. It is the default template, the reference scene for all visual/performance acceptance tests, and the focus of new first-class requirements (depth keys, picking, grid-snap camera, grid-aligned AOI).
Expand Down Expand Up @@ -245,7 +247,7 @@ Requirement IDs are tracked. `P0` = must ship in M1–M5, `P1` = M6–M7, `P2` =
| Frame time (render) | p95 ≤ 8.3 ms @ 1080p (60 FPS) | Mid-range laptop (2019–2023 class) |
| Simulation tick (10k entities, 2k dynamic bodies) | ≤ 3.0 ms avg, ≤ 5 ms p99 | Same |
| 50k visible sprites (worst-case isometric overlap), 3 parallax layers, UI | ≤ 30 draw calls; ≤ 2 ms CPU | Same |
| Isometric depth-key rebuild (10k dirty cells after terrain edit) | ≤ 0.2 ms | Same |
| Isometric depth-key rebuild (10k dirty cells after terrain edit) | ≤ 0.3 ms | Same |
| Isometric screen→grid picking | O(1), ≤ 0.01 ms per pick | Same |
| Steady-state heap allocations in sim loop | **0 per frame** (asserted in debug) | Debug builds |
| Engine base memory (empty scene, running) | ≤ 100 MB RSS | All P0 platforms |
Expand Down
4 changes: 2 additions & 2 deletions budgets.json
Original file line number Diff line number Diff line change
Expand Up @@ -46,8 +46,8 @@
"name": "iso_depthkey_rebuild",
"metric": "mean",
"unit": "ms",
"target": 0.2,
"measured": 0.0814067,
"target": 0.3,
"measured": 0.202204,
"workload": "10k dirty cells after a terrain edit (PRD 8.1)"
},
{
Expand Down
2 changes: 1 addition & 1 deletion docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -235,7 +235,7 @@ still to land.
storage, one allocation at creation), the O(1) zero-allocation
incremental update with its documented neighborhood (radius 0),
the bounded + logged growth, the PRD §8.1 budget (10k dirty cells
≤ 0.2 ms — the `iso_depthkey_rebuild` entry), and the
≤ 0.3 ms — the `iso_depthkey_rebuild` entry, re-baselined 2026-10-04), and the
sim-writes/render-reads threading contract (M2-ISO-02;
`laige-render`).
- [Projection modes + screen↔world
Expand Down
4 changes: 2 additions & 2 deletions docs/api/iso_depth_key.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,8 +120,8 @@ isometric depth sorting (M2-CAM-02 validates scene shears against it).
- The 10k-sprite per-frame cost is measured with the M2-PERF-01
render suite (the key step itself is a trivial fraction of the
§8.1 render CPU budget); the M2-ISO-02 table step adds the
precompute/incremental path (≤ 0.2 ms for 10k dirty cells,
`iso_depth_rebuild` budget).
precompute/incremental path (≤ 0.3 ms for 10k dirty cells,
`iso_depthkey_rebuild` budget — re-baselined 2026-10-04).
- **Common trap:** recomputing keys from screen-space coordinates, or
calling `isoDepthKey` from a getter that also runs a scene
traversal (API-003) — the key is the *result* of a world-state
Expand Down
22 changes: 13 additions & 9 deletions docs/api/iso_depth_table.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ The per-scene-chunk, precomputed, incrementally-updated mapping from the
tile grid to engine-owned isometric depth keys (M2-ISO-02; PRD §4,
FR-2.2 — "precomputed at scene build and incrementally updated on
tile/height changes (not recomputed per frame)", §8.1 — 10k dirty cells
≤ 0.2 ms, AC-4.4 base, S-5, G-R11; AGENTS ARCH-008/009, RENDER-003,
≤ 0.3 ms (re-baselined 2026-10-04), AC-4.4 base, S-5, G-R11; AGENTS ARCH-008/009, RENDER-003,
CORE-002/005, PERF-002/003, SCALE-001/003; ADR 0002). Public header:
`src/laige-render/include/laige/render/iso_depth_table.h` (header-only —
the API is a template over the SimMath backends, the
Expand Down Expand Up @@ -125,14 +125,18 @@ centers would leave the key domain.
(methodology §4), where a helper call is a real function call —
notably the cell is read through a cached raw pointer, since
`unique_ptr::operator[]` is a six-level call chain on that tree.
Budget: **10k dirty cells ≤ 0.2 ms mean** (PRD §8.1,
`iso_depthkey_rebuild`) — measured **0.087 ms (fpx16_16) /
0.086 ms (fp32_pinned)** on the canonical Debug tree (worse of
backends recorded in `budgets.json` and
[baselines/m2-iso-depth-table.md](../benchmarks/baselines/m2-iso-depth-table.md));
the gate also passes on the reference-class clang -O0 tree (0.152 ms
mean — the flat path clears the 0.2 ms bar with margin on both
backends). The zero-allocation contract is asserted by the update
Budget: **10k dirty cells ≤ 0.3 ms mean** (PRD §8.1,
`iso_depthkey_rebuild` — re-baselined from 0.2 ms on 2026-10-04 after
the CI reference lane measured 0.194–0.267 ms on unchanged engine
code: the 0.2 ms bar had zero margin on that toolchain — see
[baselines/m2-iso-depth-table-budget-rebaseline.md](../benchmarks/baselines/m2-iso-depth-table-budget-rebaseline.md));
measured **0.081 ms mean** on the canonical Debug tree and
**0.202204 ms** on the CI reference lane (the worse of the two
backends, recorded in `budgets.json`; the original runs in
[baselines/m2-iso-depth-table.md](../benchmarks/baselines/m2-iso-depth-table.md)
and
[baselines/m2-iso-depth-table-workload-fix.md](../benchmarks/baselines/m2-iso-depth-table-workload-fix.md)).
The zero-allocation contract is asserted by the update
suite where the allocation watch is live (the non-sanitizer trees).
- **`keyAt` / `covers` / `tileHeightAt` — O(1)** flat-index reads, no
allocation, no logging.
Expand Down
11 changes: 11 additions & 0 deletions docs/benchmarks/baselines/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,3 +72,14 @@ Recorded so far:
canonical Debug tree (4.4× inside the 1.0 ms gate); 0.161389 ms on
the clang -O0 cross-run. **The seventh baseline** — updates
`budgets.json` `measured` for `depth_sort_10k` to 0.225852.
- [m2-iso-depth-table-budget-rebaseline.md](m2-iso-depth-table-budget-rebaseline.md)
(M2-ISO-02 budget revision, 2026-10-04) — re-baselines the
`iso_depthkey_rebuild` budget after the CI reference lane (Clang
18.1.3 -O0, ubuntu-24.04) measured 0.194–0.267 ms on unchanged
engine code against the 0.2 ms bar (repeated zero-margin gate
failures, no regression): PRD §8.1 v0.4 `target` 0.2 → **0.3 ms**,
`budgets.json` `measured` 0.0814067 → **0.202204** (latest recorded
reference-lane value, worse backend). Workload and engine code
unchanged. **The eighth baseline** — supersedes
`m2-iso-depth-table-workload-fix.md` as the latest recorded value of
`iso_depthkey_rebuild`.
197 changes: 197 additions & 0 deletions docs/benchmarks/baselines/m2-iso-depth-table-budget-rebaseline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,197 @@
# Baseline: `m2-iso-depth-table-budget-rebaseline` — re-baselined
# `iso_depthkey_rebuild` target (0.2 → 0.3 ms)

Recorded by the **M2-ISO-02 budget revision** (2026-10-04, follow-up to
the M2-ISO-02 workload fix). This is the **eighth** baseline file; it is
immutable (methodology §4 — superseding it later adds a new file, it is
never edited). It **supersedes**
[`m2-iso-depth-table-workload-fix.md`](m2-iso-depth-table-workload-fix.md)
as the latest recorded value of `iso_depthkey_rebuild` (that baseline
stays in place as the before number of this before/after pair;
`budgets.json` `measured` is updated by this baseline, and — per the
NFR-8.1 policy — the absolute `target` is revised in the same change,
together with the PRD revision).

## Why this baseline exists

The `iso_depthkey_rebuild` gate (mean ≤ 0.2 ms, 10k dirty cells after a
terrain edit, PRD §8.1) failed the CI `Linux x64 (clang++)` reference
lane **repeatedly on unchanged engine code**:

- Run 36885653175 (2026-10-01, pre-workload-fix): mean 0.20889 /
0.209044 ms — FAIL. Root cause was a *harness* artifact (three
`div`/`idiv` per iteration in the measured window at `-O0`), removed
by the division-free workload fix (recorded in
`m2-iso-depth-table-workload-fix.md`).
- After that fix the engine code did **not change**, yet the gate kept
flapping on the same reference lane (Clang 18.1.3, CMake Debug,
ubuntu-24.04 shared runner):

| CI run (all `Linux x64 (clang++)`) | fpx16_16 mean | fp32_pinned mean | result |
|---|---|---|---|
| 37192995927 attempt 1 (2026-10-04) | 0.194142 ms | 0.201495 ms | fpx16 PASS / fp32 FAIL |
| 37192995927 attempt 2 (2026-10-04, slow shared runner) | 0.266709 ms | 0.267328 ms | both FAIL |
| 37194767648 (2026-10-04, master merge lane, commit `70830a7`) | 0.202146 / 0.202204 ms | 0.202018 / 0.194168 ms | FAIL (both entries) |

The reference lane measures this workload at **0.194–0.267 ms** —
straddling the 0.2 ms bar, so the gate's outcome is decided by how busy
the shared runner is at the moment, not by the engine. Evidence that
this is runner variance and not a regression:

- The engine's `setTile` update path is unchanged since the
2026-10-01 workload fix (five merges landed in between —
M2-ISO-03, M2-SORT-01, M2-SPRITE-01/02, the MSVC cast fix — none
touches `iso_depth_table`).
- The `Linux x64 (g++)` lane of run 37194767648 **passed** the same
workload on the same commit.
- The recorded numbers (0.0814 ms canonical g++ local, 0.139 ms local
Clang -O0) predate the gate's calibration to the CI lane: the 0.2 ms
bar was never re-verified with margin on the reference toolchain —
the exact lesson the workload-fix baseline recorded ("first-crossings
of a budget gate on a new CI toolchain must be verified there").

## The revision

Per the PRD §8.1 policy ("a PR that regresses any budget by > 10% (or
breaches absolute target) fails CI **unless the budget is revised via a
PRD revision**", NFR-8.1) and methodology §5 ("accepted budget
revisions change `budgets.json` in the same PR as the PRD revision"):

- **PRD §8.1** (v0.4): `≤ 0.2 ms` → `≤ 0.3 ms` (mean, 10k dirty cells).
- **`budgets.json`**: `target` 0.2 → **0.3**; `measured` 0.0814067
(local canonical tree) → **0.202204** — the latest recorded value,
the worse of the two backends on the CI reference lane, run
37194767648 (2026-10-04).
- **Workload, engine code, measurement method: unchanged.** This is a
budget re-baseline, not a benchmark change (methodology §1) — the
workload is still the PRD §8.1 10k-dirty-cells edit, 100 warm-up +
3 000 measured iterations, mean metric, both SimMath backends.

**Margin:** 0.3 ms sits 48% above the latest recorded reference-lane
value (0.202204) and 12.4% above the worst observed reference-lane mean
(0.267328, slow-runner outlier). The gate's metric is the **mean**
(n=3000, warmup=100): single-iteration spikes (worst observed max
0.300107 ms, same outlier run) do not drive the gate. 0.3 ms for 10k
dirty cells is 1.8% of the 16.7 ms 60 FPS frame budget — the product
promise stays tight.

## AGENTS §12 metadata (the recorded run)

| # | Field | Value |
|---|---|---|
| 1 | Hardware | GitHub Actions ubuntu-24.04 hosted runner (2 vCPU shared) — the CI reference machine |
| 2 | OS | ubuntu-24.04 |
| 3 | Compiler and version | Clang 18.1.3 (runner apt package), CMake Debug |
| 4 | Build type | `Debug` (`-O0 -g` + engine policy flags) |
| 5 | Relevant flags | Engine policy (NFR-8.10): `-Wall -Werror -fno-exceptions -fno-rtti`; SimMath pinned set (ADR 0002) on the engine TUs |
| 6 | Dataset / workload | `iso_depth_table` — 128×128 grid (16 384 cells, 64 chunks), terrain `(7gx+11gy)%5`; one iteration = 10 000 `setTile` calls over a 100×100 block at (14,14), column-major, height `(gx+gy)%5`, edit sequence precomputed outside the measured window; both SimMath backends |
| 7 | Warm-up | 100 iterations discarded |
| 8 | Sample count | `n=3000` iterations per backend (histogram capacity 3000, no truncation) |
| 9 | Summary statistics | see the verbatim reports below (per run) |
| 10 | Before / after | `before=0.0814067` (the M2-ISO-02 workload-fix recorded value, canonical local tree) · `after` (mean, worse backend, run 37194767648) = 0.202204 → **recorded 0.202204** · `target`: 0.2 → **0.3 ms** (PRD v0.4 revision) |

Commit measured on: `70830a7` (master, the PR #72 merge commit — the
run that motivated this revision).

## Verbatim run output (CI reference lane, run 37194767648)

### `iso_depth_table` ctest entry, first backend pair

```text
budget=iso_depthkey_rebuild result=FAIL metric=mean unit=ms
after=0.202146 before=0.0814067 target=0.2
stats: n=3000 min=0.1855 mean=0.202146 p50=0.201084 p95=0.208645 p99=0.211189 max=0.236467
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100
```

(second backend of the same entry: `after=0.202018`, stats
`min=0.18515 p50=0.201134 p95=0.208595 p99=0.211289 max=0.23845`)

### `laige-render_tests` (full binary) run of the same job

```text
budget=iso_depthkey_rebuild result=FAIL metric=mean unit=ms
after=0.202204 before=0.0814067 target=0.2
stats: n=3000 min=0.18561 mean=0.202204 p50=0.201334 p95=0.209597 p99=0.213092 max=0.282557
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100
```

```text
budget=iso_depthkey_rebuild result=PASS metric=mean unit=ms
after=0.194168 before=0.0814067 target=0.2
stats: n=3000 min=0.191289 mean=0.194168 p50=0.192781 p95=0.200302 p99=0.204068 max=0.285451
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100
```

(The second backend passed 0.07% under the bar in the same job — the
coin-flip shape of a zero-margin gate.)

### PR-lane evidence (run 37192995927)

Attempt 1 (job 111408940973):

```text
budget=iso_depthkey_rebuild result=PASS metric=mean unit=ms
after=0.194142 before=0.0814067 target=0.2
stats: n=3000 min=0.18234 mean=0.194142 p50=0.192915 p95=0.200226 p99=0.205123 max=0.237111
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100

budget=iso_depthkey_rebuild result=FAIL metric=mean unit=ms
after=0.201495 before=0.0814067 target=0.2
stats: n=3000 min=0.185765 mean=0.201495 p50=0.200497 p95=0.208028 p99=0.210752 max=0.233195
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100
```

Attempt 2 (job 111410540820, slow shared runner — the observed worst
case):

```text
budget=iso_depthkey_rebuild result=FAIL metric=mean unit=ms
after=0.266709 before=0.0814067 target=0.2
stats: n=3000 min=0.184 mean=0.266709 p50=0.267799 p95=0.276597 p99=0.282203 max=0.300107
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100

budget=iso_depthkey_rebuild result=FAIL metric=mean unit=ms
after=0.267328 before=0.0814067 target=0.2
stats: n=3000 min=0.254928 mean=0.267328 p50=0.266224 p95=0.279408 p99=0.283082 max=0.29891
context: workload=10k dirty cells after a terrain edit (PRD 8.1) build=Clang 18.1.3, CMake Debug, engine policy (NFR-8.10) machine= warmup=100
```

## Before/after (CORE-001)

| Quantity | Before (M2-ISO-02 workload fix) | After (this revision) |
|---|---|---|
| PRD §8.1 target (mean) | 0.2 ms | **0.3 ms** |
| `budgets.json` `target` | 0.2 | **0.3** |
| `budgets.json` `measured` | 0.0814067 (local canonical g++ -O0) | **0.202204** (CI reference lane, worse backend) |
| Reference-lane observed range | 0.20889 ms (pre-fix, harness artifact) | 0.194–0.267 ms (engine unchanged) |

## Interpretation

- **The gate is calibrated, not relaxed:** the target is now 48% above
the latest recorded reference-lane value and covers every observed
reference-lane mean with ≥ 12% margin. A future engine change that
regresses `setTile` by > 10% against `before = 0.202204` still trips
the 10% band (methodology §5), and a regression beyond 0.3 ms trips
the absolute gate.
- **Lesson (extends the workload-fix lesson):** a budget target set
from local-machine runs must be verified with margin on the CI
reference toolchain before it guards master — otherwise the gate is
a coin flip and CI failures stop tracking engine regressions.
- **Tail behavior** of the recorded runs: p99 ≈ 1.05–1.06× the mean;
the outlier run's max (0.300107 ms) is a single-iteration
scheduler/preemption spike — the mean (the gate's metric) is 12%
below the new bar even there.
- **Regression policy:** `budgets.json` `measured = 0.202204`; the 10%
regression band (methodology §5) applies to future re-measurements
on the reference platform; `target = 0.3 ms` is the absolute gate.

## Open items

- Same as `m2-iso-depth-table-workload-fix.md`: **M2-TILE-01** wires
the tilemap height grid to this table; **M2-SORT-01 /
M2-SPRITE-01/02** consume `keyAt` for batched depth-ordered
submission (both already landed).
- If the gate ever fails the reference lane above 0.3 ms, the evidence
indicates reference-runner degradation — the remedy is another PRD
revision through this same process, not a silent target move.
6 changes: 3 additions & 3 deletions docs/concepts/coordinates.md
Original file line number Diff line number Diff line change
Expand Up @@ -201,9 +201,9 @@ chunks (default 16×16 tiles — 256 cells each):
- **Bounded + logged growth** (`ensureChunk`): streamed regions extend
the covered region chunk by chunk, up to a documented cap
(`BudgetExhausted` beyond it).
- **Budget**: 10k dirty cells ≤ 0.2 ms mean (PRD §8.1,
`iso_depthkey_rebuild`) —
[baselines/m2-iso-depth-table.md](../benchmarks/baselines/m2-iso-depth-table.md).
- **Budget**: 10k dirty cells ≤ 0.3 ms mean (PRD §8.1,
`iso_depthkey_rebuild` — re-baselined 2026-10-04) —
[baselines/m2-iso-depth-table-budget-rebaseline.md](../benchmarks/baselines/m2-iso-depth-table-budget-rebaseline.md).

Tiles are grid-locked (M2-TILE-01), so the table's cells agree with
`isoDepthKey` at the same positions bit-for-bit — and across backends
Expand Down
4 changes: 2 additions & 2 deletions roadmap/M2-rendering-2.5d.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,12 +117,12 @@ if M2 slips, and its status is recorded in M2-EXIT-01.
- **Size:** ~150 lines + tests

- [x] **M2-ISO-02 · Depth key table + incremental updates**
- **Refs:** FR-2.2 (precomputed, incremental update), §8.1 (≤ 0.2 ms for 10k dirty cells)
- **Refs:** FR-2.2 (precomputed, incremental update), §8.1 (≤ 0.3 ms for 10k dirty cells — re-baselined from 0.2 ms on 2026-10-04, see Change Log)
- **Depends:** M2-ISO-01
- **Scope:**
- Per-scene-chunk depth key table (tile grid → key), built at scene load (headless-buildable: the table is sim-side data).
- Incremental update API: tile height change → only affected cells recomputed (cell + documented neighborhood radius); no full rebuild.
- Budget test: 10k dirty cells update ≤ 0.2 ms (`budgets.json` entry `iso_depth_rebuild`).
- Budget test: 10k dirty cells update ≤ 0.2 ms, re-baselined to ≤ 0.3 ms on 2026-10-04 (the CI reference lane measured 0.194–0.267 ms on unchanged code — zero-margin gate; `budgets.json` entry `iso_depthkey_rebuild`).
- Zero per-update allocation (tables pre-sized per chunk, growth bounded + logged).
- Unit tests: single-tile edit changes only documented cells; rebuild-from-scratch == incremental result (property test); budget test records baseline.
- **Verify:** `ctest -R iso_depth_table` green; baseline recorded in `docs/benchmarks/baselines/`.
Expand Down
Loading
Loading