Skip to content

fix: correct the inverted EEVEE engine id and close the coverage gap that hid it - #204

Merged
TMHSDigital merged 6 commits into
mainfrom
fix/engine-id-inversion-and-render-coverage
Sep 22, 2026
Merged

TMHSDigital merged 6 commits into
mainfrom
fix/engine-id-inversion-and-render-coverage

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

A live bug, the coverage gap that hid it, and two convention gaps the last loop surfaced. Six commits, one per item.

Binaries: .scratch/blender-4.5.11-windows-x64/blender.exe (4.5.11 LTS), .scratch/blender-5.1.2-windows-x64/blender.exe (5.1.2), .scratch/blender-5.2.1-windows-x64/blender.exe (5.2.1 LTS).


Part A1/A2 — the inverted engine id, twice

showcase/shipping-crate carried eevee_engine_id() keyed on >= (4, 2, 0), returning BLENDER_EEVEE_NEXT for everything above it. Blender 5.x reclaimed the plain BLENDER_EEVEE id and no longer offers BLENDER_EEVEE_NEXT. Reproduced verbatim on 5.1.2:

TypeError: bpy_struct: item.attr = val: enum "BLENDER_EEVEE_NEXT" not found
in ('BLENDER_EEVEE', 'BLENDER_WORKBENCH', 'CYCLES')
FATAL: bpy_struct: item.attr = val: enum "BLENDER_EEVEE_NEXT" not found
in ('BLENDER_EEVEE', 'BLENDER_WORKBENCH', 'CYCLES')
exit=1

After the fix, on 5.1.2: shipping-crate exit=0, iron-cauldron exit=0.

The audit found a second copy. showcase/iron-cauldron had the identical inversion. It was not in the brief; it was found by auditing every mapping rather than grepping for the one known file.

Audit method

Not a regex. Every expression that yields a BLENDER_EEVEE* id was extracted and evaluated twice, against a stubbed bpy.app.version of (4, 5, 11) and (5, 2, 1), then compared against the canonical mapping from examples/swatch-grid:

'BLENDER_EEVEE' if bpy.app.version >= (5, 0, 0) else 'BLENDER_EEVEE_NEXT'

A line-oriented regex is not good enough here: it reads the if/return form in snippets/version-branch-skeleton.py as two unrelated branches and flags a file that is correct.

The table

84 mappings across 83 files. Everything not listed below evaluates correctly on both eras.

File Line Verdict Action
showcase/shipping-crate/shipping_crate.py 104 INVERTED — 4.5→NEXT ok, 5.2→NEXT wrong fixed (7cce965)
showcase/iron-cauldron/iron_cauldron.py 118 INVERTED — 4.5→NEXT ok, 5.2→NEXT wrong fixed (b0a23c4)
examples/swatch-grid/swatch_grid.py 268 deliberately inverted (wrong = …) marked # engine-id-exempt:
examples/turntable/turntable.py 138 deliberately inverted (wrong = …) marked # engine-id-exempt:
examples/gn-sdf-remesh/gn_sdf_remesh.py 119 deliberately inverted (wrong = …) marked # engine-id-exempt:
snippets/version-branch-skeleton.py 29–31 correct (if/return form, not a ternary) none
tests/smoke/run_smoke.py 44 correct none
tests/smoke/tmpl_render.py 25 correct, inline at the assignment none
all other 76 — correct none

The three deliberate inversions are examples that assert the other era's id is rejected by the running build. They now say so in source rather than relying on the reader.

Both fixes carry the provenance header the copied-helper convention already uses for the hygiene combinatorics, naming examples/swatch-grid.get_eevee_engine_id as the source.


Report item 2 — a committed still is inconsistent with its code

Checked because the brief asked. The answer is yes, and it is broader than these two pieces. Filed as #200; not fixed here.

Method: render --output to lossless PNG on 5.2.1, compare against the committed hero.

Control — how much the webp encode alone accounts for:

shipping-crate  PNG vs the script's own q90 webp
  mean_abs_pixel_diff=0.00285  pixels_differing_gt_2pct=4296 (0.47%)

So ~0.5 % is the noise floor. Measured:

Piece mean abs diff pixels > 2 % Verdict
crate-stack (hero rendered yesterday, direct to webp) 0.00330 0.50 % matches
wooden-ladder 0.01260 12.39 % drifted
iron-cauldron 0.01105 10.22 % drifted
anvil 0.01361 16.81 % drifted
shipping-crate 0.01683 22.45 % drifted

It is not renderer era. The obvious hypothesis — old stills rendered on 4.5 with EEVEE Next — is ruled out. The same piece rendered on both binaries agrees with itself to within the noise floor:

shipping-crate  4.5.11 vs 5.2.1 render
  mean_abs_pixel_diff=0.00133  pixels_differing_gt_2pct=576 (0.06%)

Committed heroes are also systematically brighter (mean luminance +0.0104 on shipping-crate) and ~1.85× larger on disk. That points at a separate PNG→webp conversion step at a different quality and colour handling — .scratch/ still carries to_webp.py, to_webp_h.py, ads_webp.py.

This does not undermine the engine-id fix: the TypeError is independently reproducible. Both are true.


Part A3 — coverage, with the numbers

Measured render cost

showcase/shipping-crate --output, whole process including the check path:

Engine 4.5.11 5.1.2 5.2.1 Render-only delta
EEVEE (default) 2 914 ms 4 004 ms 3 515 ms 1 152 / 1 889 / 1 890 ms
Cycles, 32 samples CPU 14 449 ms — 12 559 ms ~11 000 ms

Check-only baseline: 1 762 / 2 115 / 1 625 ms.

Chosen: option 2, the static mapping check — and why not the canary

The cheap canary does not cover the class, and the covering canary is not cheap:

  • EEVEE at ~1.9 s/version is nearly free but CI will not run it. blender-smoke.yml says so itself: "Cycles (CPU) so this is reliable on GPU-less runners", and run_smoke.py: "an EEVEE GPU render aborts the process (no EGL)".
  • Cycles is what CI can run, and Cycles never touches the EEVEE id — a Cycles canary would not have caught either inversion. It costs ~11 s/version locally for one piece against a 20.9 s showcase suite, and CI runners are slower.

So a render canary buys nothing for the motivating class at either price. tests/check_engine_id.py covers it completely, for free, across all 84 mappings — including the 82 a canary on one piece would never touch. The remainder of the render path is #201, with these numbers.

No blender-smoke.yml change. The checker runs in validate.yml.

A3 canary, verbatim

Inversion reintroduced in shipping-crate:

1 engine-id mapping(s) disagree with the canonical one (BLENDER_EEVEE on 5.0+, BLENDER_EEVEE_NEXT on 4.2-4.5):
  ERROR: showcase/shipping-crate/shipping_crate.py:103 eevee_engine_id() 5.2.1->BLENDER_EEVEE_NEXT (want BLENDER_EEVEE)

A deliberately inverted mapping must carry '# engine-id-exempt: <reason>' on its own line.
CANARY RC=1

Reverted:

  exempt  examples/swatch-grid/swatch_grid.py:268 conditional - the wrong-era id this example asserts is rejected
  exempt  examples/turntable/turntable.py:138 conditional - the wrong-era id this example asserts is rejected
  exempt  examples/gn-sdf-remesh/gn_sdf_remesh.py:119 conditional - the wrong-era id this example asserts is rejected
engine-id checks passed: 84 mapping(s) across 83 file(s), 3 exempt.
RC=0

Part B — both gates required, each on its own evidence

Applied the rule in the brief rather than a preference.

Contact sheets: required. They have caught a real defect in showcase work. The first sheets for crate-stack and stone-archway showed both wedge pools reading as cool grey bands instead of the warm pool the house style calls for; both pieces were relit as a direct result. Nothing else in the pipeline puts a still beside its peers.

Asset sheets: also required — the brief's escape clause applies. Reading the definition shows they catch something budgets cannot. docs/VISUAL-STYLE.md § Asset quality scopes the gate to "asset-type examples (game props, kits)", and showcase is entirely game props. Budgets measure geometry conformance; the contact sheet measures staged presentation; neither removes the scene. The recorded evidence is decisive:

socket-attach-points passed all three on its first draft (edge90 0.000, 10 materials, no default names) — a model that was then judged bad by eye and rebuilt from scratch. The floors scored the bevels, not the design.

A showcase budget cannot see that failure. The asset sheet is the only gate here that can, so marking it "not applicable" would have been the comfortable answer rather than the correct one.

Also recorded: check_asset_quality returns 11, which showcase numbering already spends on the collider-triangle ceiling. Call sites remap rather than sharing a code.

Not retrofitted here. 26 of 28 pieces lack a contact sheet, 28 of 28 lack an asset sheet — #202, which is explicitly blocked on #200, because building a contact sheet from a drifted hero calibrates against the wrong image.


Part C — falsifiers must fail the budget they target

Documented in CONTRIBUTING.md and showcase/README.md, with --off-circle replacing --flat-arch as the worked example, and the rule stated plainly: fix a collision by changing the model or the falsifier, never by widening a band.

tests/check_falsifier_targets.py, one checker in two modes rather than a parallel harness:

  • static (default, no Blender, runs in validate.yml) — flags are real argparse flags, declared exit codes exist in that piece's exit-code table, every falsifier names a target budget, no falsifier-shaped flag is undocumented. 160 falsifiers across 25 pieces.
  • --run BLENDER — executes each falsifier and asserts the observed exit equals the declared one. This is the mode that catches an ill-aimed falsifier.

Table parsing is header-aware, not positional: wooden-ladder's four-column Falsifier | Budget violated | Exit | Measured failure is as valid as the three-column form, and a positional regex read it as having no exit column at all.

One design dropped on evidence. Requiring each exit-code row to name its falsifier would have forced a format change across 23 READMEs to restate what the falsifier table's own target column already carries. Replaced with the structural requirement that every falsifier row names a target budget — which the existing tables already satisfy.

Applied to existing pieces

25 of 28. Three — cart, hay-bale, stone-well — predate the convention and have no table; they are named in the checker, reported on every run, and filed as #203. A new piece without a table is an error, not an entry in that list.

Full runtime sweep, 5.2.1

RC=0  wall=276s
160 ok, 0 MISMATCH
falsifier-target checks passed: 320 falsifier(s) across 25 showcase piece(s).

Every one of the 160 declared exit codes is the code the falsifier actually produces. No existing piece was mis-aimed.

Part C canary, verbatim

--off-circle re-declared as targeting the bbox budget (exit 8):

  ok        stone-archway --float-pier: want 16, got 16
  ok        stone-archway --lift-z: want 16, got 16
  MISMATCH  stone-archway --off-circle: want 8, got 19
  ok        stone-archway --sink-keystone: want 17, got 17
  ok        stone-archway --skip-decimate: want 9, got 9
  ok        stone-archway --stray-vert: want 15, got 15
  ok        stone-archway --wide-mortar: want 18, got 18
CANARY RC=1

Reverted:

  ok        stone-archway --off-circle: want 19, got 19
falsifier-target checks passed: 167 falsifier(s) across 25 showcase piece(s).
RC=0

Constraints honoured

No budget band, exit-code value, or witness changed. gallery_framing.py, LICENSE, VERSION, CHANGELOG.md, release.yml and pages.yml untouched. blender-smoke.yml untouched — the chosen coverage did not need it.

Issues filed

# Title
#200 Committed hero stills do not match what the scripts render
#201 The --output render path has no CI coverage beyond the engine id
#202 Backfill contact sheets (26) and asset sheets (28) for existing showcase pieces
#203 Backfill falsifier tables for cart, hay-bale and stone-well

🤖 Generated with Claude Code

TMHSDigital and others added 6 commits September 22, 2026 07:07
eevee_engine_id() keyed on >= (4, 2, 0) and returned 'BLENDER_EEVEE_NEXT'
for every version at or above it. Blender 5.x reclaimed the plain
'BLENDER_EEVEE' id and no longer offers 'BLENDER_EEVEE_NEXT', so the
--output path died on 5.1 and 5.2:

  TypeError: bpy_struct: item.attr = val: enum "BLENDER_EEVEE_NEXT" not
  found in ('BLENDER_EEVEE', 'BLENDER_WORKBENCH', 'CYCLES')

Smoke stayed green because smoke never passes --output.

The correct mapping is already witnessed in the tree: examples/swatch-grid
asserts both the chosen id and the other era's id against the running
build. Use its form and name it as the source, per the copied-not-imported
convention the hygiene helpers already follow.

shipping-crate is the pilot piece the other 27 were modelled on, which is
how one inverted ternary propagated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
Second copy of the same inversion, found by auditing every engine-id
mapping in the tree rather than by grepping for the one known bad file.

Audit method: extract every expression that yields a BLENDER_EEVEE* id and
evaluate it twice, against a stubbed bpy.app.version of (4, 5, 11) and
(5, 2, 1), then compare both results against the canonical mapping. 83
expressions across 85 files; two returned 'BLENDER_EEVEE_NEXT' for 5.2 —
shipping-crate and this one. Everything else was already correct.

Like shipping-crate, its --output path raised TypeError on 5.1 and 5.2 and
smoke could not see it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
The render path is invisible to CI by construction: smoke runs the default
check path, and --output is exercised only at authoring time, on one
version, by the author. That is how an inverted engine id shipped twice.

tests/smoke/run_smoke.py already asserts the mapping against the running
build, but against its own copy of the helper. It cannot see the other
copies, and showcase and example scripts are standalone by convention, so
each carries its own. Eighty-four mappings across eighty-three files.

This checker closes that gap at zero render cost. It walks examples/,
showcase/, templates/, snippets/, scripts/ and tests/, and for every
function that resolves an EEVEE id it compiles and *calls* the function
twice against a stubbed bpy.app.version of (4, 5, 11) and (5, 2, 1);
conditional expressions yielding an id are evaluated the same way.

AST rather than regex, deliberately. A line-oriented regex misreads the
if/return form as two independent branches and flags
snippets/version-branch-skeleton.py, which is correct. Calling the code is
the only reading that cannot be fooled by formatting.

Resolver detection is narrow on purpose: only a body of returns and
version branches qualifies, so a main() that merely mentions the id is not
compiled and does not need the module's imports.

Three deliberately inverted mappings exist — swatch-grid, turntable and
gn-sdf-remesh each assert that the *other* era's id is rejected by the
build. Those now carry an explicit `# engine-id-exempt:` marker, so the
intent is stated in the source rather than inferred by the checker.

Canary-proven. With the shipping-crate inversion reintroduced:

  ERROR: showcase/shipping-crate/shipping_crate.py:103 eevee_engine_id()
         5.2.1->BLENDER_EEVEE_NEXT (want BLENDER_EEVEE)
  RC=1

Reverted:

  engine-id checks passed: 84 mapping(s) across 83 file(s), 3 exempt.
  RC=0

Runs in validate.yml beside the other tests/ checkers. No smoke cost, no
blender-smoke.yml change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
showcase/README.md said nothing about either gate, and no showcase piece
has ever shipped one — all 44 contact sheets and all 4 asset sheets in
docs/gallery/ belong to examples. The gates live under CLAUDE.md
"Quality Gates for Example Runs", so the next author had to guess.

Both are now required, each on its own evidence rather than on symmetry.

Contact sheets have caught a real defect in showcase work: the first
sheets for crate-stack and stone-archway showed both wedge pools reading
as cool grey bands instead of the warm pool the house style calls for,
and both pieces were relit. Nothing else in the pipeline puts the still
beside its peers.

Asset sheets cover what nothing else here covers. Showcase is entirely
game props, which is exactly the scope docs/VISUAL-STYLE.md names for the
gate. Budgets measure geometry conformance; the contact sheet measures
staged presentation; neither removes the scene. The recorded evidence is
socket-attach-points, which passed every measurable floor on its first
draft (edge90 0.000, ten materials, no default names) and was still
judged bad by eye and rebuilt — "the floors scored the bevels, not the
design". That is precisely the failure a showcase budget cannot see.

Also records the exit-code collision: check_asset_quality returns 11,
which showcase numbering already spends on the collider ceiling, so call
sites remap rather than sharing a code.

Existing pieces are not retrofitted here; the backlog is filed separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
Five falsifiers in the previous showcase run had to be retuned because
they tripped an earlier check instead of the budget they were built for.
A falsifier that exits non-zero on an earlier budget is red for the wrong
reason and proves nothing about its target, and the exit code is right
there to catch it.

Documents the rule in CONTRIBUTING.md and showcase/README.md, with
--off-circle replacing --flat-arch as the worked example: a lintel is
0.62 m shorter than the arch, so it failed the bounding box (exit 8) and
never reached the intrados circle fit (exit 19). Widening the bbox
tolerance would have destroyed a real budget to rescue a bad falsifier.
The rule is explicit that the fix is the model or the falsifier, never
the band.

tests/check_falsifier_targets.py enforces it in two modes, one checker
rather than a parallel harness:

- static (default, no Blender, runs in validate.yml): reads each piece's
  falsifier table and asserts the flags are real argparse flags, the
  declared exit codes appear in that piece's exit-code table, every
  falsifier names a target budget, and no falsifier-shaped flag in the
  script is undocumented. 160 falsifiers across 25 pieces.
- --run BLENDER: executes each falsifier and asserts the observed exit
  equals the declared one. This is the mode that catches an ill-aimed
  falsifier. One Blender launch per falsifier, about 150 s for the tree
  on one version, so it is an authoring and cron tool rather than a
  per-PR smoke step. No blender-smoke.yml change.

Table parsing is header-aware, not positional: wooden-ladder's four-column
"Falsifier | Budget violated | Exit | Measured failure" is as valid as the
three-column form, and an earlier positional regex silently read it as
having no exit column at all.

The bidirectional variant — requiring each exit-code row to name its
falsifier — was tried and dropped. It would have forced a format change
across 23 READMEs to state something the falsifier table's own target
column already carries.

cart, hay-bale and stone-well predate the convention and have no table.
They are listed explicitly in the checker and reported on every run; a new
piece without a table is an error rather than an entry in that list.

Canary-proven at runtime. Declaring --off-circle as targeting the bbox
budget:

  MISMATCH  stone-archway --off-circle: want 8, got 19
  RC=1

Reverted:

  ok        stone-archway --off-circle: want 19, got 19
  falsifier-target checks passed: 167 falsifier(s) across 25 pieces.
  RC=0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
The estimate said roughly 150 s. Running the full sweep measured 276 s for
160 falsifiers across 25 showcase pieces on Blender 5.2.1, with zero
mismatches — every declared exit code is the one the falsifier actually
produces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
@TMHSDigital TMHSDigital added the needs-5.1 Opt-in Blender 5.1 smoke on this PR. Default matrix stays 4.5 + 5.2. Auto-label will not apply this. label Sep 22, 2026
@github-actions github-actions Bot added examples Runnable smoke-gated examples under examples/ showcase Budget-conformance props under showcase/ documentation Improvements or additions to documentation ci labels Sep 22, 2026
@TMHSDigital
TMHSDigital merged commit c4cd7b4 into main Sep 22, 2026
13 checks passed
@TMHSDigital
TMHSDigital deleted the fix/engine-id-inversion-and-render-coverage branch September 22, 2026 11:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci documentation Improvements or additions to documentation examples Runnable smoke-gated examples under examples/ needs-5.1 Opt-in Blender 5.1 smoke on this PR. Default matrix stays 4.5 + 5.2. Auto-label will not apply this. showcase Budget-conformance props under showcase/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant