Skip to content

chore(scripts): add measure_hero_drift.py for issue #200 - #218

Merged
TMHSDigital merged 1 commit into
mainfrom
feat/200-measure-hero-drift
Sep 23, 2026
Merged

TMHSDigital merged 1 commit into
mainfrom
feat/200-measure-hero-drift

Conversation

@TMHSDigital

Copy link
Copy Markdown
Owner

Summary

Refs #200 (step 1 of its remedy: "measure the whole tree"). This adds scripts/measure_hero_drift.py, which renders each gallery entry's README --output command to lossless PNG and compares the result with the committed hero. Full results are posted on #200.

Type

  • feat
  • fix
  • docs / chore / ci / refactor — authoring tool, no release

What it does

  • For each entry in both gallery.json files, it takes the README's documented blender --background --python X.py -- ... --output ... line, keeps every other flag (e.g. --engine cycles), retargets --output at a PNG and renders.
  • Metrics: mean absolute difference and gt2pct (Committed hero stills do not match what the scripts render #200's definition, mean over channels), plus luma and vs_q90, the committed hero compared with the fresh render re-encoded at webp q90.
  • Verdict: vs_q90 > 2.5% means drifted. Renders are deterministic, but the q90 encode floor ranges from 0.03% to 6.8% by image, so the PNG comparison alone misclassifies. On the full sweep vs_q90 separates cleanly: matches ≤ 1.80%, drifted ≥ 3.83%.
  • --only, --out (default .scratch/hero-drift) and --json. Exit 0 when everything was measured, 3 when a render failed. Not wired into CI: one full render per entry.

Evidence

  • live-run-proven: E:\Blender-Developer-Tools\.scratch\blender-5.2.1-windows-x64\blender.exe (reports Blender 5.2.1 LTS).
    • Full sweep: # 79 measured, … 0 failed.
    • Reproduces Committed hero stills do not match what the scripts render #200's reference numbers exactly: shipping-crate mean_abs 0.01683 / gt2pct 22.45%, crate-stack 0.00330 / 0.50%.
    • Final verdicts on spot checks: bmesh-gear 1.80% → matches, vse-cut-list 3.91% → DRIFTED, crate-stack 0.00% → matches.
    • Determinism: 17 re-rendered entries are pixel-identical to their first render.
  • inspection-only: Windows only. It uses host Python with Pillow and numpy; not tried on Linux.

Checklist

  • Explicit paths only.
  • No content, count or manifest change.
  • No new CI check.
  • DCO Signed-off-by: present.
  • No credentials or emails. The binary path is as the template requests.

🤖 Generated with Claude Code

Renders every gallery entry's README --output command to lossless PNG and
compares it with the committed hero, using #200's metric (it reproduces
that issue's 0.01683 / 22.45% for shipping-crate and 0.00330 / 0.50% for
crate-stack exactly).

One correction to the method: renders are deterministic, but the webp
encode floor is not a constant (0.1-4% of pixels depending on the image),
so a fixed threshold on the PNG comparison misclassifies. The verdict
compares against the fresh render re-encoded at webp q90 instead; on the
full 79-entry sweep that splits the tree with no overlap (matches <= 1.8%,
drifted >= 3.8%).

Authoring tool only; not wired into CI.

Refs #200

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
@TMHSDigital
TMHSDigital merged commit 29edf01 into main Sep 23, 2026
11 checks passed
@TMHSDigital
TMHSDigital deleted the feat/200-measure-hero-drift branch September 23, 2026 01:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant