chore(scripts): add measure_hero_drift.py for issue #200 - #218
Merged
Merged
Conversation
Renders every gallery entry's README --output command to lossless PNG and compares it with the committed hero, using #200's metric (it reproduces that issue's 0.01683 / 22.45% for shipping-crate and 0.00330 / 0.50% for crate-stack exactly). One correction to the method: renders are deterministic, but the webp encode floor is not a constant (0.1-4% of pixels depending on the image), so a fixed threshold on the PNG comparison misclassifies. The verdict compares against the fresh render re-encoded at webp q90 instead; on the full 79-entry sweep that splits the tree with no overlap (matches <= 1.8%, drifted >= 3.8%). Authoring tool only; not wired into CI. Refs #200 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: TMHSDigital <154358121+TMHSDigital@users.noreply.github.com>
8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Refs #200 (step 1 of its remedy: "measure the whole tree"). This adds
scripts/measure_hero_drift.py, which renders each gallery entry's README--outputcommand to lossless PNG and compares the result with the committed hero. Full results are posted on #200.Type
featfixdocs/chore/ci/refactor— authoring tool, no releaseWhat it does
gallery.jsonfiles, it takes the README's documentedblender --background --python X.py -- ... --output ...line, keeps every other flag (e.g.--engine cycles), retargets--outputat a PNG and renders.gt2pct(Committed hero stills do not match what the scripts render #200's definition, mean over channels), plus luma andvs_q90, the committed hero compared with the fresh render re-encoded at webp q90.vs_q90 > 2.5%means drifted. Renders are deterministic, but the q90 encode floor ranges from 0.03% to 6.8% by image, so the PNG comparison alone misclassifies. On the full sweepvs_q90separates cleanly: matches ≤ 1.80%, drifted ≥ 3.83%.--only,--out(default.scratch/hero-drift) and--json. Exit 0 when everything was measured, 3 when a render failed. Not wired into CI: one full render per entry.Evidence
E:\Blender-Developer-Tools\.scratch\blender-5.2.1-windows-x64\blender.exe(reportsBlender 5.2.1 LTS).# 79 measured, … 0 failed.shipping-cratemean_abs 0.01683 / gt2pct 22.45%,crate-stack0.00330 / 0.50%.bmesh-gear1.80% → matches,vse-cut-list3.91% → DRIFTED,crate-stack0.00% → matches.Checklist
Signed-off-by:present.🤖 Generated with Claude Code