Drop-in replacement for Flutter's matchesGoldenFile with tolerance, SSIM, ignore
regions and real text in goldens (one golden for every OS), powered by the gleon Rust comparison engine (the same engine as the
gleon CLI).
Status: proof of concept (v0). Host
flutter teston macOS arm64, Linux x64/arm64 (Ubuntu 26.04+) and Windows x64. No Rust toolchain is needed: prebuilt, checksum-verified native libraries are downloaded once per project.
dev_dependencies:
gleon:
git:
url: https://github.com/gleon-rs/flutter.git
ref: v0.1.0 # a released tag: prebuilt libraries exist for tags onlyOn the first flutter test, the package's build hook downloads the native library for the test
host from the GitHub Release of that version, verifies it against the release's SHA256SUMS.txt
(releases are immutable) and caches it in .dart_tool/hooks_runner/shared/. Later runs, including
offline ones, reuse the cache until flutter clean. HTTPS_PROXY is honored.
Overrides (in the app's pubspec.yaml), e.g. for air-gapped machines:
hooks:
user_defines:
gleon:
ffi_path: path/to/libgleon_ffi.dylib # use this library file as is
# release_url: https://mirror.example/gleon/v0.1.0/ # SHA256SUMS.txt + assets
# gleon_repo: ../gleon # contributors: build from a gleon checkoutSee example/ for a counter app whose tests use gleon.
Replace one import — everything else from flutter_test stays available:
-import 'package:flutter_test/flutter_test.dart';
+import 'package:gleon/gleon.dart';Without extra parameters the behavior is Flutter's, with one difference: exact comparison
of everything but the text of a widget, which never fails by default (see
Real text); goldens resolved relative to the test file,
flutter test --update-goldens rewrites them, nothing is written when tests pass. If a file must keep both imports, add
hide matchesGoldenFile to the flutter_test import.
tolerance takes a GoldenTolerance; dot shorthands keep call sites short:
// Up to 0.5% of pixels may differ.
await expectLater(
find.byType(MyWidget),
matchesGoldenFile(
'goldens/my_widget.png',
tolerance: const .pixel(maxDiffRatio: 0.005),
),
);
// Tolerates rendering noise (anti-aliasing, sub-pixel geometry, color drift),
// still catches changed/missing content, color and alpha changes, blur.
await expectLater(
find.byType(MyWidget),
matchesGoldenFile('goldens/my_widget.png', tolerance: const .ssim()),
);
// Ignore a dynamic region (pixels of the golden PNG, origin top-left).
await expectLater(
find.byType(MyWidget),
matchesGoldenFile(
'goldens/my_widget.png',
ignoreRegions: [const Rect.fromLTWH(0, 0, 400, 80)],
),
);| Tolerance | Meaning |
|---|---|
.exact() (default) |
Every pixel must be identical, like Flutter (text: see textTolerance). |
.pixel({maxDiffRatio = 0.01}) |
At most this fraction (0.0–1.0) of pixels may differ. |
.ssim({minSimilarity = 0.8, colorTolerance = 8}) |
Min local SSIM of every neighborhood (0.0–1.0) and tolerated deviation beyond the local 3x3 envelope (0–255); see below. |
ignoreRegions (rectangles excluded from the comparison) works with every tolerance. Each
variant only has the parameters of its mode, so a setting for the wrong mode cannot be written;
out-of-range values, and ignore regions that are empty, negative or reach beyond 4294967295
pixels, throw an ArgumentError when the matcher is created.
On failure the test message contains the metric and the thresholds, and
failures/<name>_masterImage.png, _testImage.png and _gleonDiff.png are written next to the
test (like Flutter's own failure output). As in Flutter, <name> is only the golden's file name:
goldens of the same name in different folders, compared from tests in one directory, share these
files, and the later failure replaces the earlier one's. The artifacts directory (see
Failure images) keeps every golden apart. Warnings (masks clipped to the
image, a case report that cannot be written) are printed one per line and never fail a test. A
failure caused by a bug in gleon itself, not by the test or its files, ends with a request to
report it.
Suite-wide tolerances per path live in .gleon/gleon.yaml, see below.
flutter test draws every glyph as the same box (the FlutterTest font), so goldens cannot see
text. loadAppFonts() loads the app's real fonts instead: every family of its
FontManifest.json (its own fonts, those of its packages, MaterialIcons) and Roboto from the
Flutter SDK. Call it once from test/flutter_test_config.dart:
import 'dart:async';
import 'package:gleon/gleon.dart';
Future<void> testExecutable(FutureOr<void> Function() testMain) async {
await loadAppFonts();
await testMain();
}Operating systems then lay text out alike (the fonts' Windows line metrics are aligned with the ones macOS and Linux use, so line heights on Windows can differ from the real app there), but rasterize glyphs differently. Measured on CI, 40% or more of the pixels of a 16x16 square of text differ between macOS and Linux, more than a changed digit of the same width does (24%). No tolerance tells them apart, so there are two ways:
- One golden for every OS (the default): text never fails, everything else is compared exactly. Changes that move anything (a longer word, another weight, a shifted line) still fail; changes of text alone (a digit, a color) do not. That is the trade-off.
- Exact text: per-platform goldens, one per OS that renders differently, with the gleon CLI approving each from CI runs (coming to this package; today goldens are single files).
textTolerance (0.0–1.0, default 1) is the largest share of differing pixels allowed in any
16x16 square of text; lower it to compare text, at your own risk on other operating systems:
await expectLater(
find.byType(MyWidget),
matchesGoldenFile(
'goldens/my_widget.png',
textTolerance: 0.1, // text compared: right for goldens of this OS only
),
);The boxes of text come from the render tree: one per line of each paragraph and editable text,
grown by an eighth of the line's height for ink beyond it (diacritics, negative letter spacing,
outlined text), clipped like the text is, without WidgetSpans. On macOS, at 0.1, every
mutation of the package tests fails (a changed digit with 24% of a square, a frame tight around
text with 23%, the rest with 51% or more); at the default 1 only those that move pixels outside
the text do.
A widget (a Finder) is compared as raw pixels: no PNG is encoded unless the golden fails.
Text applies to widgets only, with an exact or pixel tolerance (ArgumentError with ssim; a
warning when an SSIM rule or a byte input leaves a textTolerance unused); ignoreRegions beat
it, for unstable backgrounds under text. Without a textTolerance, the text_tolerance of the
golden's .gleon/gleon.yaml rule applies, else 1.
Compared exactly, so different on other operating systems:
- Text drawn on a canvas (
TextPainterin aCustomPainter, charts): it has no render object to take boxes from. TextStyle.shadows: Flutter's tests turn elevation shadows off (debugDisableShadows), not the shadows of text.
With debugDefaultTargetPlatformOverride set to iOS or macOS, Material's theme asks for Apple's
system fonts, which no test can load: that text stays in Flutter's test font (the same on every
OS, but boxes). Bundle a font and name it in the theme instead.
Suite settings live in the gleon CLI's workspace file —
the very same file, so a project that later adopts the CLI keeps its rules. Each golden belongs to
the nearest directory above it with .gleon/gleon.yaml (its workspace; the working directory of
the test does not matter, and one test run may span several workspaces). The rule of a golden is
resolved by its path relative to that directory, with the CLI's own code; the file is read again
when it changes:
# .gleon/gleon.yaml (create it by hand or with `gleon init`)
required_version: ">=0.1.0"
exclude: "test/goldens/experimental/**"
screenshots:
# The first matching rule wins.
- include: "test/goldens/**/*.png"
mode: ssim
diff: { min_similarity: 0.8, color_tolerance: 8 }
masks:
- path: "**/clock*.png"
zones: [{ x: 0, y: 0, width: "25%", height: 40 }] # pixels or "NN%"
- include: "test/**/*.png"
mode: pixel
diff: { threshold: 0.01 } # max fraction of differing pixels; 0 = exact
text_tolerance: 1 # pixel only, see Real text (1: text never fails, the default)
metrics:
enabled: false
# Where the images of failing goldens are kept (relative to the workspace root).
artifacts: .gleon/runs/latest/artifacts- Priority: the
toleranceargument of a call beats the golden's rule, which beats exact;textTolerancebeats the rule'stext_tolerance. Masks of the rule are added to the call'signoreRegions. - A golden matched by
exclude(or inside a directory the CLI never scans, such asbuild/) or by no rule is compared exactly, like without the file. - Golden paths must be valid gleon test names (
[a-z0-9_.-]segments, case-insensitive), as forgleon stage; an invalid config or name fails the test with the file path and the parser's message. Unknown keys are rejected. required_versionis only checked for syntax here; the CLI enforces it.platformandfallback_platformare accepted and not used by the package yet.- Without
.gleon/gleon.yamleverything behaves like Flutter: exact by default, no metrics, and nothing is written on passing tests.
A failing golden covered by a rule also keeps its images in the workspace's artifacts directory,
<artifacts>/<test name>/golden.png, candidate.png and diff.png (a missing golden keeps its
candidate only, ready to approve with the gleon CLI), and its case report (see
Metrics), with or without metrics; when the golden passes again, the images are
removed and so is the report (with metrics, the report of the pass replaces it). The directory
is .gleon/runs/latest/artifacts unless
artifacts: or the environment variable GLEON_ARTIFACTS_DIR (which beats the file) names
another directory under .gleon/runs/ outside latest/ (any other value fails every golden), so
the images are always ignored by Git and removed by gleon clean. To keep them on a RAM disk,
link .gleon/runs there. The failures/ files next to the test are written as before.
Passing goldens don't show how close they came to failing. A golden covered by a rule writes a
case report to .gleon/runs/latest/cases/<test name>.json when it fails, and with metrics on
for every comparison (overwritten by each run; .gleon/.gitignore ignores runs/ and is created
like gleon init would if it is missing); metrics also print one line:
gleon ✓ test/goldens/swatch.png ssim 0.931 (≥0.800, +0.131) color 5.2 (≤8, +2.8) 12 ms
Turn them on with metrics: {enabled: true} in .gleon/gleon.yaml or with the environment
variable GLEON_METRICS (which beats the file): 1 or true turns them on, 0 or false off
(any case, surrounding spaces ignored), an empty value counts as unset, and any other value fails
every golden. metrics: {console: false} keeps the files and drops the lines. A
report that cannot be written is printed as a warning and never fails the test. A case report
records the golden SHA-256 and size, the candidate size and SHA-256 (a widget's raw pixels have one
only when the golden fails and its PNG is written), the effective tolerance and masks, the outcome
(identical, match, mismatch, dimension_mismatch, error with its kind, updated,
missing for a golden that does not exist yet), the metrics with their headroom to each threshold
(for SSIM min_ssim - min_similarity and color_tolerance - peak_excess), the paths of the
failure images, the test name, platform, Flutter version and timings. The format is a JSON Schema
in the gleon repository (gleon-model/schema/case.v2.json), shared with the CLI.
Reports of different goldens come from different test processes and stay until overwritten, so a
report does not show by itself which run wrote it. Set GLEON_RUN_ID (e.g.
GLEON_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt }} in GitHub Actions; up to 128
letters, digits, ., _ and -) to stamp every report of a run with the same id. A golden that
cannot be read or written is recorded as an io error.
Use them to set tolerances from measurements instead of guesses, e.g. by collecting the reports from CI runs on every OS.
The gleon CLI reads these case reports as one run, no
gleon diff needed:
gleon test -- flutter test # one run: sets GLEON_RUN_ID and GLEON_METRICS=1, records run.json
gleon report markdown # the PR comment of the run; also html, junit, json
gleon dashboard # adds the run to .gleon/history.json and renders dashboard.html
gleon approve # writes the candidates of failed goldens to their PNG filesgleon test passes the test command's exit code through; on Windows it finds flutter.bat.
Without it, the CLI picks the run of GLEON_RUN_ID, else the run of the newest report; reports
without a run id are read together, with a warning that they may mix runs, and after a plain
flutter test without metrics only the failures are there (approving works, the totals don't).
In CI, set GLEON_RUN_ID for the job, upload .gleon/runs/ and pass the downloaded latest/
to --from (gleon report markdown --from <dir>/latest, gleon approve --from <dir>/latest,
one --from per job).
The CLI can also keep goldens out of Git (content-addressed blobs with small JSON manifests) and
manage per-platform baselines for its own screenshots; per-platform goldens for this package are
coming.
A widget golden (a Finder), from the captured frame to the verdict: Flutter encodes the frame
as a PNG and compares it with its comparator (a pass short-cuts on equal bytes); gleon passes the
frame's raw pixels and encodes a PNG only for a failure. Rendering the frame is the same for both
and not measured. Exact, no .gleon/ workspace; Apple M3 Max, macOS, Flutter 3.47.6, mean
latency with bench_press:
| Scenario | Golden | Flutter SDK | gleon | Speedup |
|---|---|---|---|---|
| Passing | 400x300 | 9.49 ms | 0.46 ms | 20x |
| 390x844 | 15.2 ms | 1.25 ms | 12x | |
| 1170x2532 | 87.8 ms | 7.7 ms | 11x | |
| Failing (small change, failure files written) | 400x300 | 43.3 ms | 3.1 ms | 14x |
| 390x844 | 77.0 ms | 6.0 ms | 13x |
Encoding the PNG is most of Flutter's cost. The 390x844 and 1170x2532 passes resolve to 12.1x and 11.4x with 95% confidence intervals within ±1%; gleon's other samples (sub-millisecond, or writing files) varied too much for bench_press to resolve a ratio.
One golden comparison of PNG bytes (byte and ui.Image inputs), Flutter's own comparator
(LocalFileComparator, behind flutter_test's matchesGoldenFile) against gleon, with the same
golden file and candidate bytes. Measured inside flutter test, where golden tests run, with
bench_press (mean latency; every ratio has a 95% confidence
interval within ±8%). Apple M3 Max, macOS, Flutter 3.47.5:
| Scenario | Golden | Flutter SDK | gleon (exact) | Speedup | gleon ssim |
|---|---|---|---|---|---|
| Passing, identical bytes | 400x300 | 191 µs | 34 µs | 5.7x | 34 µs |
| 390x844 | 213 µs | 36 µs | 5.9x | 36 µs | |
| 1170x2532 | 423 µs | 45 µs | 9.5x | 45 µs | |
| Passing, re-encoded (same pixels) | 400x300 | 2.95 ms | 0.53 ms | 5.5x | 0.52 ms |
| 390x844 | 7.67 ms | 1.57 ms | 4.9x | 1.57 ms | |
| 1170x2532 | 56.4 ms | 9.1 ms | 6.2x | 9.3 ms | |
| Failing (small change, failure files written) | 400x300 | 34.6 ms | 1.96 ms | 17.6x | 1.85 ms |
| 390x844 | 61.1 ms | 4.44 ms | 13.8x | 4.07 ms |
Across all cases gleon is 7.7x faster (geometric mean), and tolerant ssim costs no more than
exact. A failing 1170x2532 golden takes Flutter about 0.4 s and gleon 25 ms; that case is beyond
bench_press's 200 ms limit for one operation, so it is a plain timing. Flutter decodes through the
engine but inverts both images and compares them pixel by pixel in Dart, and a failure renders two
diff images and encodes four PNGs; gleon decodes and compares in Rust and writes one diff image.
gleon here runs without a .gleon/ workspace (exact, nothing recorded, like Flutter).
Two gates, both computed only around the pixels that actually differ:
- Envelope gate (full resolution): every pixel of one image must lie within the value range
of the other image's 3x3 neighborhood (premultiplied RGBA, so alpha counts), widened by
colorToleranceplus a share of the local contrast. Re-rasterization and sub-pixel shifts only produce values between neighbors; new content, color and alpha changes don't. Unexplained pixels fail as regions of 3+ pixels or with a strong deviation. - Structural gate (half resolution, as in MS-SSIM): the minimum local SSIM (11x11 Gaussian
window) must reach
minSimilarity, catching structural loss such as blur. The minimum, not the mean, so a small change on a large golden is not averaged away.
The defaults are calibrated on a corpus of benign rendering noise vs. regressions in
gleon-engine/tests/ssim_corpus.rs in the gleon repository; off-the-shelf crates (image-compare, pixelmatch/dify,
butteraugli) could not separate that corpus with any threshold.
ssimfails when glyphs move by half a pixel or more — typical of different operating systems' font engines. For one golden on every OS, use real fonts withexactorpixelinstead (text never fails by default, see Real text).ssimcan pass a low-contrast color change of a one-pixel line. Useexact/pixelwhere every pixel matters.- Only Flutter's default
LocalFileComparatoris supported as the underlying golden store. - Web (
--platform chrome) and on-device tests are not supported: the native library does not exist there, and the first comparison fails with a message saying so.
One package with a hard internal boundary:
lib/gleon.dart exports only (flutter_test minus matchesGoldenFile, plus the gleon API)
lib/src/core/ plain Dart: never imports Flutter (dart:ui, package:flutter*)
compare/ GoldenTolerance, PixelRegion (the call's tolerances,
masks and text regions)
config/ GleonIntegration (who calls), GleonSession (the native session)
native/ @Native leaf bindings, NativeEngine (ABI check, packing), verdicts,
error kinds
hook/ native targets, release download, source build, atomic writes (used by
hook/build.dart)
lib/src/flutter/ the Flutter layer: matchesGoldenFile, the widget capture and its text
regions, the comparator, loadAppFonts, FlutterSession (this package
as an integration)
hook/build.dart thin build hook on top of lib/src/core/hook/
bin/ maintainer scripts (dart:io + crypto only), run with plain `dart`
The native engine (gleon-ffi) does the whole job of a golden: it finds the workspace, resolves
the .gleon/gleon.yaml rule, reads and compares the golden, and writes failure artifacts, case
reports and updated goldens; this package passes facts (paths, the raw pixels or PNG, the call's
tolerance, masks and text regions) and shows the verdict and texts it gets back. The C types and bindings of that contract
live together in lib/src/core/native/gleon_ffi.dart.
Rule: code in lib/src/core/ must not import Flutter, so it stays usable from dart test,
the build hook and other SDKs later; everything Flutter-specific (Rect, matchers, comparators)
lives in lib/src/flutter/ and converts to core types at the boundary. DCM enforces the rule
(avoid-banned-imports in analysis_options.yaml).
Every native call is an isLeaf: true call. That allows passing the PNG zero-copy via
Uint8List.address (and a call's strings as one UTF-8 buffer plus a Uint32List of their
lengths) and returning {ptr, len} slices by value, so the Dart side never allocates native memory
(no package:ffi, no malloc/free pairs; only the session is released by a NativeFinalizer).
The price is that the isolate group cannot reach a GC safepoint while a comparison runs
(milliseconds for typical goldens, seconds for very large SSIM comparisons). For tests that is the right trade-off: the test
awaits the result anyway, and flutter test parallelizes across processes. Apps that must stay
responsive would instead copy into malloced memory and make non-leaf calls in Isolate.run.
dart format --set-exit-if-changed .
flutter analyze --fatal-infos # also in example/
dcm analyze . # DCM 1.39.2, also in example/
flutter test # also in example/Case reports written by the tests are validated against case.v2.json when a gleon checkout sits
next to this repository (../gleon, as in CI); without it that check is skipped.
analysis_options.yaml is the single, strict configuration (analyzer lints plus DCM presets);
every disabled or narrowed rule carries its reason.
benchmark/ holds bench_press benchmarks that need the
Flutter engine (dart:ui), so they run under flutter test, not dart run bench_press run:
flutter test benchmark/ # ~3 min; results in build/benchmark/*.json
dart run bench_press report --from-json build/benchmark/golden_comparison.json
dart run bench_press report --from-json build/benchmark/widget_capture.jsonTo check a change for regressions, keep the JSON of a run before it and compare
(dart run bench_press diff old.json build/benchmark/golden_comparison.json). Run on an idle
machine: bench_press withholds ratios of unstable samples ("unresolved"). BENCH_PRESS_ARGS
passes options (--validate for a smoke run, as CI does; --trials 30). The manual
Benchmark workflow runs them on every supported host and puts the report into each job's
summary.
The native code (gleon-engine, gleon-model, gleon-ffi) lives in the
gleon repository; native/gleon_ref pins the commit CI builds.
Build the library for this machine into native/<target>/ (the hook prefers it over downloading):
dart bin/build_native.dart # host; --target all cross-builds on macOSUse plain dart, not dart run: dart run executes the build hook first. Cross builds need
cargo-zigbuild (Linux) and cargo-xwin (Windows); CI builds every target natively. A build
records the gleon commit it was made from, and the hook uses it only while native/gleon_ref pins
that commit. To try an unmerged engine change on top of the pinned commit, build with
--allow-dirty (and rebuild after every change: the hook cannot tell two dirty builds apart); to
test another checkout, use the gleon_repo user-define.
Releasing: bump version in pubspec.yaml and CHANGELOG.md, update native/gleon_ref if the
engine changed, and push the tag vX.Y.Z. The release workflow builds and tests all targets on
their own OS, attaches the libraries and SHA256SUMS.txt to an immutable GitHub Release, and then
verifies the download path on every OS.