Skip to content

Repository files navigation

gleon (Flutter)

Drop-in replacement for Flutter's matchesGoldenFile with tolerance, SSIM, ignore regions and real text in goldens (one golden for every OS), powered by the gleon Rust comparison engine (the same engine as the gleon CLI).

Status: proof of concept (v0). Host flutter test on macOS arm64, Linux x64/arm64 (Ubuntu 26.04+) and Windows x64. No Rust toolchain is needed: prebuilt, checksum-verified native libraries are downloaded once per project.

Install

dev_dependencies:
  gleon:
    git:
      url: https://github.com/gleon-rs/flutter.git
      ref: v0.1.0 # a released tag: prebuilt libraries exist for tags only

On the first flutter test, the package's build hook downloads the native library for the test host from the GitHub Release of that version, verifies it against the release's SHA256SUMS.txt (releases are immutable) and caches it in .dart_tool/hooks_runner/shared/. Later runs, including offline ones, reuse the cache until flutter clean. HTTPS_PROXY is honored.

Overrides (in the app's pubspec.yaml), e.g. for air-gapped machines:

hooks:
  user_defines:
    gleon:
      ffi_path: path/to/libgleon_ffi.dylib # use this library file as is
      # release_url: https://mirror.example/gleon/v0.1.0/ # SHA256SUMS.txt + assets
      # gleon_repo: ../gleon               # contributors: build from a gleon checkout

See example/ for a counter app whose tests use gleon.

Migrate

Replace one import — everything else from flutter_test stays available:

-import 'package:flutter_test/flutter_test.dart';
+import 'package:gleon/gleon.dart';

Without extra parameters the behavior is Flutter's, with one difference: exact comparison of everything but the text of a widget, which never fails by default (see Real text); goldens resolved relative to the test file, flutter test --update-goldens rewrites them, nothing is written when tests pass. If a file must keep both imports, add hide matchesGoldenFile to the flutter_test import.

Tolerance

tolerance takes a GoldenTolerance; dot shorthands keep call sites short:

// Up to 0.5% of pixels may differ.
await expectLater(
  find.byType(MyWidget),
  matchesGoldenFile(
    'goldens/my_widget.png',
    tolerance: const .pixel(maxDiffRatio: 0.005),
  ),
);

// Tolerates rendering noise (anti-aliasing, sub-pixel geometry, color drift),
// still catches changed/missing content, color and alpha changes, blur.
await expectLater(
  find.byType(MyWidget),
  matchesGoldenFile('goldens/my_widget.png', tolerance: const .ssim()),
);

// Ignore a dynamic region (pixels of the golden PNG, origin top-left).
await expectLater(
  find.byType(MyWidget),
  matchesGoldenFile(
    'goldens/my_widget.png',
    ignoreRegions: [const Rect.fromLTWH(0, 0, 400, 80)],
  ),
);
Tolerance Meaning
.exact() (default) Every pixel must be identical, like Flutter (text: see textTolerance).
.pixel({maxDiffRatio = 0.01}) At most this fraction (0.0–1.0) of pixels may differ.
.ssim({minSimilarity = 0.8, colorTolerance = 8}) Min local SSIM of every neighborhood (0.0–1.0) and tolerated deviation beyond the local 3x3 envelope (0–255); see below.

ignoreRegions (rectangles excluded from the comparison) works with every tolerance. Each variant only has the parameters of its mode, so a setting for the wrong mode cannot be written; out-of-range values, and ignore regions that are empty, negative or reach beyond 4294967295 pixels, throw an ArgumentError when the matcher is created.

On failure the test message contains the metric and the thresholds, and failures/<name>_masterImage.png, _testImage.png and _gleonDiff.png are written next to the test (like Flutter's own failure output). As in Flutter, <name> is only the golden's file name: goldens of the same name in different folders, compared from tests in one directory, share these files, and the later failure replaces the earlier one's. The artifacts directory (see Failure images) keeps every golden apart. Warnings (masks clipped to the image, a case report that cannot be written) are printed one per line and never fail a test. A failure caused by a bug in gleon itself, not by the test or its files, ends with a request to report it.

Suite-wide tolerances per path live in .gleon/gleon.yaml, see below.

Real text: one golden for every OS

flutter test draws every glyph as the same box (the FlutterTest font), so goldens cannot see text. loadAppFonts() loads the app's real fonts instead: every family of its FontManifest.json (its own fonts, those of its packages, MaterialIcons) and Roboto from the Flutter SDK. Call it once from test/flutter_test_config.dart:

import 'dart:async';

import 'package:gleon/gleon.dart';

Future<void> testExecutable(FutureOr<void> Function() testMain) async {
  await loadAppFonts();
  await testMain();
}

Operating systems then lay text out alike (the fonts' Windows line metrics are aligned with the ones macOS and Linux use, so line heights on Windows can differ from the real app there), but rasterize glyphs differently. Measured on CI, 40% or more of the pixels of a 16x16 square of text differ between macOS and Linux, more than a changed digit of the same width does (24%). No tolerance tells them apart, so there are two ways:

  1. One golden for every OS (the default): text never fails, everything else is compared exactly. Changes that move anything (a longer word, another weight, a shifted line) still fail; changes of text alone (a digit, a color) do not. That is the trade-off.
  2. Exact text: per-platform goldens, one per OS that renders differently, with the gleon CLI approving each from CI runs (coming to this package; today goldens are single files).

textTolerance (0.0–1.0, default 1) is the largest share of differing pixels allowed in any 16x16 square of text; lower it to compare text, at your own risk on other operating systems:

await expectLater(
  find.byType(MyWidget),
  matchesGoldenFile(
    'goldens/my_widget.png',
    textTolerance: 0.1, // text compared: right for goldens of this OS only
  ),
);

The boxes of text come from the render tree: one per line of each paragraph and editable text, grown by an eighth of the line's height for ink beyond it (diacritics, negative letter spacing, outlined text), clipped like the text is, without WidgetSpans. On macOS, at 0.1, every mutation of the package tests fails (a changed digit with 24% of a square, a frame tight around text with 23%, the rest with 51% or more); at the default 1 only those that move pixels outside the text do.

A widget (a Finder) is compared as raw pixels: no PNG is encoded unless the golden fails. Text applies to widgets only, with an exact or pixel tolerance (ArgumentError with ssim; a warning when an SSIM rule or a byte input leaves a textTolerance unused); ignoreRegions beat it, for unstable backgrounds under text. Without a textTolerance, the text_tolerance of the golden's .gleon/gleon.yaml rule applies, else 1.

Compared exactly, so different on other operating systems:

  • Text drawn on a canvas (TextPainter in a CustomPainter, charts): it has no render object to take boxes from.
  • TextStyle.shadows: Flutter's tests turn elevation shadows off (debugDisableShadows), not the shadows of text.

With debugDefaultTargetPlatformOverride set to iOS or macOS, Material's theme asks for Apple's system fonts, which no test can load: that text stays in Flutter's test font (the same on every OS, but boxes). Bundle a font and name it in the theme instead.

Configuration with .gleon/gleon.yaml

Suite settings live in the gleon CLI's workspace file — the very same file, so a project that later adopts the CLI keeps its rules. Each golden belongs to the nearest directory above it with .gleon/gleon.yaml (its workspace; the working directory of the test does not matter, and one test run may span several workspaces). The rule of a golden is resolved by its path relative to that directory, with the CLI's own code; the file is read again when it changes:

# .gleon/gleon.yaml (create it by hand or with `gleon init`)
required_version: ">=0.1.0"

exclude: "test/goldens/experimental/**"

screenshots:
  # The first matching rule wins.
  - include: "test/goldens/**/*.png"
    mode: ssim
    diff: { min_similarity: 0.8, color_tolerance: 8 }
    masks:
      - path: "**/clock*.png"
        zones: [{ x: 0, y: 0, width: "25%", height: 40 }] # pixels or "NN%"
  - include: "test/**/*.png"
    mode: pixel
    diff: { threshold: 0.01 } # max fraction of differing pixels; 0 = exact
    text_tolerance: 1 # pixel only, see Real text (1: text never fails, the default)

metrics:
  enabled: false

# Where the images of failing goldens are kept (relative to the workspace root).
artifacts: .gleon/runs/latest/artifacts
  • Priority: the tolerance argument of a call beats the golden's rule, which beats exact; textTolerance beats the rule's text_tolerance. Masks of the rule are added to the call's ignoreRegions.
  • A golden matched by exclude (or inside a directory the CLI never scans, such as build/) or by no rule is compared exactly, like without the file.
  • Golden paths must be valid gleon test names ([a-z0-9_.-] segments, case-insensitive), as for gleon stage; an invalid config or name fails the test with the file path and the parser's message. Unknown keys are rejected.
  • required_version is only checked for syntax here; the CLI enforces it. platform and fallback_platform are accepted and not used by the package yet.
  • Without .gleon/gleon.yaml everything behaves like Flutter: exact by default, no metrics, and nothing is written on passing tests.

Failure images

A failing golden covered by a rule also keeps its images in the workspace's artifacts directory, <artifacts>/<test name>/golden.png, candidate.png and diff.png (a missing golden keeps its candidate only, ready to approve with the gleon CLI), and its case report (see Metrics), with or without metrics; when the golden passes again, the images are removed and so is the report (with metrics, the report of the pass replaces it). The directory is .gleon/runs/latest/artifacts unless artifacts: or the environment variable GLEON_ARTIFACTS_DIR (which beats the file) names another directory under .gleon/runs/ outside latest/ (any other value fails every golden), so the images are always ignored by Git and removed by gleon clean. To keep them on a RAM disk, link .gleon/runs there. The failures/ files next to the test are written as before.

Metrics

Passing goldens don't show how close they came to failing. A golden covered by a rule writes a case report to .gleon/runs/latest/cases/<test name>.json when it fails, and with metrics on for every comparison (overwritten by each run; .gleon/.gitignore ignores runs/ and is created like gleon init would if it is missing); metrics also print one line:

gleon ✓ test/goldens/swatch.png  ssim 0.931 (≥0.800, +0.131)  color 5.2 (≤8, +2.8)  12 ms

Turn them on with metrics: {enabled: true} in .gleon/gleon.yaml or with the environment variable GLEON_METRICS (which beats the file): 1 or true turns them on, 0 or false off (any case, surrounding spaces ignored), an empty value counts as unset, and any other value fails every golden. metrics: {console: false} keeps the files and drops the lines. A report that cannot be written is printed as a warning and never fails the test. A case report records the golden SHA-256 and size, the candidate size and SHA-256 (a widget's raw pixels have one only when the golden fails and its PNG is written), the effective tolerance and masks, the outcome (identical, match, mismatch, dimension_mismatch, error with its kind, updated, missing for a golden that does not exist yet), the metrics with their headroom to each threshold (for SSIM min_ssim - min_similarity and color_tolerance - peak_excess), the paths of the failure images, the test name, platform, Flutter version and timings. The format is a JSON Schema in the gleon repository (gleon-model/schema/case.v2.json), shared with the CLI.

Reports of different goldens come from different test processes and stay until overwritten, so a report does not show by itself which run wrote it. Set GLEON_RUN_ID (e.g. GLEON_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt }} in GitHub Actions; up to 128 letters, digits, ., _ and -) to stamp every report of a run with the same id. A golden that cannot be read or written is recorded as an io error.

Use them to set tolerances from measurements instead of guesses, e.g. by collecting the reports from CI runs on every OS.

Reports, history and approvals with the gleon CLI

The gleon CLI reads these case reports as one run, no gleon diff needed:

gleon test -- flutter test   # one run: sets GLEON_RUN_ID and GLEON_METRICS=1, records run.json
gleon report markdown        # the PR comment of the run; also html, junit, json
gleon dashboard              # adds the run to .gleon/history.json and renders dashboard.html
gleon approve                # writes the candidates of failed goldens to their PNG files

gleon test passes the test command's exit code through; on Windows it finds flutter.bat. Without it, the CLI picks the run of GLEON_RUN_ID, else the run of the newest report; reports without a run id are read together, with a warning that they may mix runs, and after a plain flutter test without metrics only the failures are there (approving works, the totals don't). In CI, set GLEON_RUN_ID for the job, upload .gleon/runs/ and pass the downloaded latest/ to --from (gleon report markdown --from <dir>/latest, gleon approve --from <dir>/latest, one --from per job). The CLI can also keep goldens out of Git (content-addressed blobs with small JSON manifests) and manage per-platform baselines for its own screenshots; per-platform goldens for this package are coming.

Performance

A widget golden (a Finder), from the captured frame to the verdict: Flutter encodes the frame as a PNG and compares it with its comparator (a pass short-cuts on equal bytes); gleon passes the frame's raw pixels and encodes a PNG only for a failure. Rendering the frame is the same for both and not measured. Exact, no .gleon/ workspace; Apple M3 Max, macOS, Flutter 3.47.6, mean latency with bench_press:

Scenario Golden Flutter SDK gleon Speedup
Passing 400x300 9.49 ms 0.46 ms 20x
390x844 15.2 ms 1.25 ms 12x
1170x2532 87.8 ms 7.7 ms 11x
Failing (small change, failure files written) 400x300 43.3 ms 3.1 ms 14x
390x844 77.0 ms 6.0 ms 13x

Encoding the PNG is most of Flutter's cost. The 390x844 and 1170x2532 passes resolve to 12.1x and 11.4x with 95% confidence intervals within ±1%; gleon's other samples (sub-millisecond, or writing files) varied too much for bench_press to resolve a ratio.

One golden comparison of PNG bytes (byte and ui.Image inputs), Flutter's own comparator (LocalFileComparator, behind flutter_test's matchesGoldenFile) against gleon, with the same golden file and candidate bytes. Measured inside flutter test, where golden tests run, with bench_press (mean latency; every ratio has a 95% confidence interval within ±8%). Apple M3 Max, macOS, Flutter 3.47.5:

Scenario Golden Flutter SDK gleon (exact) Speedup gleon ssim
Passing, identical bytes 400x300 191 µs 34 µs 5.7x 34 µs
390x844 213 µs 36 µs 5.9x 36 µs
1170x2532 423 µs 45 µs 9.5x 45 µs
Passing, re-encoded (same pixels) 400x300 2.95 ms 0.53 ms 5.5x 0.52 ms
390x844 7.67 ms 1.57 ms 4.9x 1.57 ms
1170x2532 56.4 ms 9.1 ms 6.2x 9.3 ms
Failing (small change, failure files written) 400x300 34.6 ms 1.96 ms 17.6x 1.85 ms
390x844 61.1 ms 4.44 ms 13.8x 4.07 ms

Across all cases gleon is 7.7x faster (geometric mean), and tolerant ssim costs no more than exact. A failing 1170x2532 golden takes Flutter about 0.4 s and gleon 25 ms; that case is beyond bench_press's 200 ms limit for one operation, so it is a plain timing. Flutter decodes through the engine but inverts both images and compares them pixel by pixel in Dart, and a failure renders two diff images and encodes four PNGs; gleon decodes and compares in Rust and writes one diff image. gleon here runs without a .gleon/ workspace (exact, nothing recorded, like Flutter).

How ssim decides

Two gates, both computed only around the pixels that actually differ:

  1. Envelope gate (full resolution): every pixel of one image must lie within the value range of the other image's 3x3 neighborhood (premultiplied RGBA, so alpha counts), widened by colorTolerance plus a share of the local contrast. Re-rasterization and sub-pixel shifts only produce values between neighbors; new content, color and alpha changes don't. Unexplained pixels fail as regions of 3+ pixels or with a strong deviation.
  2. Structural gate (half resolution, as in MS-SSIM): the minimum local SSIM (11x11 Gaussian window) must reach minSimilarity, catching structural loss such as blur. The minimum, not the mean, so a small change on a large golden is not averaged away.

The defaults are calibrated on a corpus of benign rendering noise vs. regressions in gleon-engine/tests/ssim_corpus.rs in the gleon repository; off-the-shelf crates (image-compare, pixelmatch/dify, butteraugli) could not separate that corpus with any threshold.

Known PoC limitations

  • ssim fails when glyphs move by half a pixel or more — typical of different operating systems' font engines. For one golden on every OS, use real fonts with exact or pixel instead (text never fails by default, see Real text).
  • ssim can pass a low-contrast color change of a one-pixel line. Use exact/pixel where every pixel matters.
  • Only Flutter's default LocalFileComparator is supported as the underlying golden store.
  • Web (--platform chrome) and on-device tests are not supported: the native library does not exist there, and the first comparison fails with a message saying so.

Contributing

Architecture

One package with a hard internal boundary:

lib/gleon.dart            exports only (flutter_test minus matchesGoldenFile, plus the gleon API)
lib/src/core/             plain Dart: never imports Flutter (dart:ui, package:flutter*)
  compare/                GoldenTolerance, PixelRegion (the call's tolerances,
                          masks and text regions)
  config/                 GleonIntegration (who calls), GleonSession (the native session)
  native/                 @Native leaf bindings, NativeEngine (ABI check, packing), verdicts,
                          error kinds
  hook/                   native targets, release download, source build, atomic writes (used by
                          hook/build.dart)
lib/src/flutter/          the Flutter layer: matchesGoldenFile, the widget capture and its text
                          regions, the comparator, loadAppFonts, FlutterSession (this package
                          as an integration)
hook/build.dart           thin build hook on top of lib/src/core/hook/
bin/                      maintainer scripts (dart:io + crypto only), run with plain `dart`

The native engine (gleon-ffi) does the whole job of a golden: it finds the workspace, resolves the .gleon/gleon.yaml rule, reads and compares the golden, and writes failure artifacts, case reports and updated goldens; this package passes facts (paths, the raw pixels or PNG, the call's tolerance, masks and text regions) and shows the verdict and texts it gets back. The C types and bindings of that contract live together in lib/src/core/native/gleon_ffi.dart.

Rule: code in lib/src/core/ must not import Flutter, so it stays usable from dart test, the build hook and other SDKs later; everything Flutter-specific (Rect, matchers, comparators) lives in lib/src/flutter/ and converts to core types at the boundary. DCM enforces the rule (avoid-banned-imports in analysis_options.yaml).

Why leaf FFI calls

Every native call is an isLeaf: true call. That allows passing the PNG zero-copy via Uint8List.address (and a call's strings as one UTF-8 buffer plus a Uint32List of their lengths) and returning {ptr, len} slices by value, so the Dart side never allocates native memory (no package:ffi, no malloc/free pairs; only the session is released by a NativeFinalizer). The price is that the isolate group cannot reach a GC safepoint while a comparison runs (milliseconds for typical goldens, seconds for very large SSIM comparisons). For tests that is the right trade-off: the test awaits the result anyway, and flutter test parallelizes across processes. Apps that must stay responsive would instead copy into malloced memory and make non-leaf calls in Isolate.run.

Checks

dart format --set-exit-if-changed .
flutter analyze --fatal-infos     # also in example/
dcm analyze .                     # DCM 1.39.2, also in example/
flutter test                      # also in example/

Case reports written by the tests are validated against case.v2.json when a gleon checkout sits next to this repository (../gleon, as in CI); without it that check is skipped.

analysis_options.yaml is the single, strict configuration (analyzer lints plus DCM presets); every disabled or narrowed rule carries its reason.

Benchmarks

benchmark/ holds bench_press benchmarks that need the Flutter engine (dart:ui), so they run under flutter test, not dart run bench_press run:

flutter test benchmark/      # ~3 min; results in build/benchmark/*.json
dart run bench_press report --from-json build/benchmark/golden_comparison.json
dart run bench_press report --from-json build/benchmark/widget_capture.json

To check a change for regressions, keep the JSON of a run before it and compare (dart run bench_press diff old.json build/benchmark/golden_comparison.json). Run on an idle machine: bench_press withholds ratios of unstable samples ("unresolved"). BENCH_PRESS_ARGS passes options (--validate for a smoke run, as CI does; --trials 30). The manual Benchmark workflow runs them on every supported host and puts the report into each job's summary.

Maintainers

The native code (gleon-engine, gleon-model, gleon-ffi) lives in the gleon repository; native/gleon_ref pins the commit CI builds. Build the library for this machine into native/<target>/ (the hook prefers it over downloading):

dart bin/build_native.dart                 # host; --target all cross-builds on macOS

Use plain dart, not dart run: dart run executes the build hook first. Cross builds need cargo-zigbuild (Linux) and cargo-xwin (Windows); CI builds every target natively. A build records the gleon commit it was made from, and the hook uses it only while native/gleon_ref pins that commit. To try an unmerged engine change on top of the pinned commit, build with --allow-dirty (and rebuild after every change: the hook cannot tell two dirty builds apart); to test another checkout, use the gleon_repo user-define.

Releasing: bump version in pubspec.yaml and CHANGELOG.md, update native/gleon_ref if the engine changed, and push the tag vX.Y.Z. The release workflow builds and tests all targets on their own OS, attaches the libraries and SHA256SUMS.txt to an immutable GitHub Release, and then verifies the download path on every OS.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages