Skip to content

feat(gated): run the complete default module scope - #18

Merged
Disble merged 31 commits into
mainfrom
perf/module-scope-core
Sep 19, 2026
Merged

Disble merged 31 commits into
mainfrom
perf/module-scope-core

Conversation

@Disble

@Disble Disble commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Release 0.11.0

Gated() now answers the complete default Go module scope, instead of only the mutated file's package. Unsupported or custom test commands keep their ordinary configured path. Collision-safe compilation batches let modules with duplicate test-binary names gate rather than refusing.

Size exception

Maintainer-authorized exception: this cohesive performance/fidelity work spans 4,485 additions and 18 deletions from v0.10.0. It cannot be split into independently shippable slices: module-scope compilation, its scope admission/fallback, dependency closure, shared sandbox/compilation directory, collision batching, correctness tests, baseline counters, and evidence define one behavior. No comments, tests, or evidence were compressed to reduce the number.

Scope and non-goals

  • Preserves the exported API; changes documented Gated() behavior and its Gated: observables.
  • Does not claim an unmeasured speedup. Recent fail-fast and outer-parallelism experiments are recorded as refutations, not gains.
  • odd/ orchestration records and the stale lint backup remain untracked and excluded.

Verification

  • bash .githooks/pre-commit on the final commit: lint 0 issues; test.failfast 521 passed / 10 documented opt-in probes skipped; test.counters 15 passed / 1 environment-gated probe skipped.
  • Focused CLI help test: go test ./cmd/ditto/.
  • GitHub CI, CodeQL and OSV on this PR.
  • After merge, re-run and inspect CI, CodeQL, OSV and chained Mutation Tests on the exact main SHA before tagging v0.11.0.

Release artifacts

  • CHANGELOG.md
  • readme.md
  • docs/experiments/module-scope-runner.md
  • docs/experiments/gated-through-the-binary.md
  • docs/experiments/module-failfast-prototype.md
  • docs/experiments/failfast-cost-ceiling.md
  • docs/experiments/explicit-adaptive-scheduler-poc.md

The dominant cost of a release is starting the test command once per mutant,
measured at 700-1100ms of fixed overhead whatever the suite does. Gating was
supposed to remove it, but the gated path compiled only the mutated file's own
package: the default command is `go test -count=1 ./...`, so it answered a
smaller question than the caller asked.

This measures the ceiling of the honest version. Twelve real comparison mutants
in a four-package disposable module, three rotated rounds, one discarded
warm-up:

  A ordinary     13 driver starts, 52 package-test executions, 6 killed / 6 survived
  B package-only  1 driver start,  13 package-test executions, 5 killed / 7 survived
  C module-scope  1 driver start,  52 package-test executions, 6 killed / 6 survived

The sentinel is one mutant in pkg0 that only a test in the dependent pkg1 can
kill. B survives it while A and C kill it, on a fixture where every other
verdict agrees - so the narrowed scope is observable rather than inferred.
C/A wall clock was 0.1951, 0.1993 and 0.1809, all below the pre-registered 0.25.

The counter can refuse: a permanent test supplies 52 executions against a wrong
expectation of 51 and requires the exact refusal, and every selection asserts
its own package-execution growth, so a killed mutant cannot truncate the scope
and still look correct in the total.

The experiment test is behind the `experiment` build tag: it builds a module and
starts test binaries, so it is not part of the default suite.
The gated path compiled the mutated file's own package, which is a narrower
question than the default `go test -count=1 ./...`. A mutant in package P that
only a test in a package importing P can kill survived it, and nothing in the
tree could see the difference: the addresses matched, the totals could match,
and the run was faster either way.

This adds the runner that answers the configured scope instead. It discovers the
package layout once through `go list -json ./...`, compiles every package's test
binary in one `go test -c -o <directory> ./...` invocation, and then starts each
of those binaries for the unselected baseline and for every selected mutant, in
import-path order, from each package's own directory.

It fails closed rather than guessing. Two packages whose test binaries would
collide under one output directory are rejected before compilation; a discovery
or build failure, and a successful build that did not produce an expected
binary, all leave Built() false with the tool's own diagnostic, and start no
package binary at all. Packages without tests are compiled with the rest and
execute nothing, because there is no binary to execute.

A failing package does not stop the ones after it. Stopping at the first failure
would be right for a verdict and wrong for the scope: `-failfast` already makes
one running package stop at its first failing test, and a scope that truncated
there would answer for fewer packages than were asked for, which is the defect
this runner exists to remove.

Result semantics are the ones the existing runner already had: Ok means a
command or test failed, Err means everything passed.

The repository's own gate is what drove the shape of this: it refused the first
version for a dynamic error, a struct without json tags matched to Go's own
output, an init function, a print statement, and two test files that were not
`_internal_test.go` - which is this repository's convention for a test that
needs the package's own symbols.
Gated() built `go test -c` for the mutated file's package whatever the caller had
configured, so a repository asking for `./calc` silently got a narrower question
than its own command. Measured: against the four-package fixture, the
package-only path survived a mutant that only a dependent package kills, while
the complete scope killed it.

Gated() now replaces only the command sequences whose complete execution plan has
been measured - `go test -count=1 ./...` and its built-in `-json` form in either
flag order. Everything else keeps the ordinary laboratory and the runner the
caller configured, and the batched-laboratory shape is preserved so the counters
still say so: Gated() is 0 and every mutant is counted as fallen back.

The set is closed on purpose. A command that is almost the default is not the
default, and admitting one would repeat the defect this change removes. Extra
spacing, an extra flag, a package-local scope, an alternate executable and a make
target are all outside it by construction rather than by remembering to check.

This is a behavior change for anyone who paired Gated() with a custom
WithTestCommand: they now get their own command instead of a package-local
`go test -c`. That is the point - the old behavior answered a smaller question
than they asked - and it is declared here rather than left to be discovered.

The golden fixture moves to the scope Gated() is allowed to replace, so its
assertion that a gated run gated *something* still proves the path engages
rather than being relaxed to accept zero. A second run over the same fixture
proves the other half: a package-local command reports `none` with identical
verdicts. Weakening the admission table makes that guard refuse with
"a package-local command gated 4 mutants", which is how it was tested.

perf/baseline.json moves to 813, and the +24 is attributed by file rather than
to the change as a whole: 23 from the new internal/gobuildrunner/module_scope.go,
1 from internal/gatedlaboratory/gatedlaboratory.go, 0 from options.go, whose
command-scope table replaced a branchy classifier with the same number of
mutable sites. The three numbers sum to the 24 the ratchet reported.
The request behind this branch is long, it has already changed model once, and
the thing that would cost the most is a later session reconstructing intent from
a diff. This records the work as it happened: what was measured, what was
refuted, what was decided, and what is still unknown.

Eight entries so far. The two that matter most are corrections rather than
advances: the package-only path was measured to answer a narrower question than
the configured command, which redirected the whole proposal, and the repository
counter that fired at 813 was attributed per file rather than to the change as a
whole.

Every entry carries the revision, the model, the question, the prior hypothesis,
the intervention, the control, the exact counters, a wall-clock observation kept
separate from them, the verdict, what remains unknown and the next falsifiable
step. Refuted hypotheses stay in the record; the file is append-only and an
earlier entry is never rewritten to make the history look linear.

Its limits are stated in the file itself: wall clock is reported and never used
alone, a conclusion is not promoted to production until its control has run, and
the repository file wins if this ever disagrees with the copy mirrored outside
the repository.
Continuity evidence rather than prose: the four commit identities, what each one
carries, and the three times the repository's own gate refused the work before it
landed. A later session can start from the commit that made a claim instead of
from the diff of all of them.

It also records what is deliberately NOT in ditto's history: odd/, this session's
own task tracker, which is an el Gentleman convention rather than a ditto one.
A package test binary cannot emit `go test -json`: that flag belongs to the
driver that starts the binary, not to the binary itself. So a module-path kill
arrived as plain text, internal/verdict saw no event stream, and every kill
reported `unknown`.

Measured against the ordinary command over the same fixture and the same active
mutant: `unknown` where the control reported `assertion`
(docs/experiments/module-path-verdict-reason.md).

That is not a cosmetic loss. internal/confirminglaboratory re-runs a kill only
when its reason is Assertion, so on the gated path `--confirm-kills` silently
never fired and a flaky suite's false kill could not be caught - the defect
backlog entry 27 exists to remove, quietly absent from the path this change
promotes.

A failing package's output now goes through `go tool test2json -t -p <package>`
before it reaches the reporter. That is the toolchain's own conversion, so
nothing is re-implemented, and the verdict is still the binary's own exit status
rather than the converter's: a converter run over captured text has no test
status of its own to report. Output that cannot be converted is returned
untouched, which degrades to the previous behaviour rather than to a wrong
reason.

Only a failing package is converted. A reason is only ever asked of a kill, so a
selection where everything passed starts no converter at all, and a guard holds
that: a green selection asserts zero converter starts.

RED was captured behaviourally rather than as a compile error - the test failed
with `expected "assertion", actual "unknown"` before the change and passes after
it - and the guard was seen refusing when the conversion was deleted, with
ConverterStarts dropping from 1 to 0.

perf/baseline.json moves to 818, and the +5 is attributed to
internal/gobuildrunner/module_scope.go alone: 23 mutants before this change and
28 after, measured on that file.

A deadline kill is still not covered, and it is a different defect rather than
part of this one: the ordinary path writes its own marker when ditto fires the
clock, while here the clock belongs to the binary's -test.timeout, whose panic
text would convert into an Assertion. Recorded in the experiment note as open.
Every earlier measurement on this branch ran a fixture harness or a build-tagged
experiment. This one builds the binary from a disposable copy, points it at a
throwaway three-package module, and reads what a person gets:

  ordinary   23,145 ms   30 total, 24 killed, 6 survived
  --gated     7,361 ms   30 total, 24 killed, 6 survived
                         30 of 30 mutants ran from one compilation

0.318, or 3.14 times faster, with the sorted survivor addresses byte-identical
between the two modes over six real survivor reports.

That number is smaller than the 0.18-0.20 the mechanism measured in
module-scope-runner.md, and the difference is the point of running the binary: a
release also pays for parsing, instrumentation, the sandbox, the progress line,
and now one test2json conversion per killed selection.

Two things went wrong in the fixture before the numbers above were trustworthy,
and both are in the note because a measurement that cannot fail proves nothing.
The first fixture produced zero survivors, which would have made the address
comparison vacuous. It also reported `Gated: none of 24`, and the cause was the
fixture rather than the product: the generated sources had a trailing blank line,
so they were not gofmt-formatted, and schemata.Plan refuses a difference that
carries formatting. That is a real property of the product - a repository whose
sources are not gofmt'd gets no gating at all - and it is reported rather than
left as a puzzling zero.

The limits are stated in the note: a three-package module with a light suite is
the case gating is for and the case most favourable to it, and this repository's
own repository-sized gate was not run.
The instrument behind the previous commit, kept with the note it belongs to and
behind the experiment tag so it is not part of the default suite.

It answers one question three ways over a single fixture: the ordinary
`go test -count=1 -json ./...` command as the control, the module-scope runner's
own captured output, and that same output passed through `go tool test2json`.
The reason is read by `verdict.ReasonOf` - the product's own function rather
than a re-implementation of it - and the subtest that reads a kill fails rather
than reporting a survivor, so `unknown` cannot be read from an empty result.

A fourth subtest measures the compile case separately, because it is a different
path: a package that does not build leaves `Built()` false with the compiler's
own diagnostic and starts no package binary.

test2json is measured here as a ceiling, before any production code changed
shape, which is why the conversion lives in the experiment rather than in a
helper the product would then have had to be refactored onto.
…ounds

The first end-to-end number was one run per mode: 23,145 ms against 7,361 ms,
ratio 0.318. That is two absolute numbers measured minutes apart in separate
processes, which is not a measurement by this repository's own standard - its
doctrine is to prefer a ratio taken inside one window, because the ratio cancels
the load an absolute pair does not.

Repeated properly: one discarded warm-up, then three measured rounds with the
mode order rotated, the ratio computed within each round.

  round 1  A B   22,959 ms  /  7,349 ms  =  0.3201
  round 2  B A   23,074 ms  /  7,266 ms  =  0.3149
  round 3  A B   23,190 ms  /  7,424 ms  =  0.3201

The three agree to within 1.7%, and round 2 ran the gated mode first, so the
result is not one mode being favoured by position. Verdicts and survivor
addresses are unchanged.

The single pair is replaced rather than kept beside the rounds that supersede
it, and the note says so. The conclusion is also narrowed to the case it was
measured on: a light suite, where the fixed cost of starting the test command
dominates the bill. That is the case the tool is built for, and it is not every
case.
…e count

The gain was measured on a three-package module and the design was read as
removing the per-mutant driver start, which is paid once per mutant either way.
That reading was wrong, and a ten-package fixture shows it:

  3 packages, 30 mutants   ordinary 22,959 ms  gated  7,349 ms  = 0.3201
  10 packages, 40 mutants  ordinary 64,195 ms  gated 42,686 ms  = 0.6649

The gain collapsed from 3.14x to 1.50x, with identical verdicts in both
fixtures. The module path replaces one driver start per mutant with one test
binary PER PACKAGE per mutant, so its cost is selections x packages while the
ordinary path is selections alone. Recorded as a correction rather than as a new
result, because it invalidates a reading this log already published.

The arithmetic that follows - about 98 ms per binary start and a crossover near
sixteen packages - is derived from measured totals and is labelled an estimate,
not a measurement.

This is the finding that selects the next change: run only the packages whose
test binaries can observe the mutated package, taken from the Go import graph. A
package that cannot transitively import the mutated one cannot observe the
mutation, so the reduction is provable, and the cross-package sentinel that
caught the package-only defect is the guard that the closure is not too narrow.
…ackage scope

The module path costs selections x packages, so it stops winning above roughly
sixteen packages - measured at 0.6649 on ten packages against 0.3201 on three.
This measures whether that factor can be divided by running only the package test
binaries whose dependency closure contains the mutated package.

Seven-package fixture: base is mutated, mid imports it, top imports mid, and
four islands import nothing local. Closure read from go list -deps -test -json.

  packages that can observe a mutation in base: base, mid, top - three of seven
  unrestricted  35 executions  killed, killed, killed, survived
  restricted    15 executions  killed, killed, killed, survived

57% fewer executions with identical verdicts, and selection 2 - the sentinel only
mid kills - is still killed. top is included even though it never names base,
which is the case a closure built from direct imports alone would miss.

The reduction is provable rather than heuristic: a test binary compiles its own
package plus the transitive closure of what it imports, so a binary without the
mutated package in that closure contains no code that can refer to anything the
mutation changed. It is not the defect that was just fixed, either - that ran the
mutated package and nothing else, while this runs it plus everything that can
reach it.

The control ties the experiment to the product: the shipped runner started 35
package binaries over the same fixture and the same five runs, matching the
unrestricted mode exactly. Without that, the numbers would describe the harness.

Two harness defects are recorded because they were found before the numbers were
trusted: the shared environment helper ignored its selector argument and
hardcoded mutant 1, and the first closure omitted the mutated package itself.

The same commit fixes that environment helper, which the earlier reason
experiment had been carrying.
The ceiling is measured before the design: three of seven packages can observe a
mutation in the mutated one, and running only those removed 20 of 35 executions
with identical verdicts including the sentinel. The shipped runner agreed with
the unrestricted mode to the binary, so the numbers describe the runner rather
than the harness.

Recorded with what it still does not establish: the closure's own cost is not
measured, a single-chain repository gets no division at all, and nothing is
implemented yet.
The module path costs selections x packages, so it stops winning above roughly
sixteen packages: 0.3201 on a three-package module, 0.6649 on a ten-package one,
both measured in docs/experiments/gated-through-the-binary.md. This divides that
factor.

A package test binary compiles its own package plus the transitive closure of
what it imports, so a binary without the mutated package in that closure holds no
code that can refer to anything the mutation changed. Running it can only cost
time. The closure is read from `go list -deps -test -json ./...` - the
toolchain's own answer, with test imports included because a package's external
test file is a separate package that imports the subject.

Measured, on a ten-package module with forty mutants and identical verdicts:

  before  42,686 ms   ratio 0.6649
  after   25,607 ms   ratio 0.3995
  ordinary 64,103 ms

That fixture is the best case for the closure - every package is independent, so
a mutation is observed by one package of ten. A repository whose packages form a
chain gets less, and a single chain gets nothing.

It is not the package-only defect returning. That one ran the mutated package and
nothing else, so a mutant only a dependent package could kill survived. This runs
it plus every package that can reach it, and the guard is the same sentinel: a
fixture where only the dependent package kills the mutant, asserted to still be
killed.

The saving fails open. A scope that nothing resolves, a layout this cannot read,
a runner nobody scoped - all run every package, which is what shipped before. The
closure is a saving and never a licence to run less than the caller asked for.

SkippedPackages counts what the scoping removed, so the reduction is a number
rather than an impression.

RED was captured behaviourally: with ScopeTo storing the scope but not applying
it, the guard failed with 4 executions against 3. It was then seen refusing again
with the filter disabled. `discover` was split when the repository's own lint
refused it at cyclomatic complexity 17 against a maximum of 10.

perf/baseline.json moves to 846, attributed per file:
internal/gobuildrunner/module_scope.go 28 -> 54, and
internal/gatedlaboratory/gatedlaboratory.go 37 -> 39. The two sum to the 28 the
ratchet reported.
The measured ceiling reproduced through the shipped binary: on a ten-package
module with forty mutants and identical verdicts, the gated ratio fell from
0.6649 to 0.3995 against an ordinary run of 64,103 ms.

Recorded with its limits stated in the same entry: that fixture is the best case
for the closure, since every package is independent; a chain-shaped repository
gets less and a single chain gets nothing; the closure's own go list cost is
still unmeasured; and the ten-package ratio was taken once per mode rather than
across rotated rounds, so it is reported rather than relied on.
Entry 014's number was one pair per mode, which is not a measurement by this
repository's standard. Repeated with a discarded warm-up and three rotated
rounds:

  1 A B  63,619 / 25,546 = 0.4015
  2 B A  25,454 / 63,560 = 0.4005
  3 A B  63,414 / 25,618 = 0.4040

Spread 0.9%, with round 2 running the gated mode first. Identical verdicts in
every round: 40 total, 20 killed, 20 survived.

The single pair from entry 014 sat inside that band, so the earlier number was a
small sample of a stable one rather than luck. The caveat is closed rather than
carried forward, and what remains unknown is stated without it: this fixture is
still the best case for the closure.
…diction

The closure had only ever been measured on islands, which is the best case for
it. An eight-package chain puts the two extremes in reach: every package
transitively imports pkg0, and nothing imports pkg7.

Observer counts came out exactly as predicted - 8 of 8 at the head, 1 of 8 at the
tip - and the wall clock went the other way. Three rotated rounds each:

  chain-head (closure removes nothing)  0.3492, 0.3528, 0.3597
  chain-tail (closure removes 7 of 8)   0.5408, 0.5623, 0.5536

The fixture where the closure does nothing is the faster one, by about 1.5x. The
prediction in this note was that the head fixture would land near the
pre-closure worst case; it is refuted, and in the reversed direction.

Both fixtures reported identical verdicts and gated all six mutants, and the
fixtures differ only in which package holds the mutable sites.

This is reported as a refutation rather than explained away. Something other than
the number of package executions decides these two numbers, and the note names
its candidates without choosing between them, because the obvious story -
instrumenting a package everything imports forces a wider rebuild - has no
measurement behind it yet.

Recorded in docs/experiments/chain-shaped-module.md.
… it away

Two of three hypotheses held and the third died in the reversed direction: the
fixture where the closure removes nothing is 1.5x faster than the one where it
removes seven of eight, across three rotated rounds each.

Recorded as a correction, with what it puts in doubt: the claim that the closure
is what decides these numbers. The candidates are named and none is chosen,
because the obvious story has no measurement behind it yet.

The next step is named as a phase split rather than a guess: time the compile,
the baseline and the selections separately on both fixtures, and read
SkippedPackages to confirm the closure engaged on the tail fixture at all. If it
did not, entries 013 and 014 need re-reading.
The comfortable reading of the previous commit is that the closure silently did
not engage on the tail fixture, which would make the refutation a harness defect
rather than a result. It was checked rather than assumed, by printing the
runner's own counters from a disposable copy of the tree:

  chain-head  packageRuns 16 (2 selections x 8)  skipped 0   compilations 1
  chain-tail  packageRuns  2 (2 selections x 1)  skipped 14  compilations 1

The closure engaged exactly as designed. The tail fixture started eight times
fewer package binaries for the same selections and was still the slower of the
two, by about 1.9 s.

That kills the most comfortable explanation. Whatever costs the tail fixture its
time is not the number of package executions, and the note now says so instead of
leaving the reader to guess which reading of the earlier measurement was
intended.

The remaining candidates are named with a kill criterion attached to the one
that looks obvious: if the two compiles are within noise, the wider-rebuild story
is dead too.
… inverted

The chain refutation survived the check that would have explained it away. The
runner's own counters, printed from a disposable copy, show the closure working
exactly as designed on the tail fixture - 2 package runs and 14 skips where the
head fixture ran 16 and skipped none - and the tail fixture was still slower.

That leaves the question open with its candidates narrowed, and gives the
obvious one a kill criterion instead of a story.
…ce file

The chain refutation of the last three commits had its candidates narrowed to
one compile, the sandbox, the instrumentation, and the converters. A disposable
copy timed the phases separately and printed the runner's counters on every call,
which answers it:

  chain-head  discover 169  build 1,651  run 1,902  56 package runs  wall 3,828
  chain-tail  discover 141  build 1,676  run 1,364  24 unscoped
              discover 155  build 1,580  run   624   5 scoped       wall 5,718

The tail fixture prints TWO runners, and each compiled the whole module: 1,676 ms
and 1,580 ms. The shape explains why -

  chain-head  pkg0/pkg0.go - 6 mutants     one file   -> one batch -> one compile
  chain-tail  pkg0/pkg0.go - 2 mutants     two files  -> two batches -> two compiles
              pkg7/pkg7.go - 4 mutants

ditto.Release batches per source file and each batch builds its own
GatedLaboratory runner, so the module-wide compilation is paid once per source
file with mutants rather than once per release. The second 1,580 ms is the 1.9
second gap, and the package executions had nothing to do with it.

The wider-rebuild conjecture this note had named is refuted: the two compiles are
1,676 against 1,580 ms for the same eight-package module, and measured outside
the product 1,306 against 1,347 ms.

This also re-reads entry 014/015 without changing what they measured: the
ten-package fixture has mutants in every package, so it paid ten compiles and
still reached 0.3995. That number is a lower bound on what sharing one compile
would give, and the size of the gap is left as arithmetic rather than claimed as
measurement.
…pile

The chain refutation is now explained by measurement rather than by a story. The
tail fixture prints two runners because two source files had mutants, and each
batch compiled the whole module - 1,676 ms and 1,580 ms against the head
fixture's single 1,651 ms. The second compile is the 1.9 second gap.

The wider-rebuild conjecture is refuted, and the next lever is named with its
price: share one compilation across the batches of a release.

Entries 014 and 015 are re-read rather than corrected - the ten-package fixture
has mutants in every package, so it paid ten compiles and reached 0.3995 anyway,
which makes that number a lower bound.
Entry 018 found the module-wide compile is charged once per source file with
mutants. This measures what removing that charge is worth, before any code
changes shape: two loops of ten compiles over a ten-package module, differing
only in whether the output directory is fresh.

  ten fresh directories   15,466 ms
  one shared directory     3,087 ms   (1,405 first, then 166-288 each)

A saving of 12,379 ms against a pre-registered threshold of 8,500. For scale,
the ten-package gated run measured 25,607 ms, so this is 48% of that run's entire
wall clock. Every shared compile after the first is a small fraction of the
first, which is Go's own up-to-date check working as backlog entry 14 described
it - so the mechanism is the toolchain's and nothing is re-implemented.

The measurement's limit is stated in the note rather than discovered later: the
tree did not change between compiles, and in a real release each batch
instruments a different file. That makes 12,379 ms a ceiling for a fan-shaped
module and an overestimate for a chain, with the true figure somewhere between it
and zero.

Closure discovery is priced in the same run: about 180 ms per release, paid once.
Entry 019 priced sharing one compilation directory per release at 12,379 ms over
ten batches. This is that change built, measured, and found to pay nothing.

  before  0.4005 0.4040 0.3995
  after   0.4031 0.3996 0.4015

The sharing itself works - a guard was watched refusing when the rule was broken,
and a disposable build printing the runner's own state shows every batch writing
into ditto-module-compile-79694792. What does not happen is the saving: every
batch still spends about 1,750 ms compiling, against the ~170 ms a reused
directory measured in the priced loop.

The cause is the sandbox. Each batch links its own, so the package directories
have different absolute paths, and Go's build IDs cover those paths. Nothing is
up to date between batches however the output directory is chosen, so the
up-to-date check entry 019 was measuring never gets to fire. That number was a
real measurement of a situation that cannot occur - a limit the note had stated,
and which turned out to be the whole answer.

The change is reverted rather than kept. It added three mutable sites and a
counter and bought nothing measurable, and an unearned cost is written down, not
carried.

What would collect the prize is named rather than attempted at the end of a long
session: one sandbox per release instead of one per batch, which stabilises the
paths, and only then a shared compilation directory. The order matters, and this
measurement is why.
…nes it did not

The artifact the request opened with: what the numbers were, what they are, and
under which conditions each one means anything.

Nine recorded counters, eight of which did not move - because the optimised path
is opt-in and does not touch the road those counters measure. The ninth grew from
789 to 846, attributed per file at each of its three steps.

Then the measured gains with their conditions attached: the isolated mechanism
(13 to 1 driver starts, 52 package executions preserved), the shipped binary over
a three-package module (0.3149-0.3201), the ten-package module before and after
the observability closure (0.6649 to 0.4005-0.4040), and the closure's own
ceiling (3 of 7 packages observe the mutated one).

It also records what was refuted rather than only what worked: the chain fixture
where the faster end is the one the closure does nothing for, the verdict reason
that arrived as unknown and made --confirm-kills a silent no-op, and the shared
compilation directory that was priced at 12,379 ms, built, measured to pay
nothing, and reverted.

The unmeasured list is in the same file, in its own section, including the two
that matter most: a heavy suite, and this repository's own gate.
The module-wide compile was charged once per source file with mutants, because
ditto.Release batches per file and each batch built its own runner. On a
ten-package module that was ten full module builds for a forty-mutant run, and
docs/experiments/chain-shaped-module.md measured it as the entire difference
between two fixtures that differed only in where their mutable sites were.

Sharing the compilation directory was built first, on its own, and paid nothing:
0.4031 / 0.3996 / 0.4015 against 0.4005 / 0.4040 / 0.3995. The reason was the
sandbox. Every batch linked its own, so the package directories had different
absolute paths, Go's build IDs cover those paths, and nothing was ever up to
date however the output directory was chosen. That change was reverted rather
than carried.

This is both halves together, which is what the two measurements said the change
had to be:

  a ten-package module, forty mutants, three rotated rounds
    ordinary   64,583 / 65,964 / 64,006 ms
    --gated    14,620 / 14,708 / 14,798 ms
    ratio      0.2264 / 0.2230 / 0.2312

against 0.4005-0.4040 for the same fixture before this change. Identical verdicts
throughout: 40 total, 20 killed, 20 survived, 40 of 40 mutants ran from one
compilation.

Reusing the sandbox is safe because every batch restores the file it overwrote
before it returns, so the tree is pristine between batches and one batch cannot
see another's mutation. It is created through the temporary directory, so the
release's existing cleanup removes it with every other sandbox, and the
compilation directory lives inside it so it needs no lifetime of its own.

`CompilationDirectories()` is the integer that says the toolchain's up-to-date
check is allowed to work across batches, and the guard asserts one sandbox and
one directory for two batches, with the compilation-directory half seen refusing
when the sharing rule was broken.

perf/baseline.json moves to 850, attributed per file:
internal/gatedlaboratory/gatedlaboratory.go 39 -> 42 and
internal/gobuildrunner/module_scope.go 54 -> 55, summing to the 4 the ratchet
reported.
The row the last round left open is filled in: a ten-package module went from
0.4005-0.4040 to 0.2230-0.2312 once a release reused one sandbox and one
compilation directory, with identical verdicts and 14.6 s against 64.6 s.

The refuted entry is corrected in place rather than left reading as a dead end:
the directory alone paid nothing, the price was real, and the sandbox was the
half that made it collectable. The ratchet table gains its third step, 846 to
850, attributed per file.
… stopped moving

Entry 020 measured a shared compilation directory paying nothing and reverted it.
This records the order being the reason: with one sandbox per release as well,
the same fixture fell from 0.4005-0.4040 to 0.2230-0.2312, 25.6 s to 14.6 s, with
identical verdicts. The 12.4 s entry 019 priced came in at 11 s delivered.

The ratchet's third step is attributed per file, and what was not re-measured is
named: the two chain fixtures, and this repository's own gate.
The module scope named every test binary from `path.Base(importPath)` and
refused a tree where two packages produced the same name. On ditto itself that
was fatal: `github.com/Disble/ditto`, `cmd/ditto` and `internal/ditto` all
produce `ditto.test.exe`, the runner returned `Built=false` with
`Compilations=0` and `PackageRuns=0`, and GatedLaboratory fell back every file —
so module-scope gating gated 0 of 850 mutants even though schemata can express
517 of them (60.8%).

planCompileBatches now assigns every discovered package to the lowest-index
batch free of its binary name (case-folded on Windows, where the filesystem is
too), each batch compiles into its own `batch-N` directory under the run's
output directory, and prepare runs one `go test -c -o <batchDir> <that batch's
sorted import paths>` per batch. `errBinaryNameCollision` survives as an
invariant on a produced batch; an empty discovery fails closed.

Three measured facts shaped it, each reproduced in a throwaway module:

- the toolchain refuses a duplicate basename per argument list even when
  neither package has test files, so the collision key covers every package;
- `go list -deps -test -json ./...` does not type-check, so a package with no
  tests and a type error passes discovery while `go test ./...` exits 1 naming
  it — the union of batch arguments therefore equals the configured scope,
  which is what makes the module path fail closed where the command does;
- two sequential `go test -c -o <same dir>` compiles of same-named packages
  exit 0 while silently leaving one binary, a second reason each batch needs its
  own directory.

Measured through the shipped binary on a throwaway colliding module (21
mutants, `--threshold 0`): ordinary 21/15/6, gated 21/15/6 with
`Gated: 12 of 21`, and byte-identical survivor addresses (sha256
c1957c66a5268e8a2d6d52667abdd353158648d4ce134ea6d74918f4ddb7e8ec). The
pre-change binary on the same fixture reported `Gated: none of 21`. On ditto's
own tree the runner now reports `Built=true`, `Compilations=3`,
`PackageRuns=45` where it previously refused. A collision-free module still
plans exactly one batch, so the measured one-compile win is unchanged.

Ratchet 850 → 873, counted by this repository's own gate on the committed tree
and attributed to module_scope.go: internal/perfbench counts product .go files
only, so the two changed test files contribute none. An earlier state of the
change counted 874; the extraction that brought prepare's cyclomatic complexity
inside the gate's cyclop limit moved the count by one, and the number recorded
is the one the gate counted.

Evidence and limits: docs/performance-core-log.md entry 022.
…laims

Entry 022 of the why-log carries the whole measurement: the defect (test
binaries named from `path.Base`, so `ditto`, `cmd/ditto` and `internal/ditto`
all produce `ditto.test.exe`), the pre-change evidence (`Built=false`,
`Gated=0`, 0 of 850 mutants gated against a syntactic ceiling of 517), the
three throwaway-module toolchain facts that shaped the fix, the control
through the shipped binary (`Gated: none of 21` before, `Gated: 12 of 21`
after, identical totals and survivor addresses), and the runner-level result
on ditto's own tree (`Built=true`, `Compilations=3`, `PackageRuns=45`).
Its verdict states plainly what is still unmeasured: the realized gated share
on ditto itself, because the bounded release died to its own time budget
before any verdict.

`docs/performance-core-metrics.md` is reconciled rather than rewritten. Two
claims had been left behind by measurements already recorded in the same file:
section 10 still closed with the pre-shared-sandbox ten-package claim of
"approximately two fifths" while section 4 records 0.2230-0.2312, and section 2
called its fixture a three-package module while it executes four package test
binaries, which is what makes its 52 executions. Section 11 records the
collision finding and this fix, with the same unmeasured limit spelled out.

One line appended to `docs/learning-log.md`.
CI refuted the local gate, on the first pull request this branch opened.
`make lint` was green on Windows and red on Linux with one finding:

    internal/gobuildrunner/module_scope.go:593:49: validateBatches - goos
    always receives "linux" (unparam)

Nothing about the code differs by platform. What differs is the analyzer's
input: validateBatches takes the target OS so the Windows-only half of its
invariant can be checked on a Linux host, and its only varying call sites were
tests passing the literal "linux" -- which, on a Linux host, is exactly what
runtime.GOOS already gives the single production call site. A parameter that
never varies is one unparam wants deleted, and deleting it would have removed
the only route to testing Windows naming from the platform CI runs on.

So the fix is the coverage the parameter was for, not a suppression and not a
deletion: a case-only duplicate is one test binary on Windows and two files on
Linux, and both answers are now asserted against the validator itself instead
of only through planning.

Verified where it failed rather than where it passed:

- `GOOS=linux ./.bin/golangci-lint run` reproduced the cloud finding locally,
  the same single issue, and reports 0 issues after the change.
- the host gate is unchanged and green: lint 0 issues, 522 tests passed with
  the 10 documented opt-in probes skipped, 15 exact counters passed.
- manual mutation: with the case-fold deleted from binaryNameKey the new test
  fails (`expected error with "ditto: module test binary name collision" in
  chain but got nil`), and it passes again once the fold is restored.

The divergence itself is recorded in docs/learning-log.md, because the shape
recurs: a lint verdict can depend on the platform the linter runs on whenever
a platform seam is exercised only through test literals.
@sonarqubecloud

Copy link
Copy Markdown

@Disble
Disble merged commit 0d315b2 into main Sep 19, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant