perf(build): gate each wave once, brief each agent, cap change sets - #2
Merged
Merged
Conversation
What: build now runs a spec's Validation block once per wave, as the
gate before the wave's commits ("The wave gate" in parallel.md). A
change-set implementer runs only the test files it adds or edits plus
typecheck and lint. Commands marked `(end of build)` in the block, and
the e2e suite, benchmarks, and suite-repeating analyses even when
unmarked, run once on the final tree with the CI-parity gate.
ci-parity.md no longer re-runs a command that is green on the same tree.
The seen-red rule keeps its guarantee and loses its cost: free when the
test is written first, otherwise one break per slice running one test
file.
change-set-brief.py (new, build/scripts) cuts spec.md down to one change
set: every section but research and the change plan, the decisions the
change set links, its own plan, and implementation-notes.md without its
test and seam inventories. Lines are verbatim. parallel.md hands each
agent its brief instead of the whole spec and notes.
lint-spec.py fails a change set over 25 scenarios (change sets already
logged in implementation-notes.md are exempt, since they never
renumber) and prints, on a clean spec, the build waves the file lists
allow and the shared files that make a change set wait. scope asks for
a wide plan, one owner per shared file, and the `(end of build)` mark;
the scope-review feasibility lens checks both.
check-tests.py splits the `Tests added:` line only where a path::name
follows the comma, so a test name may hold commas.
plugins/ is regenerated, dev goes to 3.3.0, and validate.sh gains D01,
which runs the 21 new unit tests under dev/evals/tests. docs/decisions.md gains a Build speed
area: D-wave-gate, D-seen-red-cost, D-change-set-size, D-wave-report,
D-change-set-brief.
Why: feedback on the contexia searchable-kb build named five causes of
a slow build. The Validation block (lint, typecheck, build, unit,
integration, e2e, bench, and a complexity script that re-runs the suite
with coverage) ran more than once per change set. One change set spent
47 break-and-rerun cycles proving 75 tests red. Change set 4 carried 38
scenarios over about 60 files and took 73 minutes. Change sets 3, 4,
and 5 queued behind shared files. Every agent read about 130 KB of spec
and notes before writing anything; the brief for change set 5 is 42 KB.
The comma split was costing loops too: 75 of the 91 problems
check-tests.py reported on that plan were test names cut in two.
An old-versus-new run of the parallel-wave eval is recorded in
dev/evals/results.md: the Validation block ran 2 times instead of 4 and
no subagent ran it. Wall time did not improve on that fixture, whose
suite costs 20 seconds; a real build has to show the saving.
Considered: running the block once per build (commits in between would
be unverified); a file-count limit (file lists are prose, the count
would be a guess); failing the lint on a serial plan (no threshold
separates a careless chain from a necessary one); worktrees for
overlapping change sets (an overlap usually is a real dependency).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Makes the dev
buildstep faster by acting on five causes named in feedback on the contexia searchable-kb build.lint-spec.pyfails a change set over 25 scenarios. Change sets already built are exempt.lint-spec.pyprints the build waves the file lists allow and the files that make a change set wait. Scope is asked for a wide plan with one owner per shared file.change-set-brief.pycuts the spec down to one change set, verbatim. For one real change set that is 42 KB instead of 130 KB.Also fixes
check-tests.py, which split test names on every comma.Evidence
scripts/validate.shpasses, including the new D01 check that runs 21 unit tests for the three scripts.parallel-waveeval, recorded indev/evals/results.md: the Validation block ran 2 times instead of 4, and no subagent ran it.Notes
docs/decisions.md.factory-pluginbranch carries the same change plus the factory phase copies in commit 0130c2f.🤖 Generated with Claude Code