Skip to content

feat(factory): the factory and bootstrap plugins - #3

Merged
tobrun merged 31 commits into
mainfrom
factory-plugin
Oct 3, 2026
Merged

tobrun merged 31 commits into
mainfrom
factory-plugin

Conversation

@tobrun

@tobrun tobrun commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Summary

Brings back the factory plugin and adds the bootstrap plugin, plus the work that came with them (30 commits, about 200 files).

  • factory: an orchestrator skill (run) that drives a pipeline of typed phases declared in .factory/config.yaml, or the built-in scope, scope-review, build and ship. It judges each phase's result itself and ends in a pull request or a report. Includes factory-config.py, run-state.py, the phase bodies under factory/phases/, and the generated Codex distribution.
  • bootstrap: the agents-md skill, which probes a repository and writes its root AGENTS.md, gated by check-agents-md.py.
  • dev: ship's mutation testing is reported rather than a shippability gate; the build-speed change from perf(build): gate each wave once, brief each agent, cap change sets #2 is also carried in the factory phase copies.
  • validate.sh gains checks for the factory's unattended wording and result protocol (F01 to F03) and for the bootstrap checker (B01).

Read before merging

  • fa9cddf on main removed the factory plugin because the implementation "was not working out". This PR reverses that, so it needs a decision, not just a review.
  • The last recorded ship review of this branch (.dev/factory-plugin/review_2.md, not committed) has verdict BLOCK. It was written before the later factory fixes and has not been re-run.
  • scripts/validate.sh passes on this branch with main merged in. No end-to-end factory run is recorded in the repository.

🤖 Generated with Claude Code

tobrun and others added 30 commits September 20, 2026 11:14
What: add the factory plugin skeleton - plugin.json, README, marketplace
entries, the generator's PluginConfig, and the committed Codex distribution -
plus the nested-validate test harness under factory/evals/tests/.
Why: the factory orchestrator needs its own plugin identity in both
marketplaces and a Codex-buildable distribution before any skill exists;
docs/architecture.md and CLAUDE.md now state the invoker exception the run
skill relies on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add factory/references/factory-run.md (run state schema, result
envelope with skill_path, unattended policy, launch prompt shape), copy the
eight dev/references/*.md files with plan-layout.md extended for the new
plan files, and copy skill-metrics.py/architecture-check.py into
factory/scripts/.
Why: every phase copy and the run skill read this shared material rather
than duplicating it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: copy dev's scope, scope-review, build, and ship skills into
factory/skills/, rewrite every unattended-check hit with a factory-policy
recorded decision instead of a question to a person, add each skill's
Factory context section, and repoint scope-review's cross-references inside
factory/. Fix the one-line Codex spawn_agent naming gap in both the factory
copy and dev/skills/ship/references/orchestration.md's Native transport
prose, so a copied ship reaches Codex's real transport rung.
Why: these are the four phases the run skill orchestrates unattended; the
dev fix keeps the source skill's own transport documentation accurate too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add factory/skills/run/SKILL.md (transport and .dev/ preflight,
init/resume, inline scope, handoff, per-phase launch/judge/act loop,
closing report), references/judgment.md and references/launch.md, and
scripts/run-state.py with init/handoff/diff-spec/record/check-result/show.
Why: this is the orchestrator - the one skill that launches the phase
copies, judges their results against evidence, and survives a closed
session through the state file run-state.py owns.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add check_factory_unattended (F03, scoped to the four phase copies
plus factory-run.md, with a regex self-test), check_factory_protocol (F02,
the result-protocol literals and a guard against a phase invoking another),
and check_factory_script (F01, runs factory/evals/tests/ excluding the
nested-validate harness by name so the gate cannot recurse into itself).
Resolve C-factory-result's open verify mark in docs/contracts.md now that
check_factory_protocol exists to guarantee it.
Why: these are the deterministic guards the factory's unattended policy and
result protocol depend on; without check_factory_script excluding its own
harness by name, validate.sh would recurse into itself unboundedly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add the webhook delivery fixture (an http.server endpoint, two unit
tests, a launch command), setup.sh (a fresh git repo with .dev/ ignored and
a bare origin so a factory run can push offline), the recorded manual
end-to-end procedure per host with pass criteria and a real high-entropy
planted-secret variant, and results.md for recording outcomes.
Why: the fixture is small enough that a real factory run costs minutes and
real enough for build's unit layer; the bare origin lets the paid host runs
prove everything except the pull-request step, which stays honestly
unproven per D-codex-gh.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
do_POST and log_message are BaseHTTPRequestHandler overrides that
http.server calls, so they are deliberate API surface rather than dead
code. Record them in tools/harden/vulture_whitelist.py and drop the
noqa directives, which named rules this repo does not enable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Why: the four confirmed blockers in .dev/factory-plugin/review_1.md.

B1: check-result read `stop.kind` with no type guard, so a `stop` that was
a string or null crashed and exited 1 - a retryable failure - turning a
never-overridable hard stop into a repair and relaunch. A stopped result
now always exits 2, however loosely the phase wrote `stop`.

B2: nothing ever wrote `phases`, so resume could not tell a crashed attempt
from an unstarted one. Adds the `attempt` subcommand, the missing writer for
the attempt schema the run protocol documents, and calls it from run/SKILL.md
at launch and at close.

B3: the not-doing match was a bare `⊘` search, so any line citing the
character counted as an entry (17 instead of 12 on this repo's own spec) and
diff-spec invented `⊘ line dropped` drift. It now matches an entry opening
with the mark, as lint-spec.py parses the same notation.

B4: load_state, a missing spec.md, and record's stdin parse crashed or
returned 1, colliding with diff-spec's "drift found" code that the
orchestrator folds into a judgment. Bad calls now return 3, and every
subcommand's exit codes are documented in the module docstring.
Why: review_1.md blockers B5 and B6. The factory scope-review copy kept
dev's report template and authority rules verbatim, so it still offered
"APPROVED WITH DEFERRALS" and a "## Deferred" section that change set 3b
removed and D-scope-review-deferrals rejects. That verdict also contains
the substring the run skill's done criterion matched, so a deferrals
verdict would have advanced the run to build on an unsettled spec (B5).
The same copy queued escalations for an "interview below" that this copy
replaced with the unattended section 4 (B6).

What: the report template now writes only "Verdict: APPROVED", records
escalations as decided under policy, and carries a "## Rescope" section
for the failed/rescope ending instead of "## Deferred". judgment.md now
requires a whole line exactly equal to "Verdict: APPROVED", checked with
grep -x, so a longer verdict cannot pass. The frontmatter, overview and
authority rules repoint at section 4 and factory policy; the constraint
that refine agents never auto-apply an escalation is kept.
What: widen scripts/validate.sh's F03 pattern to catch the person-routing
phrasings it missed (plural "human calls", "question left for the user",
"the user asks/says/accepts", "prompt the user", "surface to the
maintainer", "AskUserQuestion"), guard it with positive and negative
self-test samples, skip YAML frontmatter (a description says when a person
invokes the skill, which is before the go), and rewrite every line the
widened pattern now catches in the factory phase copies to the unattended
policy. Adds two eval tests planting newly covered phrasings, and corrects
the scope comment's wrong justification for excluding ci-parity.md and
jira.md.

Why: review_1.md blocker B7 - C-factory-unattended was stated as a
guarantee but its guarantor regex missed live person-routing text already
inside its own scan root, so validate.sh stayed green while phases still
routed decisions to a person.
Why: mutation testing on run-state.py left 11 survivors after the attempt
subcommand and the new exit-code ladder landed. Five were real gaps: a
`return None` mutant on `return BAD_CALL` in cmd_handoff, cmd_record,
cmd_check_result and cmd_attempt turns exit 3 into exit 0, so a corrupt
state file or an impossible call would be reported to the orchestrator as
success, and stop_kind's `return ""` fallthrough was unobserved. No test
covered those error paths.

Adds seven subprocess tests: handoff over a truncated and over a
non-object state file, record with stdin JSON that parses but is not a
decision object, check-result on a JSON array, check-result on a stopped
result whose `stop` is neither dict nor str (exact stdout `stopped: `),
and attempt over a truncated and over a non-object state file.

Survivors now 6, all `return 0` sites where `SystemExit(None)` also exits 0.
Why: the ship re-review found round 1's `phases` writer inert - R7 nothing
read it back, so a resume relaunched as attempt 1 and check-result'd the
stale build-1.json left by the dead session; R8 the attempt was recorded
after a launch call that blocks until the phase ends, so a mid-phase crash
left the list empty and indistinguishable from never started; R9 section 2
resumed "the state file not marked done" when no writer produced any
terminal marker.

Section 2 now reads the history with `run-state.py show` before launching
anything and branches on the last attempt's status: a dangling `launched`
with its own result file present is picked up at check-result, one without
is closed `failed` with no `--result` and relaunched with guidance to start
from the artifacts on disk. Section 5 records the launch first and takes the
attempt number from what `run-state.py attempt` prints, so the state file is
the number's only owner. The run's terminal marker is the `decisions` list
ending in `end` or an `advance` on `ship`, which the state can already
answer and which factory-run.md now documents beside the attempt numbering.
Note: the content of this change is already in ac04e2d - a concurrently
running agent staged and committed the whole tree, including these files,
while this change was being staged. This commit records the intent and the
Why; nothing is added on top.

What: rewrite the factory gauntlet's threshold line so a run never amends
`docs/decisions.md` or the threshold it is judged by - a threshold it cannot
meet is a `failed` result naming the finding, the score, and the missed
threshold, or an `auto_decided` proposal for review. Rewrite mocking.md's
"a seam worth agreeing on with the user" line to the unattended policy.
Restate factory-run.md's policy prose so it states the rule instead of reading
as a routing instruction, and record that `run/SKILL.md` is the only file that
may use the interactive-only markers. Widen the F03 pattern from an allow-list
of its own samples into routing constructions over a much larger person
vocabulary (person, requester, owner, stakeholder, someone, whoever, team,
them), with a negated-clause test so "no phase skill asks a person anything"
stays clean, plus branches for gate nouns, approval-seeking and waiting for an
answer: 45 of 60 probed phrasings were missed before, 0 after. Make the
markers checkable - unbalanced, stray, and out-of-file markers are F03
failures, and a marker sharing its line no longer silences that line. Extend
SAMPLES and NON_SAMPLES to everything newly covered and newly excluded, and
add eval tests for the new shapes and every marker failure.

Why: the ship re-review, items R10-R13. Round 1 turned a human gate into
self-acceptance (a run could raise a threshold it had failed and write the
loosened gate into the committed decision ledger as precedent for later runs),
left a person-routing line inside F03's own scan root, widened the pattern
only as far as its own samples, and left an unclosed marker able to silence
the rest of a file.
Why: the ship re-review found six blockers in the exit ladder round 1
introduced. R1: argparse usage errors exited 2, the ladder's unappealable
hard stop, so a mistyped phase or status ended an unattended run as if a
secret had been found; a Parser subclass now reports usage errors as
BAD_CALL. R2: check-result called a missing or unparseable result file a
bad call while factory-run.md, run/SKILL.md and the spec's Error handling
all call it a failed attempt; it is now exit 1 with the reason, since a
subagent dying before writing its result is a phase outcome the
orchestrator can relaunch. R3: load_state validated only the top level, so
a null or list-shaped phases or scenario_texts crashed out as exit 1,
colliding with diff-spec's drift code; the shape the commands rely on is
validated and reported as BAD_CALL. R4: the hard-stop leniency covered
stop but not status, so a secret reported alongside status done or failed
was not read as a stop; a present stop kind now wins over status. R5:
attempt silently rewrote a closed attempt and recorded an unreadable
result as null while printing success; reclosing is refused and an
unusable result file is recorded as a failed attempt with its reason.
R6: tests for a second attempt on one phase and for every behavior above.
The ship gauntlet's ruff pass over the diff found 23 findings on added
lines. Fixed at the source rather than suppressed:

- re module aliases spelled out (re.S, re.M, re.I -> re.DOTALL,
  re.MULTILINE, re.IGNORECASE) in architecture-check.py, pr-evidence.py
  and aggregate-findings.py.
- explicit check= on every subprocess.run that reads returncode itself.
- skill-metrics.py numstat() leaked a file handle counting untracked
  lines; it now reads under a context manager.
- two adjacent f-strings inside a collection literal made explicit as
  one value in skill-metrics.py and pr-evidence.py. Neither was a bug:
  both sites build a one-element row and a three-element tuple would
  have crashed the consumer.
- the unused REJECTED constant in lint-spec.py deleted; the glyph is
  still matched through the ALTERNATIVE character class.

The dev and factory copies of each shared script stay byte-identical,
and plugins/ is regenerated rather than hand-edited.
The gauntlet's coverage-weighted complexity check was scoring every
script at 0% coverage, which collapses each score to complexity squared
and fails functions that are in fact well covered. The cause was the
measurement, not the code: this repository's tests drive their subject
scripts through subprocess.run([sys.executable, script, ...]), and a
plain `coverage run -m unittest` traces only the parent process.

coverage_run.sh wires coverage's documented subprocess support: a
sitecustomize.py on PYTHONPATH calls coverage.process_startup() in every
child, COVERAGE_PROCESS_START names the config, and both data_file and
source are absolute so a child running in a temp cwd still writes to one
place and still matches the sources.

run-state.py measures 99.5% with it, against 0% before, and the
README's gauntlet recipe now uses it.
test_factory_checks took 210.7s; it now takes 30.8s, and the full suite
drops from about 215s to 35.3s.

The cost was never the repo copy or the Codex rebuild, which measure
0.05s and 0.07s. It was that every test ran the whole gate, and 4.6s of
each 7.4s run was the copy's own check_factory_script re-running the
other 60 unit tests, which say nothing about F02 or F03. Twenty-seven
gate runs, three of them the same clean run, one test running validate
twice over identical state, and ten more run one after another inside a
single subtest.

Each distinct scenario is now named once in SCENARIOS and they run
concurrently, capped at the core count. Every assertion is preserved:
disabling check_factory_unattended in a copy still fails 19 of 20 tests.
The scenarios that only assert about a clean gate share one run, and no
scenario executes twice.

The futures live at module scope rather than in a class fixture because
tools/harden/flaky.py shuffles tests into one flat suite, and unittest
re-runs setUpClass every time execution crosses a class boundary. A
class fixture re-launched the whole fan-out many times per pass, which
made a 5-run sweep take 20m51s; at module scope it is once per process.
The two tests that never needed a gate run moved to their own classes.
Why: scope-review's phase-5 promotion step wrote these during the
spec-review run but did not commit them; committing now, before build,
so the ledger state build implements against is on disk and reviewable
on its own. Adds 37 decisions under "## Config-driven factory" to
docs/decisions.md, supersession notes on five factory-plugin entries,
and three drafted contracts to docs/contracts.md with "? verify" marks
that change set 7 of .dev/config-driven-factory/spec.md resolves.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 1 of config-driven-factory/spec.md: a strict shallow-YAML
reader/writer for .factory/config.yaml, the resolve/check/init/show/set/unset
CLI, and factory/references/pipeline-config.md describing the schema.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 3 of config-driven-factory/spec.md: run-state.py seeds an ordered
phase list, budgets and a ceiling, records attempt types, skill paths and
timestamps, seals arbitrary files at handoff, adds diff-config and reports
terminality. Adds the stub pipeline fixture and its tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 4 of config-driven-factory/spec.md: run/SKILL.md, judgment.md and
launch.md read phases, types, checks and budgets from factory-config.py instead
of naming four phases. Built-in check executables become real plugin-root
script paths and check now requires them to exist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 5 of config-driven-factory/spec.md: the four phase bodies move to
factory/phases/ so run is the plugin's only invocable skill. The generator,
F02 and F03, the path resolvers, docs and manifests follow the move, and F02
now holds every body and the launch wrapper to the result contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 6 of config-driven-factory/spec.md: factory-config.py inject copies
named phases and what they reference into .factory/, records provenance and
sha256s in .factory/.inject.json, refuses to overwrite edited copies without
--force, and points the config at the copies.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
…ants

Change set 7 of config-driven-factory/spec.md: docs/contracts.md carries the
three new factory contracts with their provenance resolved, and the evals
README gains the custom-pipeline and foreign-phase manual variants.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: run/SKILL.md now says explicitly that once past the go, every
remaining phase runs to completion in the same turn, one launch after
another, with no interim progress report or check-in; ending the turn
while run-state.py show still prints finished: no is a defect unless a
real host constraint forces it, in which case say so and print the
resume command. Step 7's advance also says to go straight back to
step 1 for the next phase instead of stopping.

Why: this design has no background runner anymore - run is one skill
looping through phases inline in a single orchestrator turn - and
nothing forbade ending that turn between phases. A run had stopped
after scope-review instead of continuing into build, with nothing in
the instructions to prevent the model from treating a routine phase
completion as a natural stopping point.
…-selection

What: a type (or a phase, overriding its type) may declare a `model`
string, resolved by factory-config.py and printed by show --resolved.
The orchestrator passes it to the host's subagent tool (launch.md's
transport ladder) when that tool takes a model parameter, and it is
silently unused when the host has none. check refuses model on an
interactive type, since that phase runs inline and never reaches a
subagent call to carry it to. The built-in pipeline declares no model
for any type: the valid values differ by host (Claude Code's Agent
tool vs. Codex's spawn_agent vs. opencode's task), so no single
default survives being run from a different host than the one it was
written for. Pin one per repository with factory-config.py set.

Why: a real run inherited Codex/GPT-5 for every phase, and the user
expected the removed runner's own per-phase defaults. D-model-selection
in docs/decisions.md had deferred this with an explicit reopen
condition - "a real run shows one phase needs a different model than
the session" - which this is; superseded by D-model-per-phase, which
also records why host-per-phase (mixing Claude Code and Codex in one
run) cannot come back under the current one-session orchestrator.
What: dropped "zero surviving mutants in scope" from the gauntlet's
threshold defaults in both dev/skills/ship and factory/phases/ship
(and their generated Codex copies). Fix agents still chase survivors
under the existing two-round loop and equivalent-mutant carve-out, and
a stubborn survivor is still recorded with its reason - it just no
longer fails the run on its own. The ship report's killed/survived/
uncovered counts are the evidence a person reads; there is no score
below which shipping is blocked. Recorded as D-mutation-threshold in
docs/decisions.md, since no prior decision had actually set this
default - it was unexamined.

Why: a real run scored 43% (430/1003) mutants killed on a diff-scoped
file and was ruled unshippable purely on that inherited coverage debt.
The user's own words: "we aim to do our best but zero is a bit too
harsh and unrealistic."
What: a third plugin, bootstrap, with one skill, agents-md. It probes a
consuming repository read-only (manifests, task runners, CI workflows,
test layout, git history), maps what it finds onto a closed concept map
distilled from the dev skills (test layers, mocking at boundaries, the
e2e environment, CI parity, fix-at-source quality, commits, plans, PR
evidence), interviews only for what the probe cannot settle, runs the
commands it will list, merges and prunes an existing AGENTS.md and
CLAUDE.md, and writes the root AGENTS.md after the user approves a diff.
Gaps become open items; nothing is scaffolded, CLAUDE.md is never
written, and nothing is committed.

check-agents-md.py is the gate every draft loops against: required
sections in order, the seven Commands slots filled or backed by an open
item, a closed open-item slug vocabulary, stale backticked paths and
links, commands whose package script, make target, just recipe, task,
or file does not exist, and a default branch that disagrees with
origin/HEAD. Checks it cannot settle statically print notes instead of
violations, so a false positive never traps the loop. 60 unittest cases
cover it and tie the concept map, the template, and the checker
together; validate.sh runs them as B01.

Wiring: marketplace entries for Claude Code and Codex, a generator
entry that copies only skills/, the generated plugins/bootstrap tree,
and a README row. Pi stays dev-only.

Why: the dev workflow's lessons only reach repositories that install
dev, and the repo facts its skills need (per-layer commands, the e2e
launch and environment, the merge gate, the scope vocabulary) are
re-probed on every run because no file declares them. A root AGENTS.md
written from a probe carries both to any agent. Three functional runs
on scratch copies of sibling repositories drove the checker fixes
(`bun turbo`, test path filters, naming conventions) and are recorded
in bootstrap/evals/results.md.

Considered: a dev skill (grows dev's namespace and the Pi package for a
one-time step); linking into dev/references (dangles in the generated
Codex tree); a CLAUDE.md pointer file (Claude Code reads AGENTS.md
directly).
What: docs/architecture.md now names three plugins, adds a bootstrap
component row, a "Bootstrapping a repository" flow, B01 in the
validation flow, the root AGENTS.md at the consuming-repository
boundary, and the checker as an entry point. docs/decisions.md gains a
Bootstrap plugin area: D-bootstrap-plugin, -skill-name, -reach, -merge,
-distilled-reference, -checker, and the not-doing entries -metrics,
-pi, and -dev-reads-agents-md. CLAUDE.md describes the third plugin and
says the Codex generator runs after changing any source plugin.

Why: the overview and ledger are what future specs in this repository
start from; D-packaging had rejected a third plugin for a different
product, so the ledger records why this one differs, and the deferred
follow-up of scope and build reading AGENTS.md's Commands section.
What: build now runs a spec's Validation block once per wave, as the
gate before the wave's commits ("The wave gate" in parallel.md). A
change-set implementer runs only the test files it adds or edits plus
typecheck and lint. Commands marked `(end of build)` in the block, and
the e2e suite, benchmarks, and suite-repeating analyses even when
unmarked, run once on the final tree with the CI-parity gate.
ci-parity.md no longer re-runs a command that is green on the same tree.

The seen-red rule keeps its guarantee and loses its cost: free when the
test is written first, otherwise one break per slice running one test
file.

change-set-brief.py (new, build/scripts) cuts spec.md down to one change
set: every section but research and the change plan, the decisions the
change set links, its own plan, and implementation-notes.md without its
test and seam inventories. Lines are verbatim. parallel.md hands each
agent its brief instead of the whole spec and notes.

lint-spec.py fails a change set over 25 scenarios (change sets already
logged in implementation-notes.md are exempt, since they never
renumber) and prints, on a clean spec, the build waves the file lists
allow and the shared files that make a change set wait. scope asks for
a wide plan, one owner per shared file, and the `(end of build)` mark;
the scope-review feasibility lens checks both.

check-tests.py splits the `Tests added:` line only where a path::name
follows the comma, so a test name may hold commas.

The factory phase copies carry the same edits, plugins/ is regenerated,
dev goes to 3.3.0, and validate.sh gains D01, which runs the 22 new
unit tests under dev/evals/tests. docs/decisions.md gains a Build speed
area: D-wave-gate, D-seen-red-cost, D-change-set-size, D-wave-report,
D-change-set-brief.

Why: feedback on the contexia searchable-kb build named five causes of
a slow build. The Validation block (lint, typecheck, build, unit,
integration, e2e, bench, and a complexity script that re-runs the suite
with coverage) ran more than once per change set. One change set spent
47 break-and-rerun cycles proving 75 tests red. Change set 4 carried 38
scenarios over about 60 files and took 73 minutes. Change sets 3, 4,
and 5 queued behind shared files. Every agent read about 130 KB of spec
and notes before writing anything; the brief for change set 5 is 42 KB.
The comma split was costing loops too: 75 of the 91 problems
check-tests.py reported on that plan were test names cut in two.

An old-versus-new run of the parallel-wave eval is recorded in
dev/evals/results.md: the Validation block ran 2 times instead of 4 and
no subagent ran it. Wall time did not improve on that fixture, whose
suite costs 20 seconds; a real build has to show the saving.

Considered: running the block once per build (commits in between would
be unverified); a file-count limit (file lists are prose, the count
would be a guess); failing the lint on a serial plan (no threshold
separates a careless chain from a necessary one); worktrees for
overlapping change sets (an overlap usually is a real dependency).

Directive: dev/skills/{scope,build}/scripts and their copies under
factory/phases must stay byte-identical; FactoryCopiesTest checks the
three scripts this change touches.
# Conflicts:
#	dev/evals/results.md
#	dev/evals/tests/test_build_speed.py
#	docs/architecture.md
#	docs/decisions.md
#	scripts/validate.sh
@tobrun
tobrun merged commit ff86fb6 into main Oct 3, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant