feat(factory): the factory and bootstrap plugins - #3
Merged
Merged
Conversation
What: add the factory plugin skeleton - plugin.json, README, marketplace entries, the generator's PluginConfig, and the committed Codex distribution - plus the nested-validate test harness under factory/evals/tests/. Why: the factory orchestrator needs its own plugin identity in both marketplaces and a Codex-buildable distribution before any skill exists; docs/architecture.md and CLAUDE.md now state the invoker exception the run skill relies on. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add factory/references/factory-run.md (run state schema, result envelope with skill_path, unattended policy, launch prompt shape), copy the eight dev/references/*.md files with plan-layout.md extended for the new plan files, and copy skill-metrics.py/architecture-check.py into factory/scripts/. Why: every phase copy and the run skill read this shared material rather than duplicating it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: copy dev's scope, scope-review, build, and ship skills into factory/skills/, rewrite every unattended-check hit with a factory-policy recorded decision instead of a question to a person, add each skill's Factory context section, and repoint scope-review's cross-references inside factory/. Fix the one-line Codex spawn_agent naming gap in both the factory copy and dev/skills/ship/references/orchestration.md's Native transport prose, so a copied ship reaches Codex's real transport rung. Why: these are the four phases the run skill orchestrates unattended; the dev fix keeps the source skill's own transport documentation accurate too. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add factory/skills/run/SKILL.md (transport and .dev/ preflight, init/resume, inline scope, handoff, per-phase launch/judge/act loop, closing report), references/judgment.md and references/launch.md, and scripts/run-state.py with init/handoff/diff-spec/record/check-result/show. Why: this is the orchestrator - the one skill that launches the phase copies, judges their results against evidence, and survives a closed session through the state file run-state.py owns. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add check_factory_unattended (F03, scoped to the four phase copies plus factory-run.md, with a regex self-test), check_factory_protocol (F02, the result-protocol literals and a guard against a phase invoking another), and check_factory_script (F01, runs factory/evals/tests/ excluding the nested-validate harness by name so the gate cannot recurse into itself). Resolve C-factory-result's open verify mark in docs/contracts.md now that check_factory_protocol exists to guarantee it. Why: these are the deterministic guards the factory's unattended policy and result protocol depend on; without check_factory_script excluding its own harness by name, validate.sh would recurse into itself unboundedly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: add the webhook delivery fixture (an http.server endpoint, two unit tests, a launch command), setup.sh (a fresh git repo with .dev/ ignored and a bare origin so a factory run can push offline), the recorded manual end-to-end procedure per host with pass criteria and a real high-entropy planted-secret variant, and results.md for recording outcomes. Why: the fixture is small enough that a real factory run costs minutes and real enough for build's unit layer; the bare origin lets the paid host runs prove everything except the pull-request step, which stays honestly unproven per D-codex-gh. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
do_POST and log_message are BaseHTTPRequestHandler overrides that http.server calls, so they are deliberate API surface rather than dead code. Record them in tools/harden/vulture_whitelist.py and drop the noqa directives, which named rules this repo does not enable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Why: the four confirmed blockers in .dev/factory-plugin/review_1.md. B1: check-result read `stop.kind` with no type guard, so a `stop` that was a string or null crashed and exited 1 - a retryable failure - turning a never-overridable hard stop into a repair and relaunch. A stopped result now always exits 2, however loosely the phase wrote `stop`. B2: nothing ever wrote `phases`, so resume could not tell a crashed attempt from an unstarted one. Adds the `attempt` subcommand, the missing writer for the attempt schema the run protocol documents, and calls it from run/SKILL.md at launch and at close. B3: the not-doing match was a bare `⊘` search, so any line citing the character counted as an entry (17 instead of 12 on this repo's own spec) and diff-spec invented `⊘ line dropped` drift. It now matches an entry opening with the mark, as lint-spec.py parses the same notation. B4: load_state, a missing spec.md, and record's stdin parse crashed or returned 1, colliding with diff-spec's "drift found" code that the orchestrator folds into a judgment. Bad calls now return 3, and every subcommand's exit codes are documented in the module docstring.
Why: review_1.md blockers B5 and B6. The factory scope-review copy kept dev's report template and authority rules verbatim, so it still offered "APPROVED WITH DEFERRALS" and a "## Deferred" section that change set 3b removed and D-scope-review-deferrals rejects. That verdict also contains the substring the run skill's done criterion matched, so a deferrals verdict would have advanced the run to build on an unsettled spec (B5). The same copy queued escalations for an "interview below" that this copy replaced with the unattended section 4 (B6). What: the report template now writes only "Verdict: APPROVED", records escalations as decided under policy, and carries a "## Rescope" section for the failed/rescope ending instead of "## Deferred". judgment.md now requires a whole line exactly equal to "Verdict: APPROVED", checked with grep -x, so a longer verdict cannot pass. The frontmatter, overview and authority rules repoint at section 4 and factory policy; the constraint that refine agents never auto-apply an escalation is kept.
What: widen scripts/validate.sh's F03 pattern to catch the person-routing phrasings it missed (plural "human calls", "question left for the user", "the user asks/says/accepts", "prompt the user", "surface to the maintainer", "AskUserQuestion"), guard it with positive and negative self-test samples, skip YAML frontmatter (a description says when a person invokes the skill, which is before the go), and rewrite every line the widened pattern now catches in the factory phase copies to the unattended policy. Adds two eval tests planting newly covered phrasings, and corrects the scope comment's wrong justification for excluding ci-parity.md and jira.md. Why: review_1.md blocker B7 - C-factory-unattended was stated as a guarantee but its guarantor regex missed live person-routing text already inside its own scan root, so validate.sh stayed green while phases still routed decisions to a person.
Why: mutation testing on run-state.py left 11 survivors after the attempt subcommand and the new exit-code ladder landed. Five were real gaps: a `return None` mutant on `return BAD_CALL` in cmd_handoff, cmd_record, cmd_check_result and cmd_attempt turns exit 3 into exit 0, so a corrupt state file or an impossible call would be reported to the orchestrator as success, and stop_kind's `return ""` fallthrough was unobserved. No test covered those error paths. Adds seven subprocess tests: handoff over a truncated and over a non-object state file, record with stdin JSON that parses but is not a decision object, check-result on a JSON array, check-result on a stopped result whose `stop` is neither dict nor str (exact stdout `stopped: `), and attempt over a truncated and over a non-object state file. Survivors now 6, all `return 0` sites where `SystemExit(None)` also exits 0.
Why: the ship re-review found round 1's `phases` writer inert - R7 nothing read it back, so a resume relaunched as attempt 1 and check-result'd the stale build-1.json left by the dead session; R8 the attempt was recorded after a launch call that blocks until the phase ends, so a mid-phase crash left the list empty and indistinguishable from never started; R9 section 2 resumed "the state file not marked done" when no writer produced any terminal marker. Section 2 now reads the history with `run-state.py show` before launching anything and branches on the last attempt's status: a dangling `launched` with its own result file present is picked up at check-result, one without is closed `failed` with no `--result` and relaunched with guidance to start from the artifacts on disk. Section 5 records the launch first and takes the attempt number from what `run-state.py attempt` prints, so the state file is the number's only owner. The run's terminal marker is the `decisions` list ending in `end` or an `advance` on `ship`, which the state can already answer and which factory-run.md now documents beside the attempt numbering.
Note: the content of this change is already in ac04e2d - a concurrently running agent staged and committed the whole tree, including these files, while this change was being staged. This commit records the intent and the Why; nothing is added on top. What: rewrite the factory gauntlet's threshold line so a run never amends `docs/decisions.md` or the threshold it is judged by - a threshold it cannot meet is a `failed` result naming the finding, the score, and the missed threshold, or an `auto_decided` proposal for review. Rewrite mocking.md's "a seam worth agreeing on with the user" line to the unattended policy. Restate factory-run.md's policy prose so it states the rule instead of reading as a routing instruction, and record that `run/SKILL.md` is the only file that may use the interactive-only markers. Widen the F03 pattern from an allow-list of its own samples into routing constructions over a much larger person vocabulary (person, requester, owner, stakeholder, someone, whoever, team, them), with a negated-clause test so "no phase skill asks a person anything" stays clean, plus branches for gate nouns, approval-seeking and waiting for an answer: 45 of 60 probed phrasings were missed before, 0 after. Make the markers checkable - unbalanced, stray, and out-of-file markers are F03 failures, and a marker sharing its line no longer silences that line. Extend SAMPLES and NON_SAMPLES to everything newly covered and newly excluded, and add eval tests for the new shapes and every marker failure. Why: the ship re-review, items R10-R13. Round 1 turned a human gate into self-acceptance (a run could raise a threshold it had failed and write the loosened gate into the committed decision ledger as precedent for later runs), left a person-routing line inside F03's own scan root, widened the pattern only as far as its own samples, and left an unclosed marker able to silence the rest of a file.
Why: the ship re-review found six blockers in the exit ladder round 1 introduced. R1: argparse usage errors exited 2, the ladder's unappealable hard stop, so a mistyped phase or status ended an unattended run as if a secret had been found; a Parser subclass now reports usage errors as BAD_CALL. R2: check-result called a missing or unparseable result file a bad call while factory-run.md, run/SKILL.md and the spec's Error handling all call it a failed attempt; it is now exit 1 with the reason, since a subagent dying before writing its result is a phase outcome the orchestrator can relaunch. R3: load_state validated only the top level, so a null or list-shaped phases or scenario_texts crashed out as exit 1, colliding with diff-spec's drift code; the shape the commands rely on is validated and reported as BAD_CALL. R4: the hard-stop leniency covered stop but not status, so a secret reported alongside status done or failed was not read as a stop; a present stop kind now wins over status. R5: attempt silently rewrote a closed attempt and recorded an unreadable result as null while printing success; reclosing is refused and an unusable result file is recorded as a failed attempt with its reason. R6: tests for a second attempt on one phase and for every behavior above.
The ship gauntlet's ruff pass over the diff found 23 findings on added lines. Fixed at the source rather than suppressed: - re module aliases spelled out (re.S, re.M, re.I -> re.DOTALL, re.MULTILINE, re.IGNORECASE) in architecture-check.py, pr-evidence.py and aggregate-findings.py. - explicit check= on every subprocess.run that reads returncode itself. - skill-metrics.py numstat() leaked a file handle counting untracked lines; it now reads under a context manager. - two adjacent f-strings inside a collection literal made explicit as one value in skill-metrics.py and pr-evidence.py. Neither was a bug: both sites build a one-element row and a three-element tuple would have crashed the consumer. - the unused REJECTED constant in lint-spec.py deleted; the glyph is still matched through the ALTERNATIVE character class. The dev and factory copies of each shared script stay byte-identical, and plugins/ is regenerated rather than hand-edited.
The gauntlet's coverage-weighted complexity check was scoring every script at 0% coverage, which collapses each score to complexity squared and fails functions that are in fact well covered. The cause was the measurement, not the code: this repository's tests drive their subject scripts through subprocess.run([sys.executable, script, ...]), and a plain `coverage run -m unittest` traces only the parent process. coverage_run.sh wires coverage's documented subprocess support: a sitecustomize.py on PYTHONPATH calls coverage.process_startup() in every child, COVERAGE_PROCESS_START names the config, and both data_file and source are absolute so a child running in a temp cwd still writes to one place and still matches the sources. run-state.py measures 99.5% with it, against 0% before, and the README's gauntlet recipe now uses it.
test_factory_checks took 210.7s; it now takes 30.8s, and the full suite drops from about 215s to 35.3s. The cost was never the repo copy or the Codex rebuild, which measure 0.05s and 0.07s. It was that every test ran the whole gate, and 4.6s of each 7.4s run was the copy's own check_factory_script re-running the other 60 unit tests, which say nothing about F02 or F03. Twenty-seven gate runs, three of them the same clean run, one test running validate twice over identical state, and ten more run one after another inside a single subtest. Each distinct scenario is now named once in SCENARIOS and they run concurrently, capped at the core count. Every assertion is preserved: disabling check_factory_unattended in a copy still fails 19 of 20 tests. The scenarios that only assert about a clean gate share one run, and no scenario executes twice. The futures live at module scope rather than in a class fixture because tools/harden/flaky.py shuffles tests into one flat suite, and unittest re-runs setUpClass every time execution crosses a class boundary. A class fixture re-launched the whole fan-out many times per pass, which made a 5-run sweep take 20m51s; at module scope it is once per process. The two tests that never needed a gate run moved to their own classes.
Why: scope-review's phase-5 promotion step wrote these during the spec-review run but did not commit them; committing now, before build, so the ledger state build implements against is on disk and reviewable on its own. Adds 37 decisions under "## Config-driven factory" to docs/decisions.md, supersession notes on five factory-plugin entries, and three drafted contracts to docs/contracts.md with "? verify" marks that change set 7 of .dev/config-driven-factory/spec.md resolves. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 1 of config-driven-factory/spec.md: a strict shallow-YAML reader/writer for .factory/config.yaml, the resolve/check/init/show/set/unset CLI, and factory/references/pipeline-config.md describing the schema. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 3 of config-driven-factory/spec.md: run-state.py seeds an ordered phase list, budgets and a ceiling, records attempt types, skill paths and timestamps, seals arbitrary files at handoff, adds diff-config and reports terminality. Adds the stub pipeline fixture and its tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 4 of config-driven-factory/spec.md: run/SKILL.md, judgment.md and launch.md read phases, types, checks and budgets from factory-config.py instead of naming four phases. Built-in check executables become real plugin-root script paths and check now requires them to exist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 5 of config-driven-factory/spec.md: the four phase bodies move to factory/phases/ so run is the plugin's only invocable skill. The generator, F02 and F03, the path resolvers, docs and manifests follow the move, and F02 now holds every body and the launch wrapper to the result contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
Change set 6 of config-driven-factory/spec.md: factory-config.py inject copies named phases and what they reference into .factory/, records provenance and sha256s in .factory/.inject.json, refuses to overwrite edited copies without --force, and points the config at the copies. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
…ants Change set 7 of config-driven-factory/spec.md: docs/contracts.md carries the three new factory contracts with their provenance resolved, and the evals README gains the custom-pipeline and foreign-phase manual variants. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G9Voi18VRysPejexYsV7g4
What: run/SKILL.md now says explicitly that once past the go, every remaining phase runs to completion in the same turn, one launch after another, with no interim progress report or check-in; ending the turn while run-state.py show still prints finished: no is a defect unless a real host constraint forces it, in which case say so and print the resume command. Step 7's advance also says to go straight back to step 1 for the next phase instead of stopping. Why: this design has no background runner anymore - run is one skill looping through phases inline in a single orchestrator turn - and nothing forbade ending that turn between phases. A run had stopped after scope-review instead of continuing into build, with nothing in the instructions to prevent the model from treating a routine phase completion as a natural stopping point.
…-selection What: a type (or a phase, overriding its type) may declare a `model` string, resolved by factory-config.py and printed by show --resolved. The orchestrator passes it to the host's subagent tool (launch.md's transport ladder) when that tool takes a model parameter, and it is silently unused when the host has none. check refuses model on an interactive type, since that phase runs inline and never reaches a subagent call to carry it to. The built-in pipeline declares no model for any type: the valid values differ by host (Claude Code's Agent tool vs. Codex's spawn_agent vs. opencode's task), so no single default survives being run from a different host than the one it was written for. Pin one per repository with factory-config.py set. Why: a real run inherited Codex/GPT-5 for every phase, and the user expected the removed runner's own per-phase defaults. D-model-selection in docs/decisions.md had deferred this with an explicit reopen condition - "a real run shows one phase needs a different model than the session" - which this is; superseded by D-model-per-phase, which also records why host-per-phase (mixing Claude Code and Codex in one run) cannot come back under the current one-session orchestrator.
What: dropped "zero surviving mutants in scope" from the gauntlet's threshold defaults in both dev/skills/ship and factory/phases/ship (and their generated Codex copies). Fix agents still chase survivors under the existing two-round loop and equivalent-mutant carve-out, and a stubborn survivor is still recorded with its reason - it just no longer fails the run on its own. The ship report's killed/survived/ uncovered counts are the evidence a person reads; there is no score below which shipping is blocked. Recorded as D-mutation-threshold in docs/decisions.md, since no prior decision had actually set this default - it was unexamined. Why: a real run scored 43% (430/1003) mutants killed on a diff-scoped file and was ruled unshippable purely on that inherited coverage debt. The user's own words: "we aim to do our best but zero is a bit too harsh and unrealistic."
What: a third plugin, bootstrap, with one skill, agents-md. It probes a consuming repository read-only (manifests, task runners, CI workflows, test layout, git history), maps what it finds onto a closed concept map distilled from the dev skills (test layers, mocking at boundaries, the e2e environment, CI parity, fix-at-source quality, commits, plans, PR evidence), interviews only for what the probe cannot settle, runs the commands it will list, merges and prunes an existing AGENTS.md and CLAUDE.md, and writes the root AGENTS.md after the user approves a diff. Gaps become open items; nothing is scaffolded, CLAUDE.md is never written, and nothing is committed. check-agents-md.py is the gate every draft loops against: required sections in order, the seven Commands slots filled or backed by an open item, a closed open-item slug vocabulary, stale backticked paths and links, commands whose package script, make target, just recipe, task, or file does not exist, and a default branch that disagrees with origin/HEAD. Checks it cannot settle statically print notes instead of violations, so a false positive never traps the loop. 60 unittest cases cover it and tie the concept map, the template, and the checker together; validate.sh runs them as B01. Wiring: marketplace entries for Claude Code and Codex, a generator entry that copies only skills/, the generated plugins/bootstrap tree, and a README row. Pi stays dev-only. Why: the dev workflow's lessons only reach repositories that install dev, and the repo facts its skills need (per-layer commands, the e2e launch and environment, the merge gate, the scope vocabulary) are re-probed on every run because no file declares them. A root AGENTS.md written from a probe carries both to any agent. Three functional runs on scratch copies of sibling repositories drove the checker fixes (`bun turbo`, test path filters, naming conventions) and are recorded in bootstrap/evals/results.md. Considered: a dev skill (grows dev's namespace and the Pi package for a one-time step); linking into dev/references (dangles in the generated Codex tree); a CLAUDE.md pointer file (Claude Code reads AGENTS.md directly).
What: docs/architecture.md now names three plugins, adds a bootstrap component row, a "Bootstrapping a repository" flow, B01 in the validation flow, the root AGENTS.md at the consuming-repository boundary, and the checker as an entry point. docs/decisions.md gains a Bootstrap plugin area: D-bootstrap-plugin, -skill-name, -reach, -merge, -distilled-reference, -checker, and the not-doing entries -metrics, -pi, and -dev-reads-agents-md. CLAUDE.md describes the third plugin and says the Codex generator runs after changing any source plugin. Why: the overview and ledger are what future specs in this repository start from; D-packaging had rejected a third plugin for a different product, so the ledger records why this one differs, and the deferred follow-up of scope and build reading AGENTS.md's Commands section.
What: build now runs a spec's Validation block once per wave, as the
gate before the wave's commits ("The wave gate" in parallel.md). A
change-set implementer runs only the test files it adds or edits plus
typecheck and lint. Commands marked `(end of build)` in the block, and
the e2e suite, benchmarks, and suite-repeating analyses even when
unmarked, run once on the final tree with the CI-parity gate.
ci-parity.md no longer re-runs a command that is green on the same tree.
The seen-red rule keeps its guarantee and loses its cost: free when the
test is written first, otherwise one break per slice running one test
file.
change-set-brief.py (new, build/scripts) cuts spec.md down to one change
set: every section but research and the change plan, the decisions the
change set links, its own plan, and implementation-notes.md without its
test and seam inventories. Lines are verbatim. parallel.md hands each
agent its brief instead of the whole spec and notes.
lint-spec.py fails a change set over 25 scenarios (change sets already
logged in implementation-notes.md are exempt, since they never
renumber) and prints, on a clean spec, the build waves the file lists
allow and the shared files that make a change set wait. scope asks for
a wide plan, one owner per shared file, and the `(end of build)` mark;
the scope-review feasibility lens checks both.
check-tests.py splits the `Tests added:` line only where a path::name
follows the comma, so a test name may hold commas.
The factory phase copies carry the same edits, plugins/ is regenerated,
dev goes to 3.3.0, and validate.sh gains D01, which runs the 22 new
unit tests under dev/evals/tests. docs/decisions.md gains a Build speed
area: D-wave-gate, D-seen-red-cost, D-change-set-size, D-wave-report,
D-change-set-brief.
Why: feedback on the contexia searchable-kb build named five causes of
a slow build. The Validation block (lint, typecheck, build, unit,
integration, e2e, bench, and a complexity script that re-runs the suite
with coverage) ran more than once per change set. One change set spent
47 break-and-rerun cycles proving 75 tests red. Change set 4 carried 38
scenarios over about 60 files and took 73 minutes. Change sets 3, 4,
and 5 queued behind shared files. Every agent read about 130 KB of spec
and notes before writing anything; the brief for change set 5 is 42 KB.
The comma split was costing loops too: 75 of the 91 problems
check-tests.py reported on that plan were test names cut in two.
An old-versus-new run of the parallel-wave eval is recorded in
dev/evals/results.md: the Validation block ran 2 times instead of 4 and
no subagent ran it. Wall time did not improve on that fixture, whose
suite costs 20 seconds; a real build has to show the saving.
Considered: running the block once per build (commits in between would
be unverified); a file-count limit (file lists are prose, the count
would be a guess); failing the lint on a serial plan (no threshold
separates a careless chain from a necessary one); worktrees for
overlapping change sets (an overlap usually is a real dependency).
Directive: dev/skills/{scope,build}/scripts and their copies under
factory/phases must stay byte-identical; FactoryCopiesTest checks the
three scripts this change touches.
# Conflicts: # dev/evals/results.md # dev/evals/tests/test_build_speed.py # docs/architecture.md # docs/decisions.md # scripts/validate.sh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Brings back the
factoryplugin and adds thebootstrapplugin, plus the work that came with them (30 commits, about 200 files).run) that drives a pipeline of typed phases declared in.factory/config.yaml, or the built-in scope, scope-review, build and ship. It judges each phase's result itself and ends in a pull request or a report. Includesfactory-config.py,run-state.py, the phase bodies underfactory/phases/, and the generated Codex distribution.agents-mdskill, which probes a repository and writes its root AGENTS.md, gated bycheck-agents-md.py.validate.shgains checks for the factory's unattended wording and result protocol (F01 to F03) and for the bootstrap checker (B01).Read before merging
fa9cddfon main removed the factory plugin because the implementation "was not working out". This PR reverses that, so it needs a decision, not just a review..dev/factory-plugin/review_2.md, not committed) has verdict BLOCK. It was written before the later factory fixes and has not been re-run.scripts/validate.shpasses on this branch with main merged in. No end-to-end factory run is recorded in the repository.🤖 Generated with Claude Code