Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
ac4388c
feat(factory): plugin skeleton and packaging
tobrun Sep 20, 2026
6929e39
feat(factory): protocol reference and shared copies
tobrun Sep 20, 2026
1115104
feat(factory): phase skill copies
tobrun Sep 20, 2026
8d4b12c
feat(factory): the run skill and run-state.py
tobrun Sep 20, 2026
c6af66b
feat(factory): validate.sh unattended, protocol, and script checks
tobrun Sep 20, 2026
aeca92a
feat(factory): fixture and end-to-end procedure
tobrun Sep 20, 2026
f52285f
fix(harden): declare framework entry points in a vulture whitelist
tobrun Sep 20, 2026
ab98fac
fix(factory): make run-state.py exit codes safe and record attempts
tobrun Sep 20, 2026
c0d50b2
fix(factory): drop deferrals from the scope-review copy
tobrun Sep 20, 2026
196ac73
fix(factory): enforce C-factory-unattended in the F03 pattern
tobrun Sep 20, 2026
ecb5f3e
test(factory): cover run-state BAD_CALL paths and empty stop kind
tobrun Sep 20, 2026
ac04e2d
fix(factory): make a crashed run resumable from factory-run.json alone
tobrun Sep 20, 2026
0e8f550
fix(factory): keep an unattended run from moving its own gates
tobrun Sep 20, 2026
f056f3d
fix(factory): reconcile run-state.py's exit ladder with the protocol
tobrun Sep 20, 2026
00d5226
fix(harden): clear the static analysis findings on the branch
tobrun Sep 20, 2026
a960f18
feat(harden): measure coverage of subprocess-invoked scripts
tobrun Sep 20, 2026
0725449
perf(factory): run the gate scenarios once each, concurrently
tobrun Sep 20, 2026
fb0d00c
docs: promote config-driven-factory decisions and contracts
tobrun Sep 21, 2026
f5e4332
feat(factory): add factory-config.py, the pipeline config reader and CLI
tobrun Sep 21, 2026
eaa4e7a
feat(factory): run state carries the declared pipeline
tobrun Sep 21, 2026
2d2d925
feat(factory): the orchestrator drives a declared pipeline
tobrun Sep 21, 2026
671440f
refactor(factory): phases stop being invocable skills
tobrun Sep 21, 2026
c93ac9c
feat(factory): inject phase bodies into a consuming repository
tobrun Sep 21, 2026
92723cd
docs(factory): land the config-driven contracts and the per-host vari…
tobrun Sep 21, 2026
6f57000
fix(factory): stop the orchestrator from pausing between phases
tobrun Sep 22, 2026
ef549f1
feat(factory): add an optional per-type model axis, reopening D-model…
tobrun Sep 22, 2026
c250409
fix(ship): mutation testing is reported, not a shippability gate
tobrun Sep 22, 2026
d2eec7d
feat(bootstrap): add the agents-md skill that writes a repo's AGENTS.md
tobrun Sep 25, 2026
22cbfca
docs(bootstrap): record the bootstrap plugin in the overview and ledger
tobrun Sep 25, 2026
0130c2f
perf(build): gate each wave once, brief each agent, cap change sets
tobrun Oct 1, 2026
6c62847
merge: bring main's build-speed squash into factory-plugin
tobrun Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,30 @@
"authentication": "ON_INSTALL"
},
"category": "Developer Tools"
},
{
"name": "factory",
"source": {
"source": "local",
"path": "./plugins/factory"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Developer Tools"
},
{
"name": "bootstrap",
"source": {
"source": "local",
"path": "./plugins/bootstrap"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Developer Tools"
}
]
}
10 changes: 10 additions & 0 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,16 @@
"name": "dev",
"source": "./dev",
"description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes."
},
{
"name": "factory",
"source": "./factory",
"description": "An orchestrator skill that drives a pipeline of typed phases declared in .factory/config.yaml (the built-in scope, scope-review, build, and ship by default), judging each phase's completion itself, ending in a pull request or a report."
},
{
"name": "bootstrap",
"source": "./bootstrap",
"description": "Probe a repository's stack and write or refresh its root AGENTS.md, encoding the dev workflow's SDLC lessons (test layers, mocking at boundaries, CI parity, fix-at-source quality, commit and PR conventions) with the repo's exact commands and its gaps as open items."
}
]
}
15 changes: 11 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,14 @@ This is a monorepo for Tobrun's Claude Code, Codex, and Pi skills. The root
`.claude-plugin/marketplace.json` references the Claude source plugins.
`.agents/plugins/marketplace.json` references generated Codex plugins under
`plugins/`. The root `package.json` exposes the `dev` source skills as a Pi
package. There is one plugin: `dev`, the hand-invoked development workflow.
package. There are three plugins: `dev`, the hand-invoked development workflow;
`bootstrap`, whose `agents-md` skill probes a consuming repository and writes
its root AGENTS.md from the workflow's SDLC lessons; and `factory`, whose `run`
skill is the one sanctioned exception to "skills
never invoke each other" - it drives the pipeline declared in the consuming
repository's `.factory/config.yaml`, or the built-in `scope`, `scope-review`,
`build`, and `ship` phase bodies under `factory/phases/` by path, and judges
each phase's completion itself. `run` is the plugin's only invocable skill.

## Repository Structure

Expand All @@ -18,7 +25,7 @@ package. There is one plugin: `dev`, the hand-invoked development workflow.
│ └── marketplace.json # Marketplace manifest listing all plugins
├── .agents/plugins/
│ └── marketplace.json # Codex marketplace manifest
├── {plugin-name}/ # Individual plugin directory (dev)
├── {plugin-name}/ # Individual plugin directory (dev, factory, bootstrap)
│ ├── .claude-plugin/
│ │ └── plugin.json # Plugin metadata (name, version, author)
│ ├── README.md # Plugin documentation
Expand All @@ -45,10 +52,10 @@ package. There is one plugin: `dev`, the hand-invoked development workflow.
- All changes must pass `scripts/validate.sh` before committing.
- Every plugin directory name must match its `plugin.json` name and marketplace entry name.
- Every `SKILL.md` must have YAML frontmatter with `name` and `description`.
- Every `SKILL.md` must set `disable-model-invocation: true`; all skills in this repo are human-triggered only, and skills recommend the next step instead of invoking each other.
- Every `SKILL.md` must set `disable-model-invocation: true`; all skills in this repo are human-triggered only, and skills recommend the next step instead of invoking each other - except the factory `run` skill, the one sanctioned invoker, which launches its phase bodies (`factory/phases/`, not skills) by path.
- Plan files under `.dev/` are never committed.
- Do not edit `plugins/` directly. Run `python3 scripts/build_codex_plugin.py`
after changing `dev/`; the generator builds every configured
after changing any source plugin; the generator builds every configured
plugin, removes Claude-only frontmatter, and writes Codex
`agents/openai.yaml` invocation policy. It never copies `evals/`.
- Keep the root Pi package version equal to `dev/.claude-plugin/plugin.json`.
Expand Down
10 changes: 7 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,26 +3,30 @@
| Plugin | Use When | Tools |
| ------ | -------- | ----- |
| [dev](dev/) | A test-focused development workflow for Claude Code, Codex, opencode, and Pi. | `scope`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz` |
| [factory](factory/) | Take a request from scope to a shipped pull request unattended, on Claude Code or Codex. | `run` |
| [bootstrap](bootstrap/) | Prepare any repository for agent work: probe its stack and write its root AGENTS.md from the workflow's SDLC lessons, on Claude Code or Codex. | `agents-md` |

## Claude Code

```bash
/plugin marketplace add tobrun/workflow
/plugin install dev@nurbot
/plugin install bootstrap@nurbot
```

Invoke skills as `/scope`, `/build`, and so on, or namespaced as `/dev:scope`.
Invoke skills as `/scope`, `/build`, and so on, or namespaced as `/dev:scope`; bootstrap's skill is `/bootstrap:agents-md`.

## Codex

```bash
codex plugin marketplace add tobrun/workflow
codex plugin add dev@nurbot
codex plugin add bootstrap@nurbot
```

Invoke skills as `$dev:scope`, `$dev:build`, and so on.
Invoke skills as `$dev:scope`, `$dev:build`, `$bootstrap:agents-md`, and so on.
The distribution is explicit-invocation only.
The checked-in Codex package under `plugins/` is generated from `dev/`:
The checked-in Codex packages under `plugins/` are generated from each source plugin:

```bash
python3 scripts/build_codex_plugin.py
Expand Down
8 changes: 8 additions & 0 deletions bootstrap/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"name": "bootstrap",
"version": "0.1.0",
"description": "Prepare a repository for agent work: probe its stack, map it onto the dev workflow's SDLC concepts (test layers, mocking at boundaries, CI parity, fix-at-source quality, commit and PR conventions), and write or refresh its root AGENTS.md with the repo's exact commands and its gaps as open items. Use on any repository, whether or not the dev plugin is installed.",
"author": {
"name": "Tobrun"
}
}
42 changes: 42 additions & 0 deletions bootstrap/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
# bootstrap

Prepare a repository for agent work.
The `agents-md` skill probes the repository's stack, maps what it finds onto the SDLC concepts the [dev](../dev/) workflow is built on, and writes or refreshes the root `AGENTS.md`.

The file it writes encodes those lessons as short outcome rules any agent can follow, with or without the dev plugin installed:
- tests on the cheapest layer that can fail for the right reason, and mocks only at boundaries the repo does not own;
- e2e against the built artifact in a stubbed, seeded environment, never production or shared staging;
- done means the merge gate passes locally, and findings are fixed at the source, never suppressed or retried away;
- conventional commits with a what and a why, plans kept out of git, and pull requests that carry evidence.

It fills those rules in with the repository's exact commands (setup, build, check, unit, integration, e2e, and the merge gate), verified by running them, and records every missing concept as an open item instead of scaffolding it.

## What it touches

Only the root `AGENTS.md`.
An existing AGENTS.md or CLAUDE.md is merged and pruned: repo-specific gotchas are kept, stale facts are updated, and a rule that contradicts a lesson becomes a question rather than a silent overwrite.
CLAUDE.md is read but never written; Claude Code reads AGENTS.md directly, so the closing report recommends removing a now-redundant CLAUDE.md.
Nothing is committed; the diff is shown for approval before the file is written.

Every draft loops against `skills/agents-md/scripts/check-agents-md.py` until it exits clean: required sections in order, every Commands slot filled or backed by an open item, no stale path, no command whose script, target, or file does not exist, and the stated default branch matching `origin/HEAD`.

## Install

### Claude Code

```bash
/plugin marketplace add tobrun/workflow
/plugin install bootstrap@nurbot
```

Invoke it as `/bootstrap:agents-md`.

### Codex

```bash
codex plugin marketplace add tobrun/workflow
codex plugin add bootstrap@nurbot
```

Invoke it as `$bootstrap:agents-md`.
The distribution is explicit-invocation only.
15 changes: 15 additions & 0 deletions bootstrap/evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Bootstrap evals

Two levels of verification:

1. **Unit tests, free**: `python3 -m unittest discover -s bootstrap/evals/tests -t .`, run on every `scripts/validate.sh` invocation as check B01.
`test_check_agents_md.py` proves the checker's CLI contract against scratch repositories.
`test_concept_sources.py` keeps the concept map, the template, and the checker agreeing, and fails when a dev source file the map was distilled from moves or is renamed; it cannot see a source whose content changed in place.
2. **Functional runs, paid and manual**: `agents-md.json` lists the fixtures and assertions, and `results.md` records the runs.

## The functional procedure

1. For each eval, make a scratch copy of a repository matching its fixture with `git clone --local`, so history and `origin/HEAD` come along and the live repository is never touched.
2. Install the plugin (`/plugin install bootstrap@nurbot`, or `codex plugin add bootstrap@nurbot`) and invoke `/bootstrap:agents-md` in the copy with the eval's prompt.
3. Answer the interview as the fixture's owner would.
4. Grade the run against the eval's assertions and the `every_eval` list, then record the result in `results.md`.
67 changes: 67 additions & 0 deletions bootstrap/evals/agents-md.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
{
"skill": "agents-md",
"mode": "functional",
"note": "Run each eval on a scratch copy made with git clone --local (it keeps history and origin/HEAD), never on the live repository. results.md names the reference repository and commit behind each fixture.",
"evals": [
{
"id": "merge-prune",
"prompt": "Write this repository's AGENTS.md.",
"fixture": "an Astro blog with both an AGENTS.md and a CLAUDE.md; AGENTS.md cites a config path that has since moved and contains em dashes; no test runner; one GitHub workflow that only deploys on push",
"assertions": [
"The moved path is rewritten to the file that exists now, and the table cites the probe evidence",
"No em dash remains in the written file",
"Unit, Integration, and E2E are none, each with a matching open item, and there is an open item with slug ci because the deploy workflow is not a merge gate",
"CLAUDE.md is not modified, and the closing report recommends deleting or trimming it",
"The diff and the kept/updated/pruned table are shown before AGENTS.md is written"
]
},
{
"id": "claude-only-python",
"prompt": "Bootstrap an AGENTS.md here.",
"fixture": "a Python service with a Makefile whose check target runs lint, types, and tests, a pull-request workflow that runs make check, a robot framework suite that drives the built service, and a CLAUDE.md with two repo-specific gotchas but no AGENTS.md",
"assertions": [
"The Merge gate slot is `make check`",
"The robot suite is recognized as the e2e layer, and its environment is either described in Tests or recorded as an e2e-env open item",
"Both CLAUDE.md gotchas appear in AGENTS.md in the section they concern",
"CLAUDE.md is not modified, and the closing report recommends deleting or trimming it"
]
},
{
"id": "rich-existing",
"prompt": "Refresh our AGENTS.md against the current code.",
"fixture": "a bun monorepo with a curated AGENTS.md that says tests must not run from the root, an e2e script that needs a live SSO session and a real access token, and conventional commits with scopes such as core, mapbox, server, and tui",
"assertions": [
"The rule that tests must not run from the root is kept",
"The live e2e run is recorded as an e2e-env open item rather than presented as a mocked e2e layer",
"The scope list names the scopes the history actually uses",
"A second run on the written file produces a diff limited to facts that changed"
]
},
{
"id": "bare",
"prompt": "Write an AGENTS.md for this repo.",
"fixture": "a repository holding a single index.html game and a README, no package manifest, no CI, a handful of unconventional commits",
"assertions": [
"The skill asks for commands it cannot find instead of guessing a runner",
"Every Commands slot the user cannot answer is none or n/a with a matching open item",
"The written file is under 60 lines"
]
},
{
"id": "conflict",
"prompt": "Update AGENTS.md.",
"fixture": "a JS repository whose AGENTS.md says to retry flaky e2e tests twice and whose playwright config sets retries: 2",
"assertions": [
"The retry rule is raised as an interview question with a recommendation and a confidence score",
"The retry rule is neither silently kept nor silently removed",
"The written file does not state both the retry rule and the no-retries rule"
]
}
],
"every_eval": [
"check-agents-md.py exits 0 on the written AGENTS.md",
"No file other than the root AGENTS.md changes in the repository",
"Nothing is committed",
"No question asks about a fact the probe table already resolved"
]
}
19 changes: 19 additions & 0 deletions bootstrap/evals/results.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Eval Results

Status: three functional runs recorded on 2026-09-25, before the skill's first release.

Every run executed the skill text directly in a Claude Code subagent (not through an installed plugin), on `git clone --local` scratch copies, with a simulated user who accepted every recommendation, skipped the closing newcomer question, and declined installs.

| Eval | Reference repository and commit | Host | Result | Notes |
| ---- | ------------------------------- | ---- | ------ | ----- |
| merge-prune | `~/ws/blog` at 0a55308, plus its untracked AGENTS.md | Claude Code (subagent) | pass (5/5 and every_eval) | stale `src/content/config.ts` updated, em dashes gone, test layers and ci recorded as open items, CLAUDE.md untouched and flagged, checker clean on the first iteration |
| rich-existing | `~/ws/mapcode` at 0eb9c91 | Claude Code (subagent) | pass (3/4; the re-run assertion was exercised on the blog copy below) | root-test rule kept, live e2e recorded as `e2e-env`, real scopes used; also caught `.dev/` not being gitignored and a stale models-generator claim |
| re-run (merge-prune output) | the blog copy after the first run | Claude Code (subagent) | pass | 70 of 72 lines kept verbatim; the 2 changes corrected an unconditional analytics claim that the code makes conditional |

## Fixes the runs drove

- The checker rejected `bun turbo` (a dependency's binary) as a missing script; `bun`, `yarn`, and implicit `pnpm` calls now fall back to declared dependencies and print a note.
- The checker now checks path filters on `test` subcommands, knows framework and asset extensions, notes unknown runners in Commands, and exempts naming conventions such as `kebab-case.mdx`.
- SKILL.md: the skill root is defined, installs of any kind need a question, a command whose dependencies are missing "cannot be run" rather than fails, and defects noticed while probing go to the closing report.
- References: scope fallback for short or non-conventional histories, e2e harnesses that drive source, configured retries as open items, conflicts limited to rules an agent file states, and approval tables grouped by section.
- A refresh starts from the existing file, the template's rule lines are exempt from pruning, and the approval table counts kept lines instead of listing them.
Empty file.
Loading
Loading