Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion dev/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "dev",
"version": "3.4.0",
"version": "3.5.0",
"description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.",
"author": {
"name": "Tobrun"
Expand Down
3 changes: 2 additions & 1 deletion dev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,9 @@

Development workflow skills for Claude Code, Codex, opencode, and Pi, built around two ideas: layered tests are the enforceable spec for behavior, and every phase produces something a human actually reviews as HTML, not markdown scrolling.

The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a test that has been seen to fail, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The skills chain loosely rather than as a rigid pipeline: `/scope` interviews for the real problem, argues every design decision against alternatives, and writes a self-contained spec whose change plan carries layer-tagged test scenarios; `/scope-review` puts the settled spec through a fresh-context, adversarially verified agent panel that checks the plan against the actual repo and refines the spec in place, looping without a human and closing with a short interview for the few findings only the user can decide, so a finished run hands `build` a spec ready to implement; `/build` executes the spec's change sets across unit/integration/e2e, proving every scenario with a real test at its tagged layer, in parallel waves where file lists allow, keeping a running implementation-notes log; `/ship` runs a deterministic quality gauntlet - the repo's own static analysis, security scan, dead code, duplication, dependency rules, coverage-weighted complexity, flakiness, mutation testing - looping fix agents until the checkers pass, then verifies the result with a fan-out review panel that checks spec conformance, e2e coverage, and logged deviations; `/commit` groups pending changes into granular commits with structured what/why messages; `/to-pitch` and `/to-quiz` turn finished work into a buy-in doc or a comprehension check.
The durable context is deliberately small: the code, its tests, the active spec under `.dev/{plan-name}/`, and three repo-tracked registries the skills maintain in the consuming project - `docs/decisions.md` (design decisions with their argued alternatives, read only after a review forms its findings), `docs/contracts.md` (boundary guarantees, read as premises before a review walks the diff), and `docs/dependencies.md` (machine-checkable module dependency rules, enforced by `ship`).
The files under `.dev/{plan-name}/` are written as a run goes, not when a stage closes: the spec opens during the interview, a report opens before its panel returns, the implementation notes gain an entry per change set and per fixup, and the PR body fills check by check, so a run can be followed from its files and a dead session loses only what was in flight.
Alongside them, `docs/architecture.md` is a plain high-level overview of the system - components, flows, boundaries, entry points - captured in full the first time a skill needs it and finds it absent, then kept current by build and commit whenever the structure changes, with a small checker that catches stale paths and files no component covers.
Every producing skill renders its own output as self-contained HTML under `/tmp/{project-slug}/reports/`. It publishes only when the user requests a shareable link and the host provides an artifact-publishing tool.
Every skill is explicit-invocation only: Claude Code and Pi use `disable-model-invocation: true`, the generated Codex distribution uses `agents/openai.yaml` with `allow_implicit_invocation: false`, and opencode enforces it with a `permission.skill` rule set to `ask` (see opencode installation below). Skills recommend the next step rather than launching each other.
Expand Down
2 changes: 1 addition & 1 deletion dev/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Eval definitions for the `dev` plugin's skills: realistic prompts and objective

- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 7 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz`.
- `results.md` - the record of the most recent full run: scores, methodology, and findings.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`), run by `scripts/validate.sh` as check D01.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`, `check-tests.py`), run by `scripts/validate.sh` as check D01.

`build` runs in `"functional"` mode (a real fixture, a real subagent run, assertions checked against the actual output).
`scope`, `scope-review`, `commit`, `ship`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.
Expand Down
4 changes: 3 additions & 1 deletion dev/evals/build.json
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
"The spec's Validation block is run once per change set as the gate before its commit, not repeated on an unchanged tree",
"No test is run just to be seen failing and no working code is broken to prove a test red: each test is written with its slice and run to green",
".dev/{plan-name}/implementation-notes.md is created and has one entry per change set",
"implementation-notes.md opens with a Build run entry naming the Validation commands and the waves, written before change set 1 is implemented, and change set 1's entry is in the file before change set 2's work starts",
"After both change sets are done, the user is asked to run ship rather than a review panel or gauntlet being launched automatically",
"Each implementation-notes.md entry carries a Tests added: line naming real path::test name references",
"scripts/check-tests.py is run against the plan directory and looped on until it exits clean, before the e2e pass"
Expand Down Expand Up @@ -71,6 +72,7 @@
"The failing e2e scenario is diagnosed and fixed rather than reported as an accepted failure or a limitation",
"The fix adds a test at the cheapest layer that can catch the bug, rather than only patching until the e2e passes",
"The scenario is re-run and re-captured after the fix, not flipped to pass with the original capture",
"implementation-notes.md gains a Fixup entry for the fix, written with the fix rather than at the end of the run, naming the e2e scenario that found it and the test added",
"The final report's summary counts match the scenarios array and reflect the passing re-run",
"The user is asked to run ship only after the e2e run is green"
]
Expand All @@ -85,7 +87,7 @@
"The screenshot failure is diagnosed and fixed rather than labeled pre-existing, flaky, unrelated, or an accepted deviation",
"The fix removes the time-dependent assumption without adding a retry, sleep, timeout increase, or looser assertion",
"The full screenshot-matrix command is rerun and green before the PR question",
"The discovered CI commands and outcomes are recorded in implementation-notes.md",
"The discovered CI commands and outcomes are recorded in implementation-notes.md, each as its command finishes",
"If the user approves a PR, required checks are watched to a terminal state and a deterministic failure is fixed and pushed rather than merely reported as restarted"
]
},
Expand Down
39 changes: 38 additions & 1 deletion dev/evals/results.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Eval Results

Status: no full run recorded for the current skill set; one partial run below.
Status: no full run recorded for the current skill set; two partial runs below.

The last recorded run (2026-07-24) covered the pre-pivot five-skill chain and was invalidated by the pivot to the scope pipeline; its scores were removed rather than left to invite false confidence.
Run the harness below against the current 6 skills (`scope`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz`) and replace this file with the dated results.
Expand Down Expand Up @@ -32,6 +32,43 @@ Findings:
- Wall time did not improve on this fixture: the suite costs 20 seconds, so two saved runs are under a minute, less than the difference between two model runs. The saving grows with the cost of the block and the number of change sets; a real build is needed to measure it.
- The old run's notes named a test with a comma in it, which `check-tests.py` then split in two. The checker now splits only where a `path::name` follows the comma.

## Partial run, 2026-10-02: plan files written as the run goes

Old is commit c32e4fb, new is the version that opens each plan file early and writes it as the run goes.
One subagent per skill and version answered that skill's comprehension prompts from the skill text and its linked references.

| Skill | Eval | Old | New |
| ----- | ---- | --- | --- |
| scope | `scope-confidence-threshold` | 4/4 | 4/4 |
| scope | `scope-written-as-the-run-goes` | 0/5 | 5/5 |
| scope-review | `scope-review-contract` | 3/4 | 4/4 |
| scope-review | `scope-review-loop-limits` | 4/4 | 4/4 |
| scope-review | `scope-review-report-as-the-run-goes` | 0/5 | 5/5 |
| ship | `gauntlet-policy` | 8/8 | 8/8 |
| ship | `review-policy` | 10/10 | 10/10 |
| ship | `ship-files-as-the-run-goes` | 2/6 | 6/6 |

The old `scope-confidence-threshold` run was graded against the old wording of its answer 3, and the old `scope-review-contract` run missed only the report in answer 1, which the old text did not name.
The two `ship-files-as-the-run-goes` answers the old text got right are the two whose answer is "no" in both versions.

One functional `build` run on the new text only: a Node CLI fixture with two change sets, the second consuming the first, and a launcher that strips trailing zeros in the built CLI.
A watcher logged every change of the headings in `implementation-notes.md` next to the commit count.

| Time | Commits | Notes file |
| ---- | ------- | ---------- |
| 15:56:22 | 1 (fixture) | absent |
| 15:58:41 | 1 | `## Build run` entry: Validation commands, waves `[1] [2]` |
| 15:59:20 | 2 | plus `## Change set 1` |
| 16:00:41 | 3 | plus `## Change set 2` with four deviations |
| 16:01:39 | 5 | plus `## E2E pass` and `## CI parity`, one line per command |

Findings:

- The run entry was in the file before any change set was committed, and each change set's entry was in the file before the next one started; `check-tests.py` exited clean.
- The fixup entry was not exercised: the agent found the launcher bug while exploring, fixed it inside change set 2, and logged it as a deviation there, so the e2e pass was green on its first run. Only the unit tests in `tests/` cover the fixup entry.
- The `ship` comprehension run showed that nothing forbade writing unverified findings into the open report, and that the question about the Quality rows assumed dead code and duplication run one after the other. The skill now says a finding enters the report only once verified, the gauntlet writes one row per check for the batched five, and the question was reworded.
- Answer 7 of `gauntlet-policy` still expected `commit` to be recommended at every wrap-up; it now says that holds only when phase 3 did not run.

## Re-running this harness

1. Pick a baseline commit (the last commit before the change under test) and the working tree as "new".
Expand Down
13 changes: 12 additions & 1 deletion dev/evals/scope-review.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
"id": "scope-review-contract",
"prompt": "Answer briefly, from the skill text only: (1) What may scope-review edit, and what must it never touch? (2) Can a finding reach the spec without being verified, and can the orchestrator apply its own opinion as a refinement? (3) What happens when lint-spec.py reports problems before the loop starts? (4) A verified finding shows a change set implements the S3 alternative that its linked decision rejected in favor of local disk - is that refinable, and if so what does the refine agent write?",
"assertions": [
"Answer 1: it edits spec.md and nothing else - never code, never docs/decisions.md, never docs/contracts.md",
"Answer 1: it edits spec.md, keeps its own spec-review_N.md report, and promotes settled changes to docs/decisions.md and docs/contracts.md - never code, and never any other file",
"Answer 2: no - findings reach the spec only through verification (REFUTED findings drop, a BLOCK survives only when CONFIRMED), and the orchestrator's own reading is not a lens",
"Answer 3: stop and recommend finishing the scope run - refinement presumes a mechanically settled spec, and repairing an unfinished draft is scope's job",
"Answer 4: refinable - a settled chosen decision outranks the spec's prose in the authority order, so the change set is rewritten to match the decision (local disk), not the other way around"
Expand All @@ -22,6 +22,17 @@
"Answer 3: recorded user intent, then the repo's reality, then settled chosen decisions, then the spec's prose - each level beats everything below it",
"Answer 4: APPROVED when every finding was refined or answered (build can start directly); APPROVED WITH DEFERRALS when something remains - an escalation defers only when the answer invalidates the premise or opens a genuinely new effort, when there is no interactive channel, or when the user declines to answer"
]
},
{
"id": "scope-review-report-as-the-run-goes",
"prompt": "Answer briefly, from the skill text and the references it links only: (1) When is spec-review_N.md created, and what does its Verdict line say then? (2) Round 1's refine agents have just returned with four refinements applied, and round 2's panel is about to start. What does the report hold now? (3) The user answers the first of three escalations. When is that recorded in the report? (4) What is left to write into the report at the wrap-up? (5) A later scope run finds a spec-review_2.md whose verdict is still IN PROGRESS and no scope-review run is going. May it take the report's findings as its opening agenda?",
"assertions": [
"Answer 1: in step 1, after the lint gate passes and before round 1's panel runs, at the next free index; Verdict: IN PROGRESS",
"Answer 2: round 1's finding counts and each of the four refinements, written when the refine agents returned - not held back until the end of the run",
"Answer 3: at once, when it is answered (or deferred) - not after all three escalations are done",
"Answer 4: only the verdict - everything else has been filling since step 1",
"Answer 5: no - an unfinished file left by a run that is no longer going is a draft, never a result; scope-review picks it up at the same index and no other skill builds on it"
]
}
]
}
13 changes: 12 additions & 1 deletion dev/evals/scope.json
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,18 @@
"Answer 0: no question; the table is applied and its chosen line carries `auto-applied at Confidence: 82%` in its because clause",
"Answer 1: yes, 55% is below the 75% threshold, so the user is asked",
"Answer 2: yes, whatever the score, because it changes what was asked for (a different problem than the request names)",
"Answer 3: once, in one block at the end of the interview, so a single reply can overturn any of them before the spec is written"
"Answer 3: once, in one block at the end of the interview, so a single reply can overturn any of them before the catalog builds on them"
]
},
{
"id": "scope-written-as-the-run-goes",
"prompt": "Answer briefly, from the skill text and the references it links only. (1) At what point in a run is spec.md created, and what does it hold at that moment? (2) You have just cataloged nine decisions and are about to talk the big ones through with the user. Where are those nine decisions right now? (3) The user answers a question that settles D-file-storage. When does the spec change? (4) The blind spot subagent is hunting through the repo while you catalog. What keeps it from reading your catalog? (5) The session dies halfway through the decision talk-through. What does the next scope run find, and how does it treat it?",
"assertions": [
"Answer 1: during the interview, once the first answers say what the change is about - not after the interview or the catalog; it holds the title, the date, what was understood so far, and empty Research, Scope, and Change plan sections",
"Answer 2: in the research section of spec.md, each written as an [open] entry the moment it was cataloged - not only in the conversation",
"Answer 3: at once, when the answer lands - the decision's marks change in the file as the talk-through settles it, not in a later write-up",
"Answer 4: it is handed only the original request and the relevant code, and is told to stay out of .dev/, where the catalog is being written",
"Answer 5: a spec.md that lint-spec.py does not pass, holding everything established before the session died; it is a draft that scope picks up where it stopped, and no other skill builds on it"
]
}
]
Expand Down
Loading
Loading