Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion dev/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "dev",
"version": "3.2.0",
"version": "3.3.0",
"description": "Development workflow skills: scope changes with argued decisions, build across unit/integration/e2e with every scenario proven by tests, ship with a deterministic quality gauntlet and an adversarially verified review, create structured commits that feed a decision ledger, and render pitches or comprehension quizzes.",
"author": {
"name": "Tobrun"
Expand Down
5 changes: 4 additions & 1 deletion dev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,8 @@ Every skill is explicit-invocation only: Claude Code and Pi use `disable-model-i

Specs a change by interviewing for the real problem behind the request, cataloging every design decision (with a subagent blind-spot pass on full-size changes), and arguing each one against alternatives in a `鉁揱/`鉁梎/`?`/`鈿燻/`鈯榒 notation with evidence marks.
Writes a self-contained spec at `.dev/{plan-name}/spec.md` - research decisions, scope with invariants and a Validation block of the repo's real commands, and a change plan of numbered change sets each ending in a layer-tagged `tests:` line - designed as a fresh-context handoff to `build`.
A checker (`scripts/lint-spec.py`) enforces the spec's mechanics - unique slugs, argued alternatives, echoes that match their decision, tagged test scenarios - so the prose stays about judgment.
A checker (`scripts/lint-spec.py`) enforces the spec's mechanics - unique slugs, argued alternatives, echoes that match their decision, tagged test scenarios, at most 25 scenarios per change set - so the prose stays about judgment.
On a clean spec it prints the build waves the file lists allow and the shared files that make change sets wait, so the plan is shaped for parallel work before build starts.
Promotes durable decisions to `docs/decisions.md` and cross-boundary invariants to `docs/contracts.md`, renders an expandable-card spec view, and has a reverse mode that audits the implicit decisions already embedded in existing code.

### scope-review
Expand All @@ -28,7 +29,9 @@ An APPROVED verdict means every finding was refined or answered: `build` can sta

Executes a spec's change sets at the layer each `tests:` scenario is tagged with - unit for business logic, integration for real cross-component seams, e2e for driving the actual running application.
Enforces outcomes rather than rituals: every test must have been seen to fail before its green counts, with strict failing-test-first reserved for bug fixes, where red is the proof the issue was actually reproduced.
The red is kept cheap: free when the test is written first, one break per slice and one test file otherwise.
Runs independent change sets in parallel as waves of subagents batched by disjoint file lists in spec order, committing each change set and appending to a running `implementation-notes.md` that logs any deviations forced by an edge case.
Each subagent starts from a brief that `scripts/change-set-brief.py` cuts from the spec for its change set and runs only its own tests; the spec's Validation block runs once per wave as the gate before its commits, and e2e, benchmark, and coverage commands run once at the end.
Once every change set is committed, drives the real app against a mocked environment, loops until every e2e scenario passes, then renders the e2e report: screenshots per scenario for frontend systems, Test Scenario and Data Model State tables for everything else.
Build never pushes or opens a PR; `ship` does, once the change is hardened and reviewed.

Expand Down
1 change: 1 addition & 0 deletions dev/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Eval definitions for the `dev` plugin's skills: realistic prompts and objective

- `{skill}.json` - one file per skill: the eval prompt(s), the fixture each expects, and the assertions to grade the output against. Covers all 7 skills: `scope`, `scope-review`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz`.
- `results.md` - the record of the most recent full run: scores, methodology, and findings.
- `tests/` - unit tests for the deterministic scripts the skills loop against (`lint-spec.py`, `change-set-brief.py`), run by `scripts/validate.sh` as check D01.

`build` runs in `"functional"` mode (a real fixture, a real subagent run, assertions checked against the actual output).
`scope`, `scope-review`, `commit`, `ship`, `to-pitch`, and `to-quiz` run in `"comprehension"` mode instead - each depends on either an interactive question loop, a live codebase, or prior artifacts (a finished spec, implementation notes, an e2e report) that are too expensive to stage on every iteration, so these check policy comprehension of the skill text directly.
Expand Down
7 changes: 6 additions & 1 deletion dev/evals/build.json
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@
"Tests are behavior-named (e.g. \"applies a discount\") rather than test1/test2, cover cases beyond the happy path, and no assertion recomputes the expected value the way the implementation does",
"The user is not asked about running ship after change set 1 finishes - no review prompt appears until both are done",
"Each change set gets its own commit on the current branch (at least 2 commits)",
"The spec's Validation block is run once per change set as the gate before its commit, not repeated on an unchanged tree",
"No test is proven red by its own break-and-rerun: tests are written first, or one break covers a slice and only the affected test file is run",
".dev/{plan-name}/implementation-notes.md is created and has one entry per change set",
"After both change sets are done, the user is asked to run ship rather than a review panel or gauntlet being launched automatically",
"Each implementation-notes.md entry carries a Tests added: line naming real path::test name references",
Expand All @@ -38,7 +40,9 @@
"The subagents do not commit, stage, or otherwise touch git state - the orchestrator makes the commits",
"The subagents do not write to implementation-notes.md themselves; the orchestrator appends each entry",
"The subagents are told not to launch the app or run e2e scenarios",
"The spec's Validation block commands are re-run by the orchestrator rather than the subagent's claim being taken at face value",
"scripts/change-set-brief.py is run for the wave and each subagent is handed its brief rather than told to read the whole spec.md and implementation-notes.md",
"The subagents are told to run only their own test files plus typecheck and lint, not the spec's Validation block",
"The spec's Validation block commands are run by the orchestrator, once for the wave, rather than the subagents' claims being taken at face value",
"Three commits exist, one per change set, in spec order",
"The orchestrator carries each subagent's testsAdded into that change set's Tests added: line rather than discarding it"
]
Expand Down Expand Up @@ -77,6 +81,7 @@
"fixture": "a repo whose spec Validation block runs unit/lint/build, while .github/workflows/pr.yml also requires a project-owned screenshot-matrix command. The active-month screenshot fails deterministically on the first UTC day, although the feature's tagged e2e scenarios pass.",
"assertions": [
"The pull-request workflow is inspected and the screenshot-matrix command is run before the user is offered a PR",
"A required command that already ran green on the unchanged final tree is recorded, not run a second time",
"The screenshot failure is diagnosed and fixed rather than labeled pre-existing, flaky, unrelated, or an accepted deviation",
"The fix removes the time-dependent assumption without adding a retry, sleep, timeout increase, or looser assertion",
"The full screenshot-matrix command is rerun and green before the PR question",
Expand Down
29 changes: 28 additions & 1 deletion dev/evals/results.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,37 @@
# Eval Results

Status: no run recorded for the current skill set.
Status: no full run recorded for the current skill set; one partial run below.

The last recorded run (2026-07-24) covered the pre-pivot five-skill chain and was invalidated by the pivot to the scope pipeline; its scores were removed rather than left to invite false confidence.
Run the harness below against the current 6 skills (`scope`, `commit`, `build`, `ship`, `to-pitch`, `to-quiz`) and replace this file with the dated results.

## Partial run, 2026-10-01: build `parallel-wave`, before and after the build-speed change

One functional run per version, not the full harness.
Old is the build skill before this change, new is the version that adds the wave gate, the change-set brief, and the cheap-red rule.
The fixture was a Node library with a 20 second legacy test in its suite, and a spec with change sets 1 and 2 disjoint and change set 3 editing files of both.
Every run of a Validation command was written to a log by the fixture's package scripts, so the counts are measured, not reported.

| Check | Old | New |
| ----- | --- | --- |
| Change sets 1 and 2 launched as subagents in one message | pass | pass |
| Change set 3 started only after 1 and 2 were committed | pass | pass |
| Subagents left git state and `implementation-notes.md` alone | pass | pass |
| Subagents told not to launch the app | pass | pass |
| Each subagent handed a brief from `change-set-brief.py` instead of the whole spec | fail (reads all of `spec.md`) | pass |
| Subagents run only their own test file plus syntax checks | fail (each ran the Validation block) | pass |
| Orchestrator runs the Validation block once per wave | pass | pass |
| Three change-set commits in spec order, `check-tests.py` clean | pass | pass |
| Runs of the full Validation block (log) | 4 | 2 |
| Extra break-and-rerun cycles to see a test red | 2 | 2 |
| Duration by `skill-metrics.py` | 7m 10s | 8m 55s |

Findings:

- The Validation block ran half as often, and no subagent ran it.
- Wall time did not improve on this fixture: the suite costs 20 seconds, so two saved runs are under a minute, less than the difference between two model runs. The saving grows with the cost of the block and the number of change sets; a real build is needed to measure it.
- The old run's notes named a test with a comma in it, which `check-tests.py` then split in two. The checker now splits only where a `path::name` follows the comma.

## Re-running this harness

1. Pick a baseline commit (the last commit before the change under test) and the working tree as "new".
Expand Down
2 changes: 2 additions & 0 deletions dev/evals/scope.json
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,8 @@
"The change is sized small or full, the choice is told to the user with a reason, and the user can override",
"The spec is written to .dev/{plan-name}/spec.md with a research section of D- slugged decisions using the marks (chosen, rejected, open, accepted downside, not doing)",
"The scope section ends with a Validation block listing the repo's real commands, discovered rather than guessed",
"The Validation block lists each check once, and any e2e suite, benchmark, or coverage re-run in it is marked (end of build)",
"No change set carries more than 25 scenarios, and edits to a file several change sets need are gathered into one change set rather than repeated across them",
"The spec does not implement code; no source files are modified",
"The run ends by recommending build; it is never invoked directly",
"Every change set ends with one tests: line whose scenarios carry [unit], [integration], or [e2e] tags, or tests: none with a reason",
Expand Down
Empty file added dev/evals/tests/__init__.py
Empty file.
Loading
Loading