Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-003.2.2--5aad1d64f506.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-003.2.2--5aad1d64f506",
"requirementId": "REQ-003.2.2",
"hash": "5aad1d64f506",
"verdict": "pass",
"summary": "init.ts gitignores only .2119/reviews/, verdict.ts writes plain JSON to .2119/verdicts/, and git ls-files confirms verdict files are committed and tracked",
"timestamp": "2026-07-22T18:27:54.589Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-003.4.2--020a385e7cde.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-003.4.2--020a385e7cde",
"requirementId": "REQ-003.4.2",
"hash": "020a385e7cde",
"verdict": "pass",
"summary": "README Risks section explicitly names all four required elements: agent can self-pass, committed+auditable verdicts, hash invalidation prevents stale reuse, and CI re-verification",
"timestamp": "2026-07-22T18:27:58.258Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-005.2.6--2013df807fe7.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-005.2.6--2013df807fe7",
"requirementId": "REQ-005.2.6",
"hash": "2013df807fe7",
"verdict": "pass",
"summary": "README line 73 states verbatim that verify commands execute arbitrary shell from spec files and carry the same trust level as package.json scripts",
"timestamp": "2026-07-22T18:28:05.520Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-008.1.2--b5bebba2533c.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-008.1.2--b5bebba2533c",
"requirementId": "REQ-008.1.2",
"hash": "b5bebba2533c",
"verdict": "pass",
"summary": "README has bold 'What 2119 is not:' block with all three items (test runner, CI replacement, security boundary) and a link to docs/design.md, appearing before the '## Use it in your repo' adoption section",
"timestamp": "2026-07-22T18:28:07.426Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-008.2.2--262945618870.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-008.2.2--262945618870",
"requirementId": "REQ-008.2.2",
"hash": "262945618870",
"verdict": "pass",
"summary": "Both README and scaling.md advise periodic cross-provider review --audit sweeps and targeted audits for high-consequence requirements",
"timestamp": "2026-07-22T18:27:58.700Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.1.1--0ffef5bf595d.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.1.1--0ffef5bf595d",
"requirementId": "REQ-010.1.1",
"hash": "0ffef5bf595d",
"verdict": "pass",
"summary": "All three requirement conjuncts checked: '3–8', 'workflow-level', and 'implementation-step'; removing any one fails the test; no tautology or over-mocking",
"timestamp": "2026-07-22T18:31:02.798Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.1.1--321048fefa44.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.1.1--321048fefa44",
"requirementId": "REQ-010.1.1",
"hash": "321048fefa44",
"verdict": "fail",
"summary": "Keyword theater: test checks toContain('workflow-level') but not 'implementation-step' or the preference framing; a rewrite dropping 'over implementation-step requirements' keeps both checked strings and passes while violating the stated rationale",
"timestamp": "2026-07-22T18:28:37.891Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.1.2--251ba2832873.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.1.2--251ba2832873",
"requirementId": "REQ-010.1.2",
"hash": "251ba2832873",
"verdict": "pass",
"summary": "All three required sizing smells checked with toContain that fail on removal: ~10 threshold, internal steps restatement, and quoted MUST cover/MUST test forms; string matching is appropriate for documentation-presence requirement",
"timestamp": "2026-07-22T18:31:17.129Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.1.2--251c3b25ce21.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.1.2--251c3b25ce21",
"requirementId": "REQ-010.1.2",
"hash": "251c3b25ce21",
"verdict": "fail",
"summary": "Test omits the 'MUST cover'/'MUST test' smell: removing that clause from AGENTS_MD_SECTION would pass all three assertions while violating the requirement's third explicitly-named conjunct",
"timestamp": "2026-07-22T18:28:25.353Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.2.1--85413e6bde7b.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.2.1--85413e6bde7b",
"requirementId": "REQ-010.2.1",
"hash": "85413e6bde7b",
"verdict": "pass",
"summary": "Core counterexamples pass: renamed section and SHOULD-only section both cause test failure; minor unrelated assertions (manual criteria checks) do not undermine REQ-010.2.1 coverage",
"timestamp": "2026-07-22T18:32:07.372Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.2.1--92310078f9d2.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.2.1--92310078f9d2",
"requirementId": "REQ-010.2.1",
"hash": "92310078f9d2",
"verdict": "fail",
"summary": "Test checks string presence of 'Core workflows' but never asserts that the section contains MUST requirements — the requirement's core obligation — so a Core workflows section with no MUST items would pass undetected (keyword theater)",
"timestamp": "2026-07-22T18:28:08.946Z"
}
8 changes: 8 additions & 0 deletions .2119/verdicts/REQ-010.2.2--324f15927592.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"reviewId": "REQ-010.2.2--324f15927592",
"requirementId": "REQ-010.2.2",
"hash": "324f15927592",
"verdict": "pass",
"summary": "Both violations covered: missing section fails toContain, notes before Requirements fails index ordering; SPEC_TEMPLATE is real (not mocked); ## heading level prevents sub-heading bypass",
"timestamp": "2026-07-22T18:28:24.372Z"
}
42 changes: 32 additions & 10 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,17 +5,39 @@ This repository enforces spec-driven testing with [2119](https://www.rfc-editor.

**When planning a feature**, write or update a spec in `specs/` first. Every
requirement is a numbered item under a `### REQ-NNN.M` heading with exactly one
RFC 2119 keyword. Run `npx 2119 lint` after editing specs.
RFC 2119 keyword, stating an observable outcome — not an implementation
mechanism. Run `npx rfc2119 lint` after editing specs. **Before writing tests
against a new spec**, dispatch a fresh-context reviewer to critique the draft
requirements themselves: outcome-stated, individually testable, one obligation
each. A flawed requirement steers the whole implementation wrong.

**Requirement granularity**: A first-pass feature spec should aim for around
3–8 enforced `MUST` requirements. Prefer workflow-level requirements (what the
user can observably do) over implementation-step requirements (how the code
achieves it). Spec sizing smells — reconsider the spec if: one feature produces
more than ~10 enforced requirements before tests exist; most requirements
restate internal steps rather than user-visible outcomes; a single test would
cover many requirements at once; or requirements say "MUST cover" or "MUST test"
instead of describing product behavior. Use `SHOULD` for polish and edge cases,
`[manual]` for UI-only behaviors, and notes or acceptance-checklist bullets for
implementation details rather than making every detail an enforced `MUST`.

**When implementing**, every MUST/SHALL requirement needs at least one test
annotated with a comment containing its ID, e.g. `// 2119: REQ-001.2.3`.
Write tests that would genuinely fail if the requirement were violated — a
fresh-context reviewer will judge each test's honesty, and tautological or
over-mocked tests will be rejected.
annotated with a comment containing its ID, e.g. `// 2119: REQ-001.2.3` (the
marker line must start with a comment leader). Write tests that would genuinely
fail if the requirement were violated — including its negative space: what the
requirement forbids needs a rejection test, not just what it allows. A
fresh-context reviewer judges each test's honesty; tautological or over-mocked
tests will be rejected.

**Reviewer diversity**: use reviewer models from different providers, routinely
or as periodic `npx rfc2119 review --audit` sweeps — adversarial audits of
passing verdicts. Audit especially the challenging or high-consequence
requirements; a single model family shares blind spots.

**Before finishing any task**, run `npx 2119 check`. It must exit 0. If it
reports pending judgment reviews, run `npx 2119 review` and dispatch each
instruction file in `.2119/reviews/` to a fresh-context subagent (never
review your own work in the same context). CI runs the same check, so skipping
it locally only defers the failure.
**Before finishing any task**, run `npx rfc2119 check`. It must exit 0. If it
reports pending judgment reviews, run `npx rfc2119 review --dispatch` and
dispatch each instruction file in `.2119/reviews/` to a fresh-context subagent
(never review your own work in the same context). CI runs the same check, so
skipping it locally only defers the failure.
<!-- 2119:end -->
72 changes: 72 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,78 @@ Here are some anti-patterns to avoid:
- **A keyword-grep test standing in for a judgment call** (e.g. `assert "fix" in error_message` for "errors MUST tell the user how to fix the problem") — the substring "fix" appearing proves nothing about whether the message actually explains the fix; "could not fix" passes it. The test can't fail honestly. Use `[review]` instead.
- **A review used on a machine-checkable fact** (e.g. using a reviewer to "check the version field equals 2") — that's judgment spent where a test is stronger.

## Requirement granularity

When drafting requirements for a feature, prefer a small number of user-meaningful,
workflow-level requirements over many implementation-step requirements.

A first-pass feature spec should usually contain **around 3–8 enforced `MUST`
requirements**. More granular requirements are appropriate only when each one
represents a distinct user promise, safety invariant, compatibility guarantee, or
independently valuable behavior.

Avoid turning every implementation detail into a separate `MUST`. Instead:

- Use `[manual]` for behavior that is primarily verified through UI or manual inspection.
- Use `SHOULD` for polish, edge cases, and preferred behavior that should not block the first enforcement pass.
- Use explanatory notes or acceptance-checklist bullets (in the spec's notes section) for implementation details.
- Split requirements only when each resulting requirement has a clear independent test or review value.

**Good:**

> A user MUST be able to paste an image into the composer, see it appear as an
> attachment chip, and send it as an image attachment.

**Less good as separate first-pass requirements:**

> The paste handler MUST read from the pasteboard.
> The classifier MUST create an image attachment.
> The chip area MUST become visible.
> The chip MUST show a thumbnail.
> The RPC payload MUST contain an image object.
> The transcript MUST show the image.

Those details may be useful, but they are often better expressed as notes,
manual criteria, or later hardening requirements unless each one needs independent
enforcement.

### Spec sizing smells

Reconsider the spec if:

- One feature produces more than ~10 enforced `MUST`s before tests exist.
- Most requirements are restating internal implementation steps.
- A reviewer would need to approve many requirements using the same single test.
- The requirement says "MUST cover" or "MUST test" instead of describing product behavior.
- Many requirements differ only by small UI details.

### Starter spec structure

The template that `2119 init` creates nudges toward this shape:

```markdown
### REQ-001.1: Core workflows

List 3–5 user workflows as `MUST`s.

### REQ-001.2: Safety and compatibility invariants

List only critical invariants as `MUST`s.

### REQ-001.3: Manual acceptance criteria

Use `[manual]` for visual/UI workflows that are not automated yet.

## Notes and non-goals

Put implementation details and deferred polish here instead of making them
enforced requirements.
```

This keeps 2119 adoption lightweight. Teams can start with a manageable enforced
surface, then split or harden requirements later when individual behaviors prove
important enough to test or review independently.

## Commands

| Command | Does |
Expand Down
4 changes: 0 additions & 4 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

32 changes: 32 additions & 0 deletions specs/REQ-010-granularity-guidance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# REQ-010: Requirement Granularity Guidance

## Overview

When agents draft a first-pass feature spec, 2119 should encourage workflow-level
requirements and warn against exploding a single feature into many
implementation-step `MUST`s. Without guidance, an agent can over-decompose one
feature into dozens of independently enforced requirements, making the
`cover`/`review` cycle disproportionately heavy before the team has even decided
which behaviors are core product promises versus implementation details, UI
polish, or manual acceptance criteria.

This spec governs the granularity guidance that 2119 surfaces in three places:
the generated AGENTS.md workflow section, the starter spec template, and the
README. All three should consistently nudge agents toward a small, user-meaningful
enforced surface that teams can grow incrementally.

## Requirements

### REQ-010.1: Agent instruction guidance

1. The AGENTS.md section MUST include guidance recommending around 3–8 enforced `MUST` requirements for a first-pass feature spec, with the rationale that workflow-level requirements are preferred over many implementation-step requirements.
2. The AGENTS.md section MUST include spec sizing smells — identifiable symptoms that a spec may be over-decomposed — such as a single feature producing more than ~10 enforced requirements before tests exist, requirements that restate internal implementation steps, or requirements that say "`MUST` cover" or "`MUST` test" instead of describing product behavior.

### REQ-010.2: Spec template structure

1. The starter spec template MUST include a dedicated section for core user workflows as `MUST` requirements.
2. The starter spec template MUST include a notes and non-goals section outside the `## Requirements` block, to give authors a designated place for implementation details and deferred polish rather than turning them into enforced requirements.

### REQ-010.3: README guidance

1. The README MUST include a requirement granularity section with guidance on appropriate scope for first-pass feature specs and examples of well-scoped versus over-decomposed requirements. [verify: node -e "const s = require('fs').readFileSync('README.md','utf8'); if (!s.includes('Requirement granularity')) process.exit(1)"]
38 changes: 34 additions & 4 deletions src/init.ts
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ const CONFIG_TEMPLATE = `# 2119 configuration — https://github.com/Unsupervise
# audit: "always"
`;

const SPEC_TEMPLATE = `# REQ-001: <Feature Name>
export const SPEC_TEMPLATE = `# REQ-001: <Feature Name>

## Overview

Expand All @@ -53,12 +53,31 @@ Use imperatives sparingly (RFC 2119 §6): constrain observable outcomes, not
implementation methods. Elaborate security implications of security-relevant
requirements (RFC 2119 §7).

Aim for 3–8 enforced \`MUST\` requirements in a first-pass spec. Prefer
workflow-level requirements (what the user observably can do) over
implementation-step requirements (how the code achieves it). Use \`SHOULD\` for
polish and edge cases, \`[manual]\` for UI-only behaviors, and notes or
acceptance-checklist bullets in the Notes section for implementation details.

## Requirements

### REQ-001.1: <Section Title>
### REQ-001.1: Core workflows

1. The user MUST be able to <describe the observable workflow and its outcome>.

### REQ-001.2: Safety and compatibility invariants

1. The system MUST NOT <describe what the system must never do>.

1. The system MUST <concrete, evaluable criterion>.
2. The system SHOULD <concrete, evaluable criterion>.
### REQ-001.3: Manual acceptance criteria

1. The user MUST be able to <describe visual or UI behavior that is verified by hand>. [manual]

## Notes and non-goals

Implementation details and deferred polish belong here as prose rather than
enforced requirements. Acceptance-checklist items, out-of-scope behaviors, and
internal implementation notes go here instead of as additional \`MUST\`s.
`;

export const AGENTS_MD_SECTION = `<!-- 2119:begin -->
Expand All @@ -74,6 +93,17 @@ against a new spec**, dispatch a fresh-context reviewer to critique the draft
requirements themselves: outcome-stated, individually testable, one obligation
each. A flawed requirement steers the whole implementation wrong.

**Requirement granularity**: A first-pass feature spec should aim for around
3–8 enforced \`MUST\` requirements. Prefer workflow-level requirements (what the
user can observably do) over implementation-step requirements (how the code
achieves it). Spec sizing smells — reconsider the spec if: one feature produces
more than ~10 enforced requirements before tests exist; most requirements
restate internal steps rather than user-visible outcomes; a single test would
cover many requirements at once; or requirements say "MUST cover" or "MUST test"
instead of describing product behavior. Use \`SHOULD\` for polish and edge cases,
\`[manual]\` for UI-only behaviors, and notes or acceptance-checklist bullets for
implementation details rather than making every detail an enforced \`MUST\`.

**When implementing**, every MUST/SHALL requirement needs at least one test
annotated with a comment containing its ID, e.g. \`// 2119: REQ-001.2.3\` (the
marker line must start with a comment leader). Write tests that would genuinely
Expand Down
Loading