diff --git a/site/harnesses/index.html b/site/harnesses/index.html index 6d6ffc2..f3cfa9f 100644 --- a/site/harnesses/index.html +++ b/site/harnesses/index.html @@ -85,7 +85,7 @@

Agent Looper Grok Bot

-

Grok Bot is the Grok operator that runs Agent Looper on the user's computer. It is not a --runtime enum value. Distinct from the Grok 4.7 model that Cursor uses as judge.

+

Grok Bot is the Grok operator that runs Agent Looper on the user's computer. It is not a --runtime enum value. Distinct from the Grok 4.7 model that Cursor uses as judge by default.

How to use

  1. Freeze GOAL.md and verify.sh (or pnpm exec agent-loop-prompt --out .cursor/loops/my-task).
  2. @@ -113,9 +113,10 @@

    Cursor

    Worker composer-2.5 Judge Grok 4.7 + grok-4.6 / grok-4.5 ok

    pnpm exec agent-loop run .cursor/loops/my-task --runtime cursor --review-gate

    -

    Export CURSOR_API_KEY or run under Doppler. Run pnpm exec agent-loop-init once. Set costPreset: "cursor" in loop.json to stay on Cursor for both worker and judge. Repo defaults that pin OpenCode escalateModel are not applied on --runtime cursor — judge uses reviewRuntime / reviewModel.

    +

    Export CURSOR_API_KEY or run under Doppler. Run pnpm exec agent-loop-init once. Default judge reviewModel is grok-4.7 when the worker is Cursor; Composer 2.5 is the only Cursor worker. Set costPreset: "cursor" in loop.json to stay on Cursor for both worker and judge. Repo defaults that pin OpenCode escalateModel are not applied on --runtime cursor — judge uses reviewRuntime / reviewModel.

    @@ -200,10 +201,11 @@

    Codex

    @openai/codex-sdk and codex CLI — ChatGPT / OpenAI BYO.

    Worker gpt-5.6-luna → gpt-5.6-terra + Optional gpt-6-astra Judge any runtime, optional

    pnpm exec agent-loop run .cursor/loops/my-task --runtime codex --review-gate

    -

    Install @openai/codex-sdk; ensure codex CLI is on PATH.

    +

    Install @openai/codex-sdk; ensure codex CLI is on PATH. Codex judge default stays gpt-5.6-sol; gpt-6-astra is listed but not the default judge.

    @@ -237,7 +239,7 @@

    Claude

    Judge any runtime, optional

    pnpm exec agent-loop run .cursor/loops/my-task --runtime claude --review-gate

    -

    Ensure claude 2.1.169+ is on PATH and claude login has been run. Typical mix: cheap worker + Claude as judge (default opus). See docs/claude-runtime.md.

    +

    Ensure claude 2.1.169+ is on PATH and claude login has been run. opus / sonnet / fable aliases track Claude Code latest-per-family. Typical mix: cheap worker + Claude as judge (default opus). See docs/claude-runtime.md.

diff --git a/site/harnesses/index.md b/site/harnesses/index.md index 69661dd..f145d12 100644 --- a/site/harnesses/index.md +++ b/site/harnesses/index.md @@ -10,7 +10,7 @@ You say what to build and how to know it's done. It keeps a coding agent working [Add to Grok Bot](https://x.ai/bot/AETdGbRRNWfckrRGv22LD) -Grok Bot is the Grok operator that runs Agent Looper on the user's computer. It is not a `--runtime` enum value. Distinct from the Grok 4.7 model that Cursor uses as judge. +Grok Bot is the Grok operator that runs Agent Looper on the user's computer. It is not a `--runtime` enum value. Distinct from the Grok 4.7 model that Cursor uses as judge by default. ### How to use @@ -29,7 +29,7 @@ Install once: `pnpm add -D @dancingteeth/agent-looper`, then add the SDK or CLI IDE subscription via `@cursor/sdk`. - Worker: `composer-2.5` -- Judge (when worker is Cursor): Grok 4.7 +- Judge (when worker is Cursor): `grok-4.7` (`grok-4.6` / `grok-4.5` still allowed) - Run: `pnpm exec agent-loop run .cursor/loops/my-task --runtime cursor --review-gate` Export `CURSOR_API_KEY` or run under Doppler. Run `pnpm exec agent-loop-init` once. Set `costPreset: "cursor"` in `loop.json` to stay on Cursor for both worker and judge. Repo defaults that pin OpenCode `escalateModel` are not applied on `--runtime cursor` — judge uses `reviewRuntime` / `reviewModel`. @@ -48,7 +48,7 @@ DeepSeek Harness CLI — worker is `dsh --profile headless`; `dsh-agent-looper` `@cline/sdk` — `cline-pass` for subscription quota, `cline` for credits. -- `cline-pass` worker: `cline-pass/deepseek-v4.1-flash` → `qwen3.7-plus` +- `cline-pass` worker: `cline-pass/deepseek-v4.1-flash` → `qwen3.7-plus` (`deepseek-v4-flash` still allowed) - `cline` worker: `deepseek/deepseek-chat` → `qwen/qwen3-coder-plus` - Judge: any runtime, optional - Run: `pnpm exec agent-loop run .cursor/loops/my-task --runtime cline-pass --review-gate` (or `--runtime cline` for credits) @@ -57,7 +57,7 @@ DeepSeek Harness CLI — worker is `dsh --profile headless`; `dsh-agent-looper` `@opencode-ai/sdk` and `opencode` CLI — Go quota by default, or BYOK through OpenRouter, Vercel AI Gateway, Ollama, or another OpenAI-compatible router. -- Go worker: `opencode-go/deepseek-v4.1-flash` → `qwen3.7-plus` +- Go worker: `opencode-go/deepseek-v4.1-flash` → `qwen3.7-plus` (`deepseek-v4-flash` still allowed) - BYOK: `openrouter/…`, `openrouter/…:free`, `vercel/…`, `ollama/…` - Judge: any runtime, optional - Run: `pnpm exec agent-loop run .cursor/loops/my-task --runtime opencode --review-gate` @@ -75,7 +75,8 @@ DeepSeek Harness CLI — worker is `dsh --profile headless`; `dsh-agent-looper` `@openai/codex-sdk` and `codex` CLI — ChatGPT / OpenAI BYO. - Worker: `gpt-5.6-luna` → `gpt-5.6-terra` -- Judge: any runtime, optional +- Optional: `gpt-6-astra` (listed; not the default judge) +- Judge: any runtime, optional (Codex-native judge defaults to `gpt-5.6-sol`, not `gpt-6-astra`) - Run: `pnpm exec agent-loop run .cursor/loops/my-task --runtime codex --review-gate` ### Muse (`--runtime muse`) @@ -88,7 +89,7 @@ DeepSeek Harness CLI — worker is `dsh --profile headless`; `dsh-agent-looper` ### Claude (`--runtime claude`) -PATH `claude` CLI — Claude Code subscription. `--safe-mode` so the harness prompt is the only instruction source (strips project hooks and auto-memory). Not on `costPreset` minmax. See [docs/claude-runtime.md](https://github.com/dancingteeth/agent-looper/blob/main/docs/claude-runtime.md). +PATH `claude` CLI — Claude Code subscription. `--safe-mode` so the harness prompt is the only instruction source (strips project hooks and auto-memory). Not on `costPreset` minmax. `opus` / `sonnet` / `fable` aliases track Claude Code latest-per-family. See [docs/claude-runtime.md](https://github.com/dancingteeth/agent-looper/blob/main/docs/claude-runtime.md). - Worker: `sonnet` → `opus` - Judge: any runtime, optional diff --git a/site/index.html b/site/index.html index 475da7f..f046944 100644 --- a/site/index.html +++ b/site/index.html @@ -107,7 +107,7 @@ "name": "How do Agent Looper worker and judge presets work?", "acceptedAnswer": { "@type": "Answer", - "text": "Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On runtime: cursor, repo defaults that pin OpenCode Go escalateModel are not applied — Composer plus another judge uses reviewRuntime / reviewModel, not escalateModel." + "text": "Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok — default judge grok-4.7). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On runtime: cursor, repo defaults that pin OpenCode Go escalateModel are not applied — Composer plus another judge uses reviewRuntime / reviewModel, not escalateModel." } }, { @@ -171,7 +171,15 @@ "name": "What do the spend numbers mean?", "acceptedAnswer": { "@type": "Answer", - "text": "Watch and the report card show two numbers when they differ: list (public API rates, including prompt-cache) and billed (what the runtime invoice says). $0 on a subscription quota is billed $0, not free. Budget caps use billed when you are on PAYG and list when the invoice is $0. OpenCode Go and Cline Pass model lists and list prices come from models.dev; unknown but well-formed slugs soft-gate — warn at loop start, show $0 until priced. Budget caps still reject unpriced worker or escalate models." + "text": "Watch and the report card show two numbers when they differ: list (public API rates, including prompt-cache) and billed (what the runtime invoice says). $0 on a subscription quota is billed $0, not free. Budget caps use billed when you are on PAYG and list when the invoice is $0. OpenCode Go and Cline Pass model lists and list prices come from models.dev — sync picks up Grok 4.7, GPT-6 Luna, MiMo V2.6 Flash/Pro, and Space Bunny Free; unknown but well-formed slugs soft-gate — warn at loop start, show $0 until priced. Budget caps still reject unpriced worker or escalate models." + } + }, + { + "@type": "Question", + "name": "How do I harden verify between runs?", + "acceptedAnswer": { + "@type": "Answer", + "text": "Before you add another line to verify.sh or REVIEWS.md, walk a diverse sample of past runs under .cursor/loop-exports/ so the scoreboard matches real failures — not one lucky green. Procedure: docs/unknowns-preflight.md on the repo. Freeze during a run is unchanged." } } ] @@ -342,7 +350,7 @@

Which coding agents does Agent Looper work with?

How do Agent Looper worker and judge presets work?

- Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On runtime: cursor, repo defaults that pin OpenCode Go escalateModel are not applied — Composer plus another judge uses reviewRuntime / reviewModel, not escalateModel. + Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok — default judge grok-4.7). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On runtime: cursor, repo defaults that pin OpenCode Go escalateModel are not applied — Composer plus another judge uses reviewRuntime / reviewModel, not escalateModel.

@@ -395,7 +403,13 @@

What if verify fails because my environment is broken?

What do the spend numbers mean?

- Watch and the report card show two numbers when they differ: list (public API rates, including prompt-cache) and billed (what the runtime invoice says). $0 on a subscription quota is billed $0, not “free.” Budget caps use billed when you are on PAYG and list when the invoice is $0. OpenCode Go and Cline Pass model lists and list prices come from models.dev; unknown but well-formed slugs soft-gate — warn at loop start, show $0 until priced. Budget caps still reject unpriced worker or escalate models. + Watch and the report card show two numbers when they differ: list (public API rates, including prompt-cache) and billed (what the runtime invoice says). $0 on a subscription quota is billed $0, not “free.” Budget caps use billed when you are on PAYG and list when the invoice is $0. OpenCode Go and Cline Pass model lists and list prices come from models.dev — sync picks up Grok 4.7, GPT-6 Luna, MiMo V2.6 Flash/Pro, and Space Bunny Free; unknown but well-formed slugs soft-gate — warn at loop start, show $0 until priced. Budget caps still reject unpriced worker or escalate models. +

+
+
+

How do I harden verify between runs?

+

+ Before you add another line to verify.sh or REVIEWS.md, walk a diverse sample of past runs under .cursor/loop-exports/ so the scoreboard matches real failures — not one lucky green. Procedure: docs/unknowns-preflight.md. Freeze during a run is unchanged.

@@ -410,6 +424,7 @@

How it works

Optional setup.sh (or setup in loop.json) runs once before the first worker — setup failure does not spawn a worker. A fresh worker loops until that check passes — optional judge / reviewGate only on serious findings. After each visit the harness restores frozen specs if a worker edited them — and removes any frozen basename they planted (like setup.sh) so it cannot carry into the next run. + Between runs, label export packs before you tighten the check (unknowns preflight). Progress lives in git and files, not chat memory.

diff --git a/site/index.md b/site/index.md index 6df0af4..f90992b 100644 --- a/site/index.md +++ b/site/index.md @@ -12,7 +12,7 @@ Agent Looper uses the coding agents you already pay for: Cursor, Cline, OpenCode ## How do Agent Looper worker and judge presets work? -Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On `runtime: cursor`, repo defaults that pin OpenCode Go `escalateModel` are not applied — Composer plus another judge uses `reviewRuntime` / `reviewModel`, not `escalateModel`. +Named presets: minmax (efficiency — cheap capable worker, strongest included judge), balanced (spend more on the worker, same judge), cursor (stay on Cursor: Composer + Grok — default judge `grok-4.7`). Or you, your agent, or Looper wires a pair from what's installed. The pair stays for the whole loop. Not Auto. On `runtime: cursor`, repo defaults that pin OpenCode Go `escalateModel` are not applied — Composer plus another judge uses `reviewRuntime` / `reviewModel`, not `escalateModel`. ## How does Agent Looper keep cost down for indie builders? @@ -50,11 +50,15 @@ When the check fails because something on your machine is missing or broken — ## What do the spend numbers mean? -Watch and the report card show two numbers when they differ: **list** (public API rates, including prompt-cache) and **billed** (what the runtime invoice says). `$0` on a subscription quota is billed `$0`, not “free.” Budget caps use billed when you are on PAYG and list when the invoice is `$0`. OpenCode Go and Cline Pass model lists and list prices come from [models.dev](https://models.dev/api.json); unknown but well-formed slugs soft-gate — warn at loop start, show `$0` until priced. Budget caps still reject unpriced worker or escalate models. +Watch and the report card show two numbers when they differ: **list** (public API rates, including prompt-cache) and **billed** (what the runtime invoice says). `$0` on a subscription quota is billed `$0`, not “free.” Budget caps use billed when you are on PAYG and list when the invoice is `$0`. OpenCode Go and Cline Pass model lists and list prices come from [models.dev](https://models.dev/api.json) — sync picks up Grok 4.7, GPT-6 Luna, MiMo V2.6 Flash/Pro, and Space Bunny Free; unknown but well-formed slugs soft-gate — warn at loop start, show `$0` until priced. Budget caps still reject unpriced worker or escalate models. + +## How do I harden verify between runs? + +Before you add another line to `verify.sh` or `REVIEWS.md`, walk a diverse sample of past runs under `.cursor/loop-exports/` so the scoreboard matches real failures — not one lucky green. Procedure: [docs/unknowns-preflight.md](https://github.com/dancingteeth/agent-looper/blob/main/docs/unknowns-preflight.md). Freeze during a run is unchanged. ## How it works -Your agent writes `GOAL.md` and a deterministic check. Optional `setup.sh` (or `setup` in `loop.json`) runs once before the first worker — setup failure does not spawn a worker. A fresh worker loops until the check passes. After each visit the harness restores frozen specs if a worker edited them — and removes any frozen basename they planted (like `setup.sh`) so it cannot carry into the next run. +Your agent writes `GOAL.md` and a deterministic check. Optional `setup.sh` (or `setup` in `loop.json`) runs once before the first worker — setup failure does not spawn a worker. A fresh worker loops until the check passes. After each visit the harness restores frozen specs if a worker edited them — and removes any frozen basename they planted (like `setup.sh`) so it cannot carry into the next run. Between runs, label export packs before you tighten the check ([unknowns preflight](https://github.com/dancingteeth/agent-looper/blob/main/docs/unknowns-preflight.md)). ## How is Agent Looper different from looping in chat? diff --git a/site/llms.txt b/site/llms.txt index 19bf993..c42f23e 100644 --- a/site/llms.txt +++ b/site/llms.txt @@ -6,7 +6,7 @@ When to use Agent Looper: a human wants a feature finished without babysitting C When not to use Agent Looper: there is no deterministic check (pure taste with no script or vision gate); you need a team-ops factory, shared org cloud, or multi-repo orchestration; you want the model to grade itself as “done.” -How an agent should call it: install `@dancingteeth/agent-looper`, ask the human what they want, then implement the feature with Agent Looper (set up the loop, write GOAL.md and verify.sh, run until green). Or from the terminal: `pnpm exec agent-loop-setup` (one-screen setup wizard — recap plus current question) then `pnpm exec agent-loop-prompt --out .cursor/loops/` — the scaffold judge stops as soon as frozen GOAL.md and verify.sh pass freeze (10-minute wall, not a long grind) — then `pnpm exec agent-loop run .cursor/loops/ --runtime `. On `runtime: cursor`, repo defaults that pin OpenCode Go `escalateModel` are not applied; Composer plus another judge uses `reviewRuntime` / `reviewModel`, not `escalateModel`. DSH headless: `~/.dsh/.credentials.yaml` must be a flat `KEY: "string"` map — a wrapped store (`version` / `refs` / `records`) or YAML integer fails `agent-check dsh` and `agent-loop run` before spawning the worker; quoting `version` alone is not enough (the harness does not rewrite that file). Optional `--review-gate` re-opens the loop only on blocking review findings. If verify fails because the environment is broken, the loop waits instead of burning another worker. After each visit the harness restores frozen specs and removes any frozen basename a worker planted (e.g. `setup.sh`) so it cannot persist. OpenCode Go and Cline Pass model lists and list prices come from models.dev; unknown well-formed slugs warn at start and read `$0` until priced — budget caps still reject unpriced worker or escalate models. CLI binaries: `agent-loop`, `agent-loop-prompt`. Current npm: **0.7.0** (supported line **0.7.x**). +How an agent should call it: install `@dancingteeth/agent-looper`, ask the human what they want, then implement the feature with Agent Looper (set up the loop, write GOAL.md and verify.sh, run until green). Or from the terminal: `pnpm exec agent-loop-setup` (one-screen setup wizard — recap plus current question) then `pnpm exec agent-loop-prompt --out .cursor/loops/` — the scaffold judge stops as soon as frozen GOAL.md and verify.sh pass freeze (10-minute wall, not a long grind) — then `pnpm exec agent-loop run .cursor/loops/ --runtime `. On `runtime: cursor`, repo defaults that pin OpenCode Go `escalateModel` are not applied; Composer worker plus judge uses `reviewRuntime` / `reviewModel` (default judge `grok-4.7`; `grok-4.6` / `grok-4.5` still allowed), not `escalateModel`. DeepSeek Flash defaults are 4.1: DSH `deepseek-official/deepseek-flash`, OpenCode Go `opencode-go/deepseek-v4.1-flash`, Cline Pass `cline-pass/deepseek-v4.1-flash`; older `deepseek-v4-flash` slugs still parse. DSH headless: `~/.dsh/.credentials.yaml` must be a flat `KEY: "string"` map — a wrapped store (`version` / `refs` / `records`) or YAML integer fails `agent-check dsh` and `agent-loop run` before spawning the worker; quoting `version` alone is not enough (the harness does not rewrite that file). Optional `--review-gate` re-opens the loop only on blocking review findings. If verify fails because the environment is broken, the loop waits instead of burning another worker. After each visit the harness restores frozen specs and removes any frozen basename a worker planted (e.g. `setup.sh`) so it cannot persist. Before tightening `verify.sh` or `REVIEWS.md` between runs, label a diverse sample of `.cursor/loop-exports/` (see docs/unknowns-preflight.md). OpenCode Go and Cline Pass model lists and list prices come from models.dev (sync picks up Grok 4.7, GPT-6 Luna, MiMo V2.6 Flash/Pro, Space Bunny Free); unknown well-formed slugs warn at start and read `$0` until priced — budget caps still reject unpriced worker or escalate models. CLI binaries: `agent-loop`, `agent-loop-prompt`. Current npm: **0.7.0** (supported line **0.7.x**). ## Developer resources diff --git a/src/site/landingAgentReadiness.test.ts b/src/site/landingAgentReadiness.test.ts index aab6832..93723b5 100644 --- a/src/site/landingAgentReadiness.test.ts +++ b/src/site/landingAgentReadiness.test.ts @@ -652,7 +652,7 @@ describe('landing agent readiness', () => { } }) - it('names 0.6.3 setup wizard, env-wait, harness setup, frozen restore, plant block, models.dev soft-gate, Cursor escalate, scaffold wall, DSH flat credentials, and DSH 4.1 Flash opt-in', () => { + it('names 0.7.0 setup wizard, env-wait, harness setup, frozen restore, plant block, models.dev soft-gate, Cursor Grok 4.7, Flash 4.1 defaults, loop-exports preflight, Cursor escalate, scaffold wall, DSH flat credentials, Codex astra, and Claude CLI aliases', () => { const html = readSite('index.html') const md = readSite('index.md') const llms = readSite('llms.txt') @@ -686,8 +686,13 @@ describe('landing agent readiness', () => { const plantBeat = 'removes any frozen basename' const modelsDevBeat = 'models.dev' const softGateBeat = 'soft-gate' - const flashOptIn = 'deepseek-flash' - const flashLabel = '4.1 Flash' + const flashDefault = 'deepseek-flash' + const flash41Slug = 'deepseek-v4.1-flash' + const grok47Beat = 'grok-4.7' + const loopExportsBeat = '.cursor/loop-exports/' + const unknownsPreflightBeat = 'unknowns-preflight.md' + const codexAstraBeat = 'gpt-6-astra' + const claudeAliasBeat = 'latest-per-family' const cursorEscalateBeat = 'are not applied' const scaffoldWallBeat = '10-minute wall' const dshFlatBeat = 'must be a flat' @@ -712,17 +717,23 @@ describe('landing agent readiness', () => { expect(surface).not.toMatch(/Current npm.*0\.6\.1/i) expect(surface).not.toMatch(/Current npm.*0\.6\.2/i) expect(surface).not.toMatch(/Current npm.*0\.6\.3/i) + expect(surface).not.toMatch(/supported line \*\*0\.6\.x\*\*/i) expect(surface).toContain(cursorEscalateBeat) expect(surface).toContain(scaffoldWallBeat) + expect(surface).toContain(loopExportsBeat) + expect(surface).toContain(unknownsPreflightBeat) } expect(llms).toContain('Current npm: **0.7.0**') expect(llms).toContain('0.7.x') expect(llms).not.toContain('0.5.0') expect(llms).not.toContain('Current npm: **0.6.3**') + expect(llms).not.toMatch(/supported line \*\*0\.6\.x\*\*/) expect(llms).not.toContain('Current npm: **0.6.2**') expect(llms).not.toContain('Current npm: **0.6.1**') expect(llms).not.toContain('Current npm: **0.6.0**') + expect(llms).toContain(grok47Beat) + expect(llms).toContain(flash41Slug) expect(llms).toContain(dshPreflightBeat) expect(llms).toContain(dshQuoteNotEnoughBeat) expect(llms).toContain(setupWizardBeat) @@ -746,6 +757,20 @@ describe('landing agent readiness', () => { )?.[0] ?? '' expect(cursorCard).toContain('reviewRuntime') expect(cursorCard).toContain('escalateModel') + expect(cursorCard).toContain(grok47Beat) + + const codexCard = + harnessHtml.match( + /
[\s\S]*?<\/article>/, + )?.[0] ?? '' + expect(codexCard).toContain(codexAstraBeat) + expect(codexCard).toContain('gpt-5.6-sol') + + const claudeCard = + harnessHtml.match( + /
[\s\S]*?<\/article>/, + )?.[0] ?? '' + expect(claudeCard).toContain(claudeAliasBeat) const dshCard = harnessHtml.match( @@ -754,11 +779,14 @@ describe('landing agent readiness', () => { expect(dshCard).toContain(dshFlatBeat) expect(dshCard).toContain(dshPreflightBeat) expect(dshCard).not.toContain('version: "1"') - expect(dshCard).toContain(flashOptIn) - expect(dshCard).toContain(flashLabel) + expect(dshCard).toContain(flashDefault) expect(dshCard).toContain('deepseek-v4-flash') - expect(harnessMd).toContain(flashOptIn) + expect(harnessMd).toContain(flashDefault) expect(harnessMd).toContain('deepseek-v4-flash') + expect(harnessMd).toContain(flash41Slug) + expect(harnessMd).toContain(grok47Beat) + expect(harnessMd).toContain(codexAstraBeat) + expect(harnessMd).toContain(claudeAliasBeat) expect(harnessMd).toContain(dshFlatBeat) expect(harnessMd).toContain(dshPreflightBeat) expect(harnessMd).toContain(dshQuoteNotEnoughBeat)