From 7af21a5b77b32831af3ec77a49ce283b4178b8cc Mon Sep 17 00:00:00 2001 From: Eric Lee Date: Wed, 23 Sep 2026 14:06:42 -0700 Subject: [PATCH] =?UTF-8?q?release:=20v1.7.0=20=E2=80=94=20the=20multi-age?= =?UTF-8?q?nt=20system,=20rebuilt?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The news item and README highlight for the multi-agent refactor (#915, #950); the version bump across pyproject, the package fallback, install.sh, install.ps1 and uv.lock; and the changelog cut. The Unreleased block, backfilled with the PRs since v1.6.0 it had missed (nano among them), becomes 1.7.0, dated today, and the compare-link ladder gains 1.7.0. docs/NEWS.md gains the three items it had fallen behind on (v1.6.0 and both 2026-08-09 entries); the ZH README syncs to the EN top 10. The nano copy now reports both of its latest full-suite runs and drops the cost ratio until the pricing basis of the 08-16 run is checked; RUN_NANO_TB21.md records the round-3 (#900) run. The Bash pipe-drain fix that shipped with nano is listed under Fixed, since it changed every mode. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 311 +++++++++++++++++++++++++++-------- README.md | 31 +++- docs/NEWS.md | 4 + docs/i18n/README_ZH.md | 18 +- docs/nano.md | 4 +- eval/harbor/RUN_NANO_TB21.md | 13 +- install.ps1 | 2 +- install.sh | 2 +- pyproject.toml | 2 +- src/__init__.py | 2 +- uv.lock | 4 +- 11 files changed, 303 insertions(+), 90 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7d42c501b..7a3962a26 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,33 +7,108 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] -### Changed - -- **A model or effort pick is saved as your default for new sessions, on - every interface.** Choosing a model in the TUI's `/model` picker, the web - client's model chip, or the desktop's model menu — and choosing an effort - level beside it — now writes the choice to your settings, so the next - session on the CLI, the web client and the desktop all start on it, and - the surface says so: `Set model to deepseek-flash and saved as your - default for new sessions`, matching Claude Code's `/model`. Before this - the TUI persisted the model silently while labelling the picker - "persist: session", effort never persisted anywhere, and the web and - desktop pickers were deliberately session-scoped — so a pick made in one - place had to be re-made in every other. A pick from another provider - moves `default_provider` too (the persisted pair was otherwise never read - back), and the welcome-screen chip and `/api/model/info` now report the - saved choice rather than the provider's configured default. The typed - `/model --session` form keeps a switch to this session only; the - picker's `^g` global/session toggle is gone. Persisting happens only on - the host user's own transports — the TUI's stdio child and `clawcodex - serve`'s desktop/web sessions — never from a `--http` peer, which gets - `persisted: false` and a session-scoped switch. +## [1.7.0] - 2026-09-23 ### Added +- **Persistent agent teams, and background workers that really resume + (#950).** A follow-up sent to a finished background worker used to flip it + back to `running` without ever starting a model loop, and the team tools + (`TeamCreate`, `SendMessage`, the shared task board) existed with no + production path that ran a teammate. Both lifecycles are now wired end to + end: + - **Resume is real.** A follow-up launches a managed thread that reloads the + worker's typed transcript history and reuses its ID and settings. A + correction accepted while the worker was writing its final answer + continues the loop instead of going unread. + - **Persistent teammates.** `TeamCreate` establishes the leader and roster; + a named `Agent` call creates a teammate that stays available for further + assignments, keeping its context and file-read fingerprints between them. + `SendMessage` delivers findings to a named peer or to the leader, a + teammate's final prose stays private, and idle/exit notices report who is + available. + - **A shared task board.** Team members share one locked, persisted board; + automatic pickup honors dependencies (the end-to-end suite runs 24 + concurrent claimers over 12 tasks), and a completion hook can veto a + completion. + - **Plan and permission control.** Only a matching leader approval changes a + teammate's permission mode — a rejection, a stale response, or a mailbox + record the runtime never issued cannot. A worker's permission request + carries its identity and an abort signal, so interrupting the worker + denies and removes the pending prompt. + - **One supervisor for every worker.** Local, team and workflow workers + share admission, progress, transcript persistence, session- and + parent-scoped notification routing (a process-wide queue could deliver + another session's result), and a bounded shutdown when the session ends. + - **Worktree isolation is honored.** An `Agent` or workflow worker asked for + isolation runs in a real Git worktree or fails before the model runs — it + never silently edits the parent checkout. Edits and commits are preserved, + and a checkout still used by a background descendant is kept. + - **Races closed:** concurrent named launches (the task is published + before its name is claimed, and collisions are rejected), workflow + startup/stop (an immediate `TaskStop` works and a late start cannot + resurrect the task), workflow budget accounting (checked after a slot is + acquired, every attempt charged), and replay checkpoints. + + Scope, as documented in `docs/multi-agent-runtime-verification.md`: teams + are in-process — the reference's tmux/iTerm pane, remote and UDS backends + are not implemented; one team per workspace; no automatic crash recovery; + a workflow's observed token budget stops new work from starting but is not a + hard spending cap. Verified by new end-to-end suites that drive real query + loops, tools, registries, worktrees and WebSocket connections against a + scripted provider, plus one live DeepSeek smoke run (background Read, + same-ID resume, peer-to-leader delivery, approved shutdown and + `TeamDelete`) — a smoke test, not a benchmark. +- **Agent control plane — live subagent status, pause and interrupt (#915).** A + session-scoped supervisor now sees every subagent from both spawn paths + (foreground delegations previously registered nowhere, so nothing could list + or stop them). The TUI agents overlay's status readout, pause key (which + stops new spawns; running agents continue) and kill key are wired to it — + they had no backend at all and silently did nothing. The browser client + gains live agent status with per-agent interrupt and a session-wide pause on + new spawns (first as an **Agents** tab, since folded into the header's + subagent list — see Changed). An interrupted agent reports itself + as interrupted, not as a completed delegation with partial output. +- Two admission limits, both configurable (#915): + `CLAWCODEX_MAX_CONCURRENT_AGENTS` (default 32 — a runaway backstop, not a + scheduling budget) and `CLAWCODEX_MAX_AGENT_DEPTH` (default 3). A refused + spawn returns a tool error the model can act on rather than failing the + turn. +- **`clawcodex --nano` — a pi-shaped minimal harness profile (#879–#885, + #889–#891, #894, #897–#902).** Six tools (Read, Bash, Edit, Write, Grep, + Glob), a ~300-token system prompt, one-to-three-sentence tool docs, no + per-turn injections and `/eco` on: a ≈2,000-token fixed payload against the + default's ≈17,000. Skills are listed rather than tooled; Edit gains pi's + multi-edit ladder with a fuzzy near-miss match (#880); a truncation guard + refuses tool calls carried by a `max_tokens`-cut response, and compaction + summaries end with read/modified path ledgers (#881); the advisor never + activates under nano (#885); `vision_analyze` and `WebSearch` join as + conditional tools (#890, #894); long-running work gets no Bash timeout + ceiling and anti-poll guidance (#897), and an operational + constraint-checklist verification guideline lands (#898) — both held back + by #899 and relanded in #900, which adds an output-idle watchdog and pi-style + tail-keeping truncation with a full-output spill file. Runs headless, in the + TUI with a `nano` chip in the status line (#883, #884), and under `clawcodex + serve`/`web` with a badge in the browser (#901, #902). Default-off: without + the flag the harness is unchanged, apart from the Bash fix under Fixed. On + the full Terminal-Bench 2.1 suite, head to head with the pi harness + (`deepseek-v4-flash`, vision and web search on both sides, k=1 per run), + nano scored **64/89** at #896 (before the #897–#900 round) and **63/89** on + the #900 branch (measured before its final, nano-only tail-truncation + commit), against pi's **63/89**: parity on score. + The runs' recorded costs are in `eval/harbor/RUN_NANO_TB21.md`; the matched + Harbor runner and its fixes are #886, #887, #892, #893, #895. See + `docs/nano.md`. +- **Cost-aware auto-compaction (#903).** `compact.mode = "cost_aware"` (the + default stays `token_threshold`) compacts only when the estimated savings — + tokens shed, at the model's effective input and cache-read rates — repay + the summary call within `compact.break_even_turns` (default 10), falling + back to the token threshold when pricing is unknown. Compaction telemetry + measures the first post-compaction turn's cache-hit rate and net cost, and + `/context` warns when a compaction cost more than it saved. - **DeepSeek-V4.1-Flash (`deepseek-flash`), and the DeepSeek line collapses - onto it.** 1M context, 384K max output, thinking on by default — and the - first DeepSeek model that accepts an image, folding in the retired + onto it (#925).** 1M context, 384K max output, thinking on by default — and + the first DeepSeek model that accepts an image, folding in the retired `deepseek-v4-flash-vision-exp`, so a screenshot no longer needs a fusion model on this provider. It is now the `deepseek` provider's default and its whole subagent tier table (opus/sonnet/haiku), @@ -41,18 +116,24 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 performance, cost, speed, and total time" and is retiring Pro onto it. `deepseek-v4-pro`, `deepseek-v4-flash`, `deepseek-chat` and `deepseek-reasoner` all still resolve, so a pinned session keeps working. -- **`/cost` follows DeepSeek's re-card and its V4 Pro retirement.** V4.1 Flash - is cheaper than the V4 flash line it replaces — **$0.15 / $0.60 per MTok - off-peak and $0.003 cache-hit**, against $0.22 / $0.66 / $0.007 — and the - retired flash ids are billed at that same card, since DeepSeek serves them - from V4.1 Flash. `deepseek-v4-pro` keeps its own card until 2026-09-14 +- **`/cost` follows DeepSeek's re-card and its V4 Pro retirement (#925).** V4.1 + Flash is cheaper than the V4 flash line it replaces — **$0.15 / $0.60 per + MTok off-peak and $0.003 cache-hit**, against $0.22 / $0.66 / $0.007 — and + the retired flash ids are billed at that same card, since DeepSeek serves + them from V4.1 Flash. `deepseek-v4-pro` keeps its own card until 2026-09-14 12:00 Beijing (04:00 UTC) and prices as Flash from that instant, on top of - the existing peak/off-peak schedule. Both axes read the request's + the peak/off-peak schedule (see Changed). Both axes read the request's timestamp, so re-opening an August session still shows what it actually cost rather than restating it 4.4× low. - -- **Web: the reference's session stats strip.** One centred line under the - composer — `2 turns · 106 steps | LLM 6m28s · Tool call 23.7s | TTFT avg +- **ChatGPT-subscription model discovery (#913, #917).** A subscription login + now discovers the models that account can actually use instead of showing a + hardcoded list, caches them per account and token, and hides a model the + backend rejects for five minutes (#917). The subscription catalog adds + GPT-6 Astra and GPT-5.6 Sol, Terra and Luna (Astra joins the API catalog + too), and `/model` groups each fusion model under its base model's provider + rather than the active one (#913). +- **Web: the reference's session stats strip (#923).** One centred line under + the composer — `2 turns · 106 steps | LLM 6m28s · Tool call 23.7s | TTFT avg 1.3s · 258 tok/s | Cache hit 99% | Input 11.5M tok · Output 65.9K tok` — replacing the two pills and their dialogs. A group with nothing measured drops out whole. Stored assistant messages now keep each step's token @@ -60,58 +141,147 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 forwards them in the live `step.complete` shape, so a resumed session totals its cost exactly instead of hiding the figure; only TTFT and output speed, which the file cannot record, stay off the line. -- **Web: subagents in the header, and a child view per run.** A session that - delegates shows **N subagents ▾** beside its title, with a live dot while any - still run; the list behind it names each delegation with its type, model, - state, tokens and duration, and opens the run in the conversation column — - its prompt, everything it did as tool rows, and a read-only seat in place of - the composer — with the parent's title as the way back. Fed by two sources - that used to be dropped on the floor: the Agent tool's per-message progress - now reaches the browser as `subagent.progress` (the gateway translated - nothing out of `agent_progress` frames before), and the Agent tool's result - envelope (`agent_id`, status, model, duration, tokens, tool count) rides the - completion as `result.agent` and is persisted beside the stored result, so a - resumed session lists its subagents exactly as the live one did. -- Foreground subagents now keep the same sidechain transcript background ones - do (`~/.clawcodex/transcripts/.jsonl`); a new `subagent.transcript` - gateway method reads one in the stored-message shape `session.resume` uses. -- **Model-written session titles.** After the heuristic first-line name lands, - the session's own provider is asked for a short title (`generate_title` - control) and its answer replaces the heuristic — on any provider, not the - Anthropic-only path `generate_llm_title` was pinned to. An explicit rename - in the meantime wins. -- **Agent control plane — live subagent status, pause and interrupt.** A - session-scoped supervisor now sees every subagent from both spawn paths - (foreground delegations previously registered nowhere, so nothing could list - or stop them). The TUI agents overlay's status readout, pause key and kill - key are wired to it — they had no backend at all and silently did nothing. - The browser client gains an **Agents** tab listing live agents with their - model, tool count, age and status, with per-agent interrupt and a - session-wide pause. An interrupted agent reports itself as interrupted, not - as a completed delegation with partial output. -- Two admission limits, both configurable: `CLAWCODEX_MAX_CONCURRENT_AGENTS` - (default 32 — a runaway backstop, not a scheduling budget) and - `CLAWCODEX_MAX_AGENT_DEPTH` (default 3). A refused spawn returns a tool error - the model can act on rather than failing the turn. +- **Web: subagents in the header, and a child view per run (#922).** A + session that delegates shows **N subagents ▾** beside its title, with a + live dot while any still run; the list behind it names each delegation with + its type, model, state, tokens and duration, and opens the run in the + conversation column — its prompt, everything it did as tool rows, and a + read-only seat in place of the composer — with the parent's title as the + way back. Fed by two sources that used to be dropped on the floor: the + Agent tool's per-message progress now reaches the browser as + `subagent.progress` (the gateway translated nothing out of + `agent_progress` frames before), and the Agent tool's result envelope + (`agent_id`, status, model, duration, tokens, tool count) rides the + completion as `result.agent` and is persisted beside the stored result, so + a resumed session lists its subagents exactly as the live one did. + Foreground subagents now keep the same sidechain transcript background ones + do (`~/.clawcodex/transcripts/.jsonl`), and a new + `subagent.transcript` gateway method reads one in the stored-message shape + `session.resume` uses. **Model-written session titles** land in the same + change: after the heuristic first-line name, the session's own provider is + asked for a short title (`generate_title` control) — on any provider, not + the Anthropic-only path `generate_llm_title` was pinned to — and an + explicit rename in the meantime wins. +- **Web: a sidebar of tabs, with the workspace readable in it (#918).** The + right column becomes tabs: **Session** (the old panel), **Files** (the + workspace, listed a level at a time over the gateway), and a tab per opened + file, read a page at a time — a `Read` row opens the file at the line the + agent was looking at. Its two stats pills were replaced the next day by the + one-line strip above (#923). +- **Web: the reference's composer menu, sidebar start page, and whole-file + previews (#929).** The `+` button and a typed `/` open one ranked menu (an + Add section — image, plan, goal — and the commands in usage order). A `+` in + the right column's tab strip opens a Start page; files open in a viewer + that fits them — Markdown as prose, highlighted code, HTML in a sandboxed + frame, images, PDF — read through two new workspace-confined gateway calls + (`fs.read_bytes`, `fs.read_related`); and `@file` mentions in sent messages + become chips that open the file beside the conversation. +- **Web: attach files of any type (#949).** The composer's Add menu gains + **File** beside **Image**; an attached file becomes a `[File #N]` chip plus + a card under the text, and drops and pastes sort images from other files. ### Changed +- **A model or effort pick is saved as your default for new sessions, on + every interface (#930).** Choosing a model in the TUI's `/model` picker, the + web client's model chip, or the desktop's model menu — and choosing an + effort level beside it — now writes the choice to your settings, so the + next session on the CLI, the web client and the desktop all start on it, + and the surface says so: `Set model to deepseek-flash and saved as your + default for new sessions`, matching Claude Code's `/model`. Before this + the TUI persisted the model silently while labelling the picker + "persist: session", effort never persisted anywhere, and the web and + desktop pickers were deliberately session-scoped — so a pick made in one + place had to be re-made in every other. A pick from another provider + moves `default_provider` too (the persisted pair was otherwise never read + back), and the welcome-screen chip and `/api/model/info` now report the + saved choice rather than the provider's configured default. The typed + `/model --session` form keeps a switch to this session only; the + picker's `^g` global/session toggle is gone. Persisting happens only on + the host user's own transports — the TUI's stdio child and `clawcodex + serve`'s desktop/web sessions — never from a `--http` peer, which gets + `persisted: false` and a session-scoped switch. +- **Agent turn budgets raised (#906):** subagents without an explicit limit + 30 → 100 turns, `/goal` runs 20 → 100, and the query loop's `max_turns` + 50 → 200 — long multi-step runs were being cut off mid-task. +- **DeepSeek pricing gains a peak/off-peak axis (#905).** DeepSeek's current + card is published as peak/off-peak, which neither prompt size nor a + response's `service_tier` could express; `/cost` now prices by the + request's timestamp. +- The advisor's activation helper now defaults to inactive when a caller + omits the enablement flag (#907). User-facing behavior is unchanged: the + advisor was already off unless `advisor_enabled` or `/advisor` turned it on. +- **`clawcodex serve` refuses a non-loopback bind without `--allow-remote` + (#921),** as `clawcodex web` always did: its `GET /` hands out the session + token by design. `web --allow-remote` now forwards the flag to the server + it spawns. +- **`clawcodex agent-server` refuses a non-loopback bind without `--token` + (#924),** since `POST /sessions` authenticates only when a token is set. + The bearer comparison is now constant-time; `--stdio` is exempt. +- Web: one surface for subagents (#933–#935). The Agents tab is gone: Stop + moves onto each running row of the header's subagent list and into a + running child's seat, **Pause spawning** sits at the list's foot beside the + running count and cap, the list stays whole in a narrow column, and the + title gives way before the chip does. - Web: an `Agent` row reads `Agent · `, with the run's activity (running) or `N tools · duration` (done) at its right edge, and a body of - prompt, report and a button into the run — it was a generic IN/OUT card. + prompt, report and a button into the run — it was a generic IN/OUT card + (#922). - Web: a `Bash` row's summary is the model's one-line description of the command when it gave one, with the command itself in the terminal card; the - raw command line was the summary before. + raw command line was the summary before (#922). - Web: the sidebar lists a blank session only while it is the one on screen, as **New session**; every other never-used runtime session is hidden, and - project counts count what is shown. + project counts count what is shown (#922). +- Web: the right column may take up to 70% of the frame (#931); the + trajectory inspector is resizable and remembers its width (#942); only the + active session's workspace starts expanded (#946). +- The installer's completion message names the Web UI command and address + (#914). ### Fixed +- **Bash no longer stalls on a command that prints more than the ~64 KB pipe + buffer (#900).** Output was read only after the command exited, so a + command writing more than the pipe could hold blocked on its own write until + the timeout killed it, and came back reported as timed out with its output + cut at 64 KB. Output is now drained while the command runs, in every mode — + a fix that shipped with nano's stuck-command detection. +- The TUI banner showed a stale version (v1.4.0 on a v1.6.0 build); it now + shows the running backend's version from the init frame (#917). +- **Web: a saved session opens in milliseconds, not ~45 s (#947, #948).** + The transcript of a 1.8 MB session now shows in 85–150 ms instead of 48 s: a + recursive workspace walk built twice per resume was the bulk of it, and a + repeat click reuses the runtime instead of spawning a new one. The same + change adds a **New session** dialog that can start in a new workspace or a + fresh Git worktree, with **Add workspace…** pinned below the list (#948). +- Web: a reload lands back on the same session (#932); the file reader keeps + its place across a reload (#919); non-Git sessions group by their own + workspace instead of all falling under Home (#910); a late `session.clear` + reply can no longer clear the chat you navigated to (#911); column drags + over an embedded document no longer stick or re-render the app per pixel + (#936–#938). +- Web: image attachments stay visible in sent and reloaded messages (#940). + An uploaded image is retained in the session's artifact directory — it was + deleted right after attaching, so a later `Read` or `vision_analyze` failed + with "No such image file" — and a vision-capable main model gets the image + natively instead of being offered a second model to look at it (#941). +- The tool-failure-loop guard no longer conflates distinct Bash failures that + share a startup banner: its fallback category now keys off the tail of the + result, where the traceback or compiler error is (#928). +- `install.sh` stopped appending its PATH block to your shell rc on every run + (#943; the regression test skips on Windows, #945). +- TUI: a stray glyph at the tail of a row the renderer repaints is now + erased instead of sticking for the rest of the session. The fix is partial, + as its PR states: a glyph anywhere else still needs `ctrl+L` (#888). +- TUI: a collapsed paste echoes its full text in the transcript, and + `/retry` and prompts queued while the agent is busy send the paste rather + than its `[[ … ]]` label (#951). - `anthropic` is capped below 1.0. Version 1.0.0 moved the SDK onto `httpx2`, which rejects the `http_client=httpx.Client(...)` hook the provider layer depends on — the same migration `openai` is already capped for. An unpinned - resolve installed 1.x in CI while pinned local environments stayed green. + resolve installed 1.x in CI while pinned local environments stayed green + (#916, which also supplies the Harbor adapter's `_websearch` test fixture). ## [1.6.0] - 2026-08-15 @@ -989,7 +1159,8 @@ The focus was on building a solid foundation with clean architecture, comprehens --- -[Unreleased]: https://github.com/agentforce314/clawcodex/compare/v1.6.0...HEAD +[Unreleased]: https://github.com/agentforce314/clawcodex/compare/v1.7.0...HEAD +[1.7.0]: https://github.com/agentforce314/clawcodex/compare/v1.6.0...v1.7.0 [1.6.0]: https://github.com/agentforce314/clawcodex/compare/v1.5.0...v1.6.0 [1.5.0]: https://github.com/agentforce314/clawcodex/compare/v1.4.0...v1.5.0 [1.4.0]: https://github.com/agentforce314/clawcodex/compare/v1.3.0...v1.4.0 diff --git a/README.md b/README.md index 37bd65ab5..ec850f094 100644 --- a/README.md +++ b/README.md @@ -26,6 +26,24 @@
+# 🤝🦞 Multi-Agent Teams + +# Persistent teammates, and **one supervisor** for every agent + +### New in v1.7.0: `TeamCreate` makes your session the leader. Named teammates stay alive between assignments, message each other and the leader, and pull work from a shared, dependency-aware task board — while one session-scoped supervisor admits, tracks and interrupts every subagent, foreground or background, and can pause new spawns for the whole session. + +Background workers genuinely resume with their history under the same ID, a worker's pending permission +prompt is withdrawn when it is interrupted, and worktree isolation runs in a real Git checkout or fails +before the model runs. In-process teams, one per workspace — verified end to end with a scripted provider +driving real query loops, tools and WebSocket connections, plus a live DeepSeek smoke run. +**[Read the verification notes →](docs/multi-agent-runtime-verification.md)** + +
+ +*** + +
+ # 🏆 Terminal-Bench 2.1 # **80.9%** on Claude Opus 5 — a top-tier open-source result @@ -83,11 +101,13 @@ rate during peak hours (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri) — stil # The pi-style minimal harness, built in -### Full Terminal-Bench 2.1, same model, head-to-head with the [pi harness](https://pi.dev) (its own TB setup, vision+websearch matched): **nano 64/89 (71.9%) at $1.31 vs pi 63/89 (70.8%) at $2.01** — equal-or-better score, **35% cheaper**, k=1. +### Full Terminal-Bench 2.1, same model, head-to-head with the [pi harness](https://pi.dev) (its own TB setup, vision+websearch matched): **nano 64/89 and 63/89 in its two latest runs vs pi 63/89** — level on score, k=1 per run. -Six tools, a **~2K-token** fixed payload (vs ~16K default), zero per-turn injections, /eco on. +Six tools, a **~2K-token** fixed payload (vs ~17K default), zero per-turn injections, /eco on. `clawcodex --nano -p ""` ports the pi harness's edit ladder (multi-edit + fuzzy match), -truncation guard, and compaction file-ledger — while non-nano behavior stays byte-identical. +truncation guard, and compaction file-ledger — while non-nano behavior is unchanged, apart from +a Bash fix every mode got: output is now drained while a command runs, so one that prints over ~64 KB +no longer stalls until the timeout. **[docs/nano.md](docs/nano.md)**
@@ -168,6 +188,7 @@ The `session`, `settings`, and `env` blocks are optional — sensible defaults a ## 📰 News +- **2026-09-23 (v1.7.0):** **The multi-agent system, rebuilt — persistent teams, workers that really resume, and one supervisor for every agent (#915, #950)** — ClawCodex used to spawn subagents down two paths that never met, and only one of them was observable. A foreground delegation registered nowhere, so nothing could list or stop it; the TUI's agents overlay called three RPCs that had no backend; a follow-up sent to a finished background worker flipped it back to `running` without ever starting a model loop; and the team tools existed with no production path that actually ran a teammate. v1.7.0 rebuilds that layer. **One supervisor (#915, #950):** a session-scoped supervisor admits every worker — foreground, background, team and workflow — behind two backstops (`CLAWCODEX_MAX_CONCURRENT_AGENTS`, default 32; `CLAWCODEX_MAX_AGENT_DEPTH`, default 3), with live status, per-agent interrupt and a session-wide pause on new spawns in the TUI's agents overlay and the web client's subagent list (#933); a refused spawn comes back as a tool error the model can act on. **Persistent teams (#950):** `TeamCreate` makes your session the leader; each named `Agent` call adds a teammate that stays alive between assignments with its context intact; `SendMessage` routes findings to a named peer or to the leader; a shared, locked task board hands out work in dependency order; plan approvals and shutdowns follow a matched request/response protocol; and interrupting a worker withdraws its pending permission prompt. **Lifecycles that hold (#950):** resume reloads the worker's history under the same ID, a correction accepted while it was finishing is no longer left unread, notifications reach the session and parent that own them, isolation runs in a real Git worktree or fails before the model runs, and ending a session interrupts every worker it owns and waits a bounded time for them to stop. Verified by end-to-end suites that drive real query loops, tools, worktrees and WebSocket connections against a scripted provider, plus a live DeepSeek smoke run — and scoped plainly: teams are in-process (no tmux/iTerm pane or remote backends), one per workspace, a worker resumes only within the session that started it, and there is no automatic crash recovery ([verification notes](docs/multi-agent-runtime-verification.md)). **Also in v1.7.0:** `clawcodex --nano`, the pi-style minimal harness — a ≈2K-token fixed payload instead of ≈17K, and level with the pi harness on the full Terminal-Bench 2.1 suite (`deepseek-v4-flash`, k=1): 64/89 and 63/89 in nano's two latest runs, against pi's 63/89 (#879–#885, #889–#891, #894, #896–#902); a Bash fix for every mode, where a command printing more than ~64 KB stalled until the timeout and came back cut off (#900); DeepSeek-V4.1-Flash as the DeepSeek default, with `/cost` following its new peak/off-peak card (#905, #925); model and effort picks saved as your default on every interface (#930); cost-aware auto-compaction (#903); ChatGPT-subscription model discovery (#913, #917); and a round of web-client work — subagents in the header with a child view per run (#922), a sidebar of tabs that reads the workspace (#918), attachments of any file type (#949), and saved sessions that open in milliseconds instead of ~45 s (#947). **Upgrade notes:** `clawcodex serve` now refuses a non-loopback bind without `--allow-remote` (#921), and `clawcodex agent-server` refuses one without `--token` (#924). - **2026-08-15 (v1.6.0):** **ClawCodex Web — the whole agent in a browser tab (`ui-web/`, `clawcodex web`)** — one command serves the full agent at a stable, typeable address: **`http://127.0.0.1:8081`**. It is deliberately *not* a second server: the browser drives the **same in-process agent, JSON-RPC gateway, and durable session store as the TUI and the Desktop app** — start a session in one surface, resume it in another. The UI is a three-column shell (session tree | conversation | details) with streaming replies and collapsible reasoning, live tool cards (ANSI-colored terminal transcripts, unified diffs, line-numbered file reads), permission approvals as a composer takeover with once/session/always grants, a prompt queue for follow-ups typed mid-turn, slash-command completion, a context meter with per-category breakdown, a To-dos panel folded live from the transcript, and light/dark/system themes. Two views per session: **Chat** is the conversation; **Trajectory** is the same run as a metered ledger — every model request and tool call on a three-lane timeline with per-step tokens and timing (TTFT, throughput, cache-hit rate), model time and tool time reported separately. A **Settings** page covers what used to need slash commands or hand-editing `config.json`: provider **API keys** (add / replace / disconnect, stored where `clawcodex login` stores them), the **default provider** new sessions start on, the approvals default (with a Full-access confirmation that means it), response language, output style, and the end-of-turn recap toggle. `GET /` inlines the session token so the bare URL works after launch — and because of that, `clawcodex web` binds loopback only unless you explicitly pass `--allow-remote`. - **2026-08-09 (later that day):** **ClawCodex Desktop runs and packages on Windows** — `npm run dist:win:nsis` in `ui-desktop/` now produces a working unsigned NSIS installer (`ClawCodex--win-x64.exe`, crab icon and version info stamped via rcedit, node-pty **conpty** binaries staged), and `clawcodex desktop` launches the dev app natively. The packaged app boots the same `%USERPROFILE%\.clawcodex\clawcodex` backend that `install.ps1` creates — **one shared install, config, and session store across CLI, TUI, and Desktop**. Verified live on Windows 11: silent NSIS install → app boot → backend spawned from the shared venv → healthy loopback gateway → real DeepSeek turns on the shared keys. Also fixed along the way: the in-app updater's relaunch resolver pointed at the upstream `apps/desktop/` layout (never matching `ui-desktop/`), its bash relaunch handoff now honestly reports manual-restart on Windows, and a hashbang broke vitest collection of the packaging tests. - **2026-08-09:** **Native Windows support for the CLI** — ClawCodex now runs first-class on Windows 10/11 (PowerShell / cmd / Windows Terminal, no WSL required), installed with one line: `irm https://clawcodex.app/install.ps1 | iex`. The new `install.ps1` mirrors `install.sh` end to end — uv install, Python provisioning, lock-pinned deps, PATH registration, TUI build, plus the same `doctor` / `verify` / `update` / `uninstall` lifecycle. Under the hood the port is structural, not cosmetic: a shell platform layer resolves **Git Bash** for the Bash tool (explicitly refusing the WSL `System32\bash.exe` shim), so shell commands keep their POSIX semantics everywhere; process-tree kills go through `taskkill /T`; the mailbox/transcript/lockfile gain real `msvcrt` locking (the Windows CRT's `O_APPEND` emulation can silently *lose* concurrent writes); the permission layer folds NTFS case-insensitivity and refuses whole-drive grants and drive-relative escapes (deny-side only); and `@`-mentions, `CLAWCODEX.md` `@includes`, and persistent-`cd` tracking all round-trip real `C:\` paths. The full test suite now runs on `windows-latest` alongside Ubuntu in CI. @@ -177,10 +198,6 @@ The `session`, `settings`, and `env` blocks are optional — sensible defaults a - **2026-07-13:** **`/eco` token compression — -80% Bash-output tokens, measured, now a headline (#708, #712)** — a new session toggle compresses the model-bound rendering of every Bash result with deterministic filters ported from [RTK](https://github.com/rtk-ai/rtk)'s method set: failure-focused test summaries (kept error lines are never rewritten), `git`/`pip`/`npm` ceremony stripping, log dedup with `[×N]` counts, and a recoverable head-cap — all behind a **never-worse** guard, with every lossy compression teeing the full output to disk behind a runnable recovery hint (#708). A reproducible benchmark (`eval/eco/`) replays 27 real command outputs through the exact production pipeline and counts real tokenizer tokens: **92,989 → 17,767 (-80%)** corpus-wide, -88% on filter hits, plus an honestly conservative recompute of RTK's own 30-minute-session model (-19% under their averaged assumptions — real sessions are fat-tailed) (#712). Full tables: the [`/eco` section](#eco-benchmark) and [`eval/eco/results/`](eval/eco/results/results.md). - **2026-07-12 (v1.1.0):** **ClawCodex v1.1.0 — run OpenAI *and* Claude models on your subscription, not metered API billing** — the headline of 1.1.0 is **subscription auth for the two biggest model families**, so you can point ClawCodex at a plan you already pay for. **Sign in with ChatGPT (#698):** `clawcodex login → openai → subscription` (browser, device-code, or import from an existing Codex CLI login) routes requests through the ChatGPT Codex backend's Responses API — `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, and `gpt-5.3-codex-spark` on your Plus/Pro allowance, with encrypted-reasoning replay across turns and **$0** metered cost. **Claude Pro/Max (#697):** `clawcodex login → anthropic → subscription` connects a Claude subscription via OAuth (PKCE) with automatic token refresh and the same $0 accounting; follow-ups repaired the login after Anthropic moved its OAuth endpoints to `platform.claude.com` (#702) and stopped sending adaptive thinking to models that don't support it (#699). A configured API key always wins, and subscription usage reports `billing_mode: subscription`. **More models:** a Meta (`api.meta.ai`) provider with the 1M-context `muse-spark-1.1` reasoning model (#692) and refreshed MiniMax parameters (#696). **Workflow & TUI:** `/plan` mode with implicit plan-mode entry/exit (#676), `--worktree/-w` session isolation for parallel runs in separate git worktrees (#672), the `/memory` picker + `$EDITOR` spawn (#693), config/state directories rebranded `.claude → .clawcodex` with a one-time migration (#678), `/logo` startup color schemes (#677), plus TUI polish — Tab accepts the suggested placeholder (#690), past inputs get the Claude-Code highlight band (#691), clickable agent URLs (#694), and a per-terminal link-open affordance (#701). **Quality:** semantic tool-input coercion with parity validation errors (#700) and looser, Claude-Code-faithful permission granting (#673). - **2026-07-07:** **`/loop` scheduled tasks now actually fire — full port of Claude Code's session-scoped scheduler (#680)** — the bundled `/loop` skill finally has a real engine behind it: a new `src/scheduled_tasks` module parses standard 5-field cron expressions and fires due prompts **between turns** from the agent-server's idle poll. `CronCreate`/`CronList`/`CronDelete` register real firing jobs (8-char IDs, 50-job cap, deterministic jitter, 7-day recurring expiry with one final fire), and the new **`ScheduleWakeup`** tool drives self-paced `/loop` mode — the model picks each next delay (1 min–1 hr), `stop: true` ends the loop, and a ~20-minute fallback wakeup catches iterations that forget to reschedule. Typed skill slash commands now reach the backend (new `skill_command` control), so `/loop 5m check ci` works from the composer with completion + argument hint; the TUI shows a live countdown indicator (`⟳ loop wakeup in 2m 14s · ⏰ 1 scheduled`) and **Esc while idle stops a waiting loop**. `/clear` drops session tasks, `--resume` restores unexpired ones, `CLAWCODEX_DISABLE_CRON=1` disables the scheduler. 117 new tests; verified live over stdio NDJSON and a real PTY TUI drive (typed dispatch → CronCreate → a real wakeup fire between turns → Esc-stop). -- **2026-07-07:** **Bounded the ESC-cancel chunk queue in OpenAI-compatible streaming (#278)** — `OpenAICompatibleProvider.chat_stream_response`'s worker-thread queue (added in #148) was an unbounded `queue.Queue`. A non-graceful disconnect from a proxy that keeps sending bytes after ESC (and never closes the SDK iterator) let the orphaned worker thread accumulate chunks in memory indefinitely. The queue is now capped at 64 chunks, so `put()` blocks the worker once full instead of growing without bound. -- **2026-07-06 (v1.0.0):** **ClawCodex v1.0.0 — the 1.0 release: goal-directed autonomy, hooks & MCP wired for production, and a hardened permission system** — 86 commits since v0.7.0 (#580–#668) finish wiring the big subsystems end-to-end and graduate ClawCodex to 1.0. **Goal-directed autonomy:** the `/goal` + `/subgoal` completion-condition loop keeps the agent working until an LLM judge confirms the goal is actually met (#664), the new Monitor tool streams long-running shell output with backpressure (#665), background-bash completion notifications (#663), coordinator mode wired end-to-end on the live paths (#634), and `/advisor` token-efficient worker/reviewer pairing restored on the Ink TUI (#668). **Hooks live in production:** configured hooks now actually fire — bootstrap Hooks abstraction (#583), UserPromptSubmit (#597), multi-scope + lifecycle hooks (#595), `if` pre-filters (#643), PreToolUse `permissionDecision` (#655), PermissionRequest hooks at the ask seam (#637), MCP elicitation hooks (#659), and teammate TaskCompleted / TeammateIdle stop hooks (#642). **MCP completion:** OAuth server auth via the `/mcp` flow (#662), live `tools/list_changed` refresh (#598, #604), server instructions injected into the system prompt (#654), and `clawcodex mcp serve` re-exposes ClawCodex tools as an MCP stdio server (#635). **Permission hardening:** readable approval boxes with broadenable, persistent session grants (#608–#611), compound-command permission parity (#622), Bash normalization hardening (#626), `disableBypassPermissionsMode` lockdown (#660), an honest refuse-to-start unsandboxed guard (#658), subprocess secret-scrubbing (#650), and a flag-gated LLM security-classifier lane for auto mode (#589). **TUI maturity:** faithful Claude Code look & feel — diff rendering, tool-call transcript, task list, composer + permission-mode badge, busy line (#612–#616) — plus a minimal vim editing engine (#667), Esc-interrupt with a defanged Ctrl+C (#625), fully editable multi-line input (#621), slash-command argument hints (#631), a persistent session-stats line (#657), and restored `/cost`, `/skills`, and `/model` (#627, #629, #630). **Reliability:** the production compaction pipeline is wired and auto-compact actually applies its result (#587, #607), full retry lane + model fallback + message-history caching (#586), parallel Agent fan-out with the concurrency-cap deadlock fixed (#590), killing a background agent really stops the run (#606), and output styles work end-to-end (#640). Codebase stats: 1,170 Python files, **256,909 lines** (up from 233,520 lines on 2026-06-11). -- **2026-06-30 (v0.7.0):** **ClawCodex v0.7.0 — TUI auto-theming, faithful inline rendering & a Claude-Code-style tool trail** — the Ink TUI now detects your terminal's background color (OSC 11) on startup and selects the light/dark theme to match, so text stays readable on any terminal with no env var needed (#577). Inline mode renders *truly* inline like Claude Code: no screen wipe on launch, and no overlap with prior terminal output on startup or with the returning shell prompt on exit (#573, #575). The tool trail reads Claude-style — workspace-relative paths (`Read(src/foo.ts)`), `Grep(pattern)` labels, and a `Read N lines` result collapse (#574) — and the banner gains a 🦞 mascot with brighter secondary text on dark themes (#576). -- **2026-06-23:** **One-click installer** — `curl -fsSL https://clawcodex.app/install.sh | bash` installs uv (no sudo), provisions Python 3.10+, clones to `~/.clawcodex`, creates a lock-pinned venv, and registers `clawcodex` on PATH; ships status / doctor / verify / update / uninstall subcommands, is safe to re-run, and works on macOS / Linux / WSL. 📚 Older items have moved to the full **[News archive](docs/NEWS.md)**. *** diff --git a/docs/NEWS.md b/docs/NEWS.md index 5f4d9a946..9ca96c739 100644 --- a/docs/NEWS.md +++ b/docs/NEWS.md @@ -2,6 +2,10 @@ Full news history for ClawCodex. The [README News section](../README.md#-news) keeps only the 10 most recent items. +- **2026-09-23 (v1.7.0):** **The multi-agent system, rebuilt — persistent teams, workers that really resume, and one supervisor for every agent (#915, #950)** — ClawCodex used to spawn subagents down two paths that never met, and only one of them was observable. A foreground delegation registered nowhere, so nothing could list or stop it; the TUI's agents overlay called three RPCs that had no backend; a follow-up sent to a finished background worker flipped it back to `running` without ever starting a model loop; and the team tools existed with no production path that actually ran a teammate. v1.7.0 rebuilds that layer. **One supervisor (#915, #950):** a session-scoped supervisor admits every worker — foreground, background, team and workflow — behind two backstops (`CLAWCODEX_MAX_CONCURRENT_AGENTS`, default 32; `CLAWCODEX_MAX_AGENT_DEPTH`, default 3), with live status, per-agent interrupt and a session-wide pause on new spawns in the TUI's agents overlay and the web client's subagent list (#933); a refused spawn comes back as a tool error the model can act on. **Persistent teams (#950):** `TeamCreate` makes your session the leader; each named `Agent` call adds a teammate that stays alive between assignments with its context intact; `SendMessage` routes findings to a named peer or to the leader; a shared, locked task board hands out work in dependency order; plan approvals and shutdowns follow a matched request/response protocol; and interrupting a worker withdraws its pending permission prompt. **Lifecycles that hold (#950):** resume reloads the worker's history under the same ID, a correction accepted while it was finishing is no longer left unread, notifications reach the session and parent that own them, isolation runs in a real Git worktree or fails before the model runs, and ending a session interrupts every worker it owns and waits a bounded time for them to stop. Verified by end-to-end suites that drive real query loops, tools, worktrees and WebSocket connections against a scripted provider, plus a live DeepSeek smoke run — and scoped plainly: teams are in-process (no tmux/iTerm pane or remote backends), one per workspace, a worker resumes only within the session that started it, and there is no automatic crash recovery ([verification notes](multi-agent-runtime-verification.md)). **Also in v1.7.0:** `clawcodex --nano`, the pi-style minimal harness — a ≈2K-token fixed payload instead of ≈17K, and level with the pi harness on the full Terminal-Bench 2.1 suite (`deepseek-v4-flash`, k=1): 64/89 and 63/89 in nano's two latest runs, against pi's 63/89 (#879–#885, #889–#891, #894, #896–#902); a Bash fix for every mode, where a command printing more than ~64 KB stalled until the timeout and came back cut off (#900); DeepSeek-V4.1-Flash as the DeepSeek default, with `/cost` following its new peak/off-peak card (#905, #925); model and effort picks saved as your default on every interface (#930); cost-aware auto-compaction (#903); ChatGPT-subscription model discovery (#913, #917); and a round of web-client work — subagents in the header with a child view per run (#922), a sidebar of tabs that reads the workspace (#918), attachments of any file type (#949), and saved sessions that open in milliseconds instead of ~45 s (#947). **Upgrade notes:** `clawcodex serve` now refuses a non-loopback bind without `--allow-remote` (#921), and `clawcodex agent-server` refuses one without `--token` (#924). +- **2026-08-15 (v1.6.0):** **ClawCodex Web — the whole agent in a browser tab (`ui-web/`, `clawcodex web`)** — one command serves the full agent at a stable, typeable address: **`http://127.0.0.1:8081`**. It is deliberately *not* a second server: the browser drives the **same in-process agent, JSON-RPC gateway, and durable session store as the TUI and the Desktop app** — start a session in one surface, resume it in another. The UI is a three-column shell (session tree | conversation | details) with streaming replies and collapsible reasoning, live tool cards (ANSI-colored terminal transcripts, unified diffs, line-numbered file reads), permission approvals as a composer takeover with once/session/always grants, a prompt queue for follow-ups typed mid-turn, slash-command completion, a context meter with per-category breakdown, a To-dos panel folded live from the transcript, and light/dark/system themes. Two views per session: **Chat** is the conversation; **Trajectory** is the same run as a metered ledger — every model request and tool call on a three-lane timeline with per-step tokens and timing (TTFT, throughput, cache-hit rate), model time and tool time reported separately. A **Settings** page covers what used to need slash commands or hand-editing `config.json`: provider **API keys** (add / replace / disconnect, stored where `clawcodex login` stores them), the **default provider** new sessions start on, the approvals default (with a Full-access confirmation that means it), response language, output style, and the end-of-turn recap toggle. `GET /` inlines the session token so the bare URL works after launch — and because of that, `clawcodex web` binds loopback only unless you explicitly pass `--allow-remote`. +- **2026-08-09 (later that day):** **ClawCodex Desktop runs and packages on Windows** — `npm run dist:win:nsis` in `ui-desktop/` now produces a working unsigned NSIS installer (`ClawCodex--win-x64.exe`, crab icon and version info stamped via rcedit, node-pty **conpty** binaries staged), and `clawcodex desktop` launches the dev app natively. The packaged app boots the same `%USERPROFILE%\.clawcodex\clawcodex` backend that `install.ps1` creates — **one shared install, config, and session store across CLI, TUI, and Desktop**. Verified live on Windows 11: silent NSIS install → app boot → backend spawned from the shared venv → healthy loopback gateway → real DeepSeek turns on the shared keys. Also fixed along the way: the in-app updater's relaunch resolver pointed at the upstream `apps/desktop/` layout (never matching `ui-desktop/`), its bash relaunch handoff now honestly reports manual-restart on Windows, and a hashbang broke vitest collection of the packaging tests. +- **2026-08-09:** **Native Windows support for the CLI** — ClawCodex now runs first-class on Windows 10/11 (PowerShell / cmd / Windows Terminal, no WSL required), installed with one line: `irm https://clawcodex.app/install.ps1 | iex`. The new `install.ps1` mirrors `install.sh` end to end — uv install, Python provisioning, lock-pinned deps, PATH registration, TUI build, plus the same `doctor` / `verify` / `update` / `uninstall` lifecycle. Under the hood the port is structural, not cosmetic: a shell platform layer resolves **Git Bash** for the Bash tool (explicitly refusing the WSL `System32\bash.exe` shim), so shell commands keep their POSIX semantics everywhere; process-tree kills go through `taskkill /T`; the mailbox/transcript/lockfile gain real `msvcrt` locking (the Windows CRT's `O_APPEND` emulation can silently *lose* concurrent writes); the permission layer folds NTFS case-insensitivity and refuses whole-drive grants and drive-relative escapes (deny-side only); and `@`-mentions, `CLAWCODEX.md` `@includes`, and persistent-`cd` tracking all round-trip real `C:\` paths. The full test suite now runs on `windows-latest` alongside Ubuntu in CI. - **2026-08-08 (v1.5.0):** **ClawCodex Desktop — the whole agent in a native app (#802–#808)** — ClawCodex now ships a real desktop application (`ui-desktop/`): streaming chat with a live tool trail and reasoning, permission approvals with once/session/always grants, a session sidebar that lists and resumes the **same durable sessions as the TUI**, side-by-side previews, settings, and the official pixel-art crab as the dock icon and in-app brand mark. `clawcodex desktop` launches it from a checkout (first run installs the UI deps; `--no-dev` builds once and launches Electron directly). The architecture is the interesting part: the app spawns **`clawcodex serve`** — one loopback port serving `/api/*` REST plus a JSON-RPC WebSocket gateway at `/api/ws` — and sessions run on the **same in-process agent core the TUI uses**, so both surfaces share one config, one session store, one skills set, and one permission system; the wire contract is the TUI's own gateway vocabulary, adapted server-side. The port itself is one of the largest single features ClawCodex has landed: ~310K lines of TypeScript across ~1,500 files brought over from the reference desktop implementation, rebranded end to end, with every quality gate held to the reference's own baseline (typecheck green, lint identical, unit-test failure set byte-identical) and the whole loop verified live — boot → real chat turn → streamed reply rendered in the window. macOS packaging works today (`npm run dist:mac` → DMG/zip via electron-builder, hardened-runtime config in place). Ship-week fixes landed the same day: the root `.gitignore`'s Python-oriented `lib/` pattern had silently kept 180 renderer source files out of the initial import — fresh clones failed at boot until #806; the default UI scale moved from the reference's dense 90% preset to Chromium's 100% actual size (#807); and a second `clawcodex desktop` now gets a friendly "already running" message instead of a vite stack trace (#808). - **2026-08-02 (v1.4.0):** **Fusion models — give a text-only model vision (#771, #787)** — several strong reasoning models cannot see images at all: `deepseek-v4-pro` rejects an image content block outright (`400 unknown variant \`image_url\``), so pasting a screenshot, `@`-mentioning one, or letting `Read` return one ended the turn. A **fusion model** pairs that base model with a second, vision-capable one — every image is described by the vision model first, and the base model reads the description. `/fusion create ` saves one; it then behaves like a normal model in the `/model` picker, as `--model `, in `-p`, and across restarts. Ported from [claude-code-router](https://ccrdesk.top/en/configuration/fusion-models/)'s Fusion Model concept, with one deliberate difference: CCR is a proxy, so it can only offer vision as a *tool* the model may choose to call — which cannot help a pasted image, already on the wire before the model gets a turn. ClawCodex owns the agent loop, so it substitutes images in place, covering paste, `@file.png`, `Read`, and Bash image output at once. Verified end to end on Terminal-Bench 2.1's `code-from-image` task — transcribing handwritten pseudocode from a PNG and reproducing its output — with `deepseek-v4-flash` + `openai:gpt-5.6-luna` (#787); the base model alone returns a 400 on the same image. **Also in v1.4.0:** GPT-5.6 Sol/Terra/Luna (#773); four more OpenAI-compatible providers — groq, cerebras, baseten, xai — taking the registry to 30 (#784); `/mode` becomes `/permissions` with a three-level picker and Full Access by default (#768); `AskUserQuestion` finally renders a real picker instead of returning JSON to the model (#774); the OpenAI provider now picks its wire protocol from the model rather than the auth mode, which is what makes `gpt-5.6-luna` usable on an API key (#783); cached prompt tokens are billed at the cache rate instead of the full input rate, and OpenRouter's streamed reasoning is no longer discarded (#785, #786); and headless runs stop reporting a cut-short run as a success (#777–#782). - **2026-07-29 (v1.3.0):** **ClawCodex scores 80.9% on Terminal-Bench 2.1 — a top-tier open-source result on Opus 5 (#720–#725, #747–#754)** — running headless on `claude-opus-5` at `effort=xhigh`, ClawCodex solved **72 of 89** Terminal-Bench 2.1 tasks: **80.9% pass@1** on a single run. On the [public 2.1 leaderboard](https://www.tbench.ai/leaderboard/terminal-bench/2.1) (k=5 averages) that would slot **around third** — behind Claude Code / Fable 5 (83.8%) and Codex / GPT-5.5 (83.1%), statistically level with the 79–80% cluster, and **ahead of Claude Code on Opus 4.8 (78.9%) and Sonnet 5 (74.6%)**. Getting there was open, unglamorous parity work: a Harbor eval adapter (`eval/harbor/`) for three-way ClawCodex-vs-openclaude-vs-Claude-Code runs (#720, #724, #725), then a run of prompt- and reliability-parity fixes — restored task-tool skip conditions and parallel-tool guidance, deferred nonessential initial tools, and recovery of trials lost to empty turns and transport drops (#747–#754). **Also in v1.3.0:** `claude-opus-5` support with an interactive `/effort` fix (#746), bounded persistent memory with a background self-improvement review (#731), a VS Code extension driving the agent-server over stdio (#727), image-paste input with an `[Image #N]` un-attach chip (#761, #762), the `CLAUDE.md → CLAWCODEX.md` context-file rebrand (#732), and transport-retry hardening (#757, #760). Stated plainly: this is a single k=1 pass (binomial 1σ ±4.2pp) against the board's k=5 ± ~1.2pp averages, benchmarked on `main` at #756 (before the v1.3.0 tag), so read it as directional rather than a ranked submission. diff --git a/docs/i18n/README_ZH.md b/docs/i18n/README_ZH.md index 9c9b52ff6..12610b68d 100644 --- a/docs/i18n/README_ZH.md +++ b/docs/i18n/README_ZH.md @@ -64,16 +64,16 @@ clawcodex --dangerously-skip-permissions # 启动 REPL ## 📰 新闻 +- **2026-09-23(v1.7.0):** **多智能体系统重构 —— 持久化团队、真正能续跑的后台 worker,以及统管所有 agent 的监管者(#915、#950)** —— 过去 ClawCodex 通过两条互不相通的路径派生子 agent,只有其中一条可被观测。前台委派不在任何地方登记,因此无法列出或停止;TUI 的 agents 面板调用的三个 RPC 根本没有后端;向已完成的后台 worker 发送后续消息,会让它重新显示为 `running`,却从未真正启动模型循环;团队工具虽然存在,却没有任何真正运行队友的生产路径。v1.7.0 重建了这一层。**统一监管者(#915、#950):** 会话级监管者负责准入每一个 worker —— 前台、后台、团队与工作流 —— 并设有两道保护上限(`CLAWCODEX_MAX_CONCURRENT_AGENTS`,默认 32;`CLAWCODEX_MAX_AGENT_DEPTH`,默认 3),在 TUI 的 agents 面板和 Web 客户端的子 agent 列表中提供实时状态、单个 agent 中断,以及会话级的新派生暂停(#933);被拒绝的派生会以模型可以处理的工具错误返回。**持久化团队(#950):** `TeamCreate` 让当前会话成为领导者;每次具名的 `Agent` 调用都会加入一名队友,它在多次任务之间保持存活、上下文完整保留;`SendMessage` 把结论发送给指定的同伴或领导者;共享且加锁的任务板按依赖顺序分派工作;计划审批与关闭遵循请求与响应一一匹配的协议;中断某个 worker 会同时撤回它尚未处理的权限请求。**可靠的生命周期(#950):** 续跑会以同一 ID 重新加载 worker 的历史,worker 收尾期间已被接收的更正不再被搁置未读,通知会送达拥有它的会话与父 agent,隔离要么运行在真实的 Git worktree 中、要么在模型运行前就失败,结束会话时会中断其拥有的所有 worker,并在限定时间内等待它们停止。验证方式:端到端测试套件以脚本化 provider 驱动真实的查询循环、工具、worktree 与 WebSocket 连接,外加一次使用真实 DeepSeek provider 的冒烟测试 —— 适用范围也写得很清楚:团队运行在进程内(不支持 tmux/iTerm 面板或远程后端),每个工作区一个团队,worker 只能在启动它的会话内续跑,不提供崩溃后自动恢复([验证说明](../multi-agent-runtime-verification.md))。**v1.7.0 还包括:** `clawcodex --nano`,pi 风格的极简 harness —— 固定负载约 2K token(默认约 17K),在完整 Terminal-Bench 2.1 套件上与 pi harness 持平(`deepseek-v4-flash`,k=1):nano 最近两次运行分别为 64/89 与 63/89,pi 为 63/89(#879–#885、#889–#891、#894、#896–#902);一个作用于所有模式的 Bash 修复:输出超过约 64 KB 的命令过去会卡到超时,并返回被截断的输出(#900);DeepSeek-V4.1-Flash 成为 DeepSeek 默认模型,`/cost` 跟随其新的高峰/低谷价目(#905、#925);在所有界面上选择的模型与 effort 都会保存为新会话的默认值(#930);成本感知的自动压缩(#903);ChatGPT 订阅模型自动发现(#913、#917);以及一轮 Web 客户端改进 —— 标题栏中的子 agent 与每次运行的子视图(#922)、可浏览工作区的标签页侧栏(#918)、任意文件类型的附件(#949),以及已保存会话从约 45 秒缩短到毫秒级打开(#947)。**升级须知:** `clawcodex serve` 现在拒绝在没有 `--allow-remote` 的情况下绑定非回环地址(#921),`clawcodex agent-server` 在没有 `--token` 时同样拒绝(#924)。 +- **2026-08-15(v1.6.0):** **ClawCodex Web —— 在浏览器标签页中运行完整的 agent(`ui-web/`、`clawcodex web`)** —— 一条命令即可在一个稳定、便于手动输入的地址上提供完整的 agent:**`http://127.0.0.1:8081`**。它刻意*不是*第二个服务端:浏览器驱动的是**与 TUI 和桌面应用相同的进程内 agent、JSON-RPC 网关与持久化会话存储** —— 在一个界面开始的会话,可以在另一个界面续接。界面是三栏布局(会话树 | 对话 | 详情),提供流式回复与可折叠的推理过程、实时工具卡片(ANSI 着色的终端记录、统一 diff、带行号的文件读取)、以接管输入框形式呈现的权限审批(支持单次/本会话/始终授权)、可在回合中途输入后续消息的提示队列、斜杠命令补全、按类别细分的上下文用量表、从对话记录实时汇总的待办面板,以及亮色/暗色/跟随系统主题。每个会话有两个视图:**Chat** 是对话本身;**Trajectory** 把同一次运行呈现为计量账本 —— 每一次模型请求和工具调用都排布在三泳道时间线上,附带逐步的 token 与耗时(TTFT、吞吐量、缓存命中率),模型时间与工具时间分开统计。**设置**页涵盖了过去需要斜杠命令或手动编辑 `config.json` 才能完成的事:provider **API key**(添加/替换/断开,存储位置与 `clawcodex login` 相同)、新会话使用的**默认 provider**、审批默认值(选择完全访问时会给出名副其实的确认)、回复语言、输出风格,以及回合结束总结开关。`GET /` 会内联会话 token,使启动后直接访问裸 URL 即可使用 —— 也正因如此,除非显式传入 `--allow-remote`,`clawcodex web` 只绑定回环地址。 +- **2026-08-09(当天稍晚):** **ClawCodex Desktop 可在 Windows 上运行与打包** —— 在 `ui-desktop/` 中执行 `npm run dist:win:nsis` 现在能生成可用的未签名 NSIS 安装程序(`ClawCodex--win-x64.exe`,通过 rcedit 写入螃蟹图标与版本信息,并附带 node-pty 的 **conpty** 二进制文件),`clawcodex desktop` 也能原生启动开发版应用。打包后的应用会启动由 `install.ps1` 创建的同一个 `%USERPROFILE%\.clawcodex\clawcodex` 后端 —— **CLI、TUI 与桌面端共享同一套安装、配置与会话存储**。已在 Windows 11 上实机验证:静默 NSIS 安装 → 应用启动 → 从共享 venv 拉起后端 → 回环网关健康 → 使用共享密钥完成真实的 DeepSeek 对话回合。顺带修复:应用内更新器的重启解析器指向上游的 `apps/desktop/` 目录结构(永远匹配不到 `ui-desktop/`),其 bash 重启交接现在会在 Windows 上如实提示需要手动重启,以及一行 hashbang 导致 vitest 无法收集打包测试的问题。 +- **2026-08-09:** **CLI 原生支持 Windows** —— ClawCodex 现已在 Windows 10/11 上获得一等支持(PowerShell / cmd / Windows Terminal,无需 WSL),一行命令即可安装:`irm https://clawcodex.app/install.ps1 | iex`。新的 `install.ps1` 端到端对齐 `install.sh` —— 安装 uv、准备 Python、按锁文件固定依赖、注册 PATH、构建 TUI,并提供同样的 `doctor` / `verify` / `update` / `uninstall` 生命周期。底层移植是结构性的,而非表面功夫:shell 平台层为 Bash 工具解析 **Git Bash**(明确拒绝 WSL 的 `System32\bash.exe` 垫片),因此 shell 命令在各平台都保持 POSIX 语义;进程树终止改用 `taskkill /T`;mailbox/transcript/锁文件获得真正的 `msvcrt` 加锁(Windows CRT 对 `O_APPEND` 的模拟可能会悄无声息地*丢失*并发写入);权限层处理 NTFS 大小写不敏感,并拒绝整盘授权与相对盘符逃逸(仅作用于拒绝侧);`@` 引用、`CLAWCODEX.md` 的 `@includes` 与持久化 `cd` 跟踪都能正确往返真实的 `C:\` 路径。完整测试套件现已在 CI 中与 Ubuntu 一起运行于 `windows-latest`。 +- **2026-08-08(v1.5.0):** **ClawCodex Desktop —— 原生应用中的完整 agent(#802–#808)** —— ClawCodex 现已提供真正的桌面应用(`ui-desktop/`):带实时工具轨迹与推理过程的流式对话、支持单次/本会话/始终授权的权限审批、可列出并续接**与 TUI 相同的持久化会话**的会话侧栏、并排预览、设置页,以及作为程序坞图标和应用内品牌标识的官方像素风螃蟹。`clawcodex desktop` 可从源码检出直接启动它(首次运行会安装 UI 依赖;`--no-dev` 只构建一次并直接启动 Electron)。架构才是有意思的部分:应用会拉起 **`clawcodex serve`** —— 一个回环端口同时提供 `/api/*` REST 与位于 `/api/ws` 的 JSON-RPC WebSocket 网关 —— 会话运行在**与 TUI 相同的进程内 agent 核心**上,因此两个界面共享同一套配置、会话存储、技能集与权限系统;线协议沿用 TUI 自己的网关词汇,在服务端完成适配。这次移植本身是 ClawCodex 迄今落地的最大单项功能之一:从参考桌面实现迁入约 31 万行 TypeScript、约 1,500 个文件,端到端完成品牌替换,每一道质量关卡都以参考实现自身的基线为准(类型检查通过、lint 结果一致、单元测试失败集合逐字节一致),并实机验证了完整链路 —— 启动 → 真实对话回合 → 流式回复在窗口中渲染。macOS 打包现已可用(`npm run dist:mac` → 通过 electron-builder 生成 DMG/zip,已配置 hardened runtime)。发布周的修复也在同一天落地:根目录 `.gitignore` 中面向 Python 的 `lib/` 规则曾悄无声息地把 180 个渲染进程源文件挡在首次导入之外 —— 在 #806 之前,全新克隆会在启动时失败;默认 UI 缩放从参考实现偏密的 90% 预设调整为 Chromium 的 100% 实际大小(#807);再次运行 `clawcodex desktop` 时,现在会得到友好的「已在运行」提示,而不是 vite 堆栈跟踪(#808)。 - **2026-08-02(v1.4.0):** **融合模型 —— 让纯文本模型拥有视觉能力(#771、#787)** —— 一些强推理模型完全无法读图:`deepseek-v4-pro` 会直接拒绝图像内容块(`400 unknown variant \`image_url\``),因此粘贴截图、用 `@` 引用图片或让 `Read` 返回图片都会中断当前回合。**融合模型**把这样的基础模型与另一个具备视觉能力的模型配对 —— 每张图片先由视觉模型描述,基础模型读到的是描述文本。用 `/fusion create <名称> <基础模型> <视觉模型>` 保存后,它在 `/model` 选择器、`--model <名称>`、`-p` 以及重启后都表现得与普通模型一致。移植自 [claude-code-router](https://ccrdesk.top/en/configuration/fusion-models/) 的 Fusion Model 概念,但有一处刻意的差异:CCR 是代理,只能把视觉暴露成模型「可以选择调用」的工具 —— 这救不了已经在请求里的粘贴图片。ClawCodex 拥有整个 agent 循环,因此直接就地替换图像块,一次性覆盖粘贴、`@file.png`、`Read` 与 Bash 图像输出。已在 Terminal-Bench 2.1 的 `code-from-image` 任务上端到端验证 —— 从 PNG 中转写手写伪代码并复现其输出 —— 使用 `deepseek-v4-flash` + `openai:gpt-5.6-luna`(#787);同一张图片下基础模型单独运行会返回 400。**v1.4.0 其他更新:** GPT-5.6 Sol/Terra/Luna(#773);新增 groq、cerebras、baseten、xai 四个 OpenAI 兼容供应商,注册表增至 30 个(#784);`/mode` 更名为 `/permissions`,提供三档选择器并默认 Full Access(#768);`AskUserQuestion` 终于会渲染真正的选择框,而不是把 JSON 返回给模型(#774);OpenAI 供应商改为按模型而非认证方式选择传输协议,这正是 `gpt-5.6-luna` 能在 API key 下可用的原因(#783);缓存的 prompt token 按缓存价计费而非全额输入价,OpenRouter 的流式推理内容不再被丢弃(#785、#786);headless 运行不再把中途终止的回合报告为成功(#777–#782)。 - **2026-07-29(v1.3.0):** **ClawCodex 在 Terminal-Bench 2.1 上取得 80.9% —— Opus 5 上的顶尖开源成绩(#720–#725、#747–#754)** —— 在 `claude-opus-5`(`effort=xhigh`)上无头运行,ClawCodex 解决了 **89 个 Terminal-Bench 2.1 任务中的 72 个**:单次运行 **80.9% pass@1**。在[公开的 2.1 排行榜](https://www.tbench.ai/leaderboard/terminal-bench/2.1)(按 k=5 平均)上这大约排在**第 3 左右** —— 落后于 Claude Code / Fable 5(83.8%)与 Codex / GPT-5.5(83.1%),与 79–80% 的一档在统计上难分伯仲,并**领先 Claude Code 搭配 Opus 4.8(78.9%)与 Sonnet 5(74.6%)**。取得这一成绩靠的是公开而不起眼的对齐工作:一个用于 ClawCodex-vs-openclaude-vs-Claude-Code 三方对比的 Harbor 评测适配器(`eval/harbor/`,#720、#724、#725),以及一批提示词与可靠性对齐修复 —— 恢复 task 工具的跳过条件与并行工具指引、延迟加载非必要初始工具、回收因空轮次与瞬时传输中断而丢失的试次(#747–#754)。**v1.3.0 还包含:** `claude-opus-5` 支持及交互式 `/effort` 修复(#746)、带后台自我改进评审的有界持久记忆(#731)、通过 stdio 驱动 agent-server 的 VS Code 扩展(#727)、带 `[Image #N]` 取消附加芯片的图像粘贴输入(#761、#762)、`CLAUDE.md → CLAWCODEX.md` 上下文文件更名(#732),以及传输重试加固(#757、#760)。诚实说明:这是单次 k=1(二项 1σ ±4.2pp),对照的是榜单的 k=5 ± 约 1.2pp 平均值,基准运行在 v1.3.0 打标签之前的 `main`(#756)上 —— 属方向性参考,而非正式排名提交。 - **2026-07-13:** **`/eco` token 压缩 —— Bash 输出 token 实测 -80%,现为(英文版)头条之一(#708、#712)** —— 新的会话开关用一组从 [RTK](https://github.com/rtk-ai/rtk) 方法集移植的确定性过滤器压缩每个 Bash 结果的模型侧渲染:聚焦失败的测试摘要(保留的错误行从不改写)、`git`/`pip`/`npm` 仪式性输出裁剪、带 `[×N]` 计数的日志去重、可恢复的头部截断 —— 全部处于**绝不更差**守卫之下,所有有损压缩都会把完整输出 tee 到磁盘并附一条可直接运行的恢复提示(#708)。可复现的基准测试(`eval/eco/`)将 27 个真实命令输出经由生产管线逐字节重放并统计真实分词器 token:语料整体 **92,989 → 17,767(-80%)**,过滤器命中子集 -88%,另附对 RTK 自身 30 分钟会话模型的保守重算(在其平均化假设下为 -19% —— 真实会话是重尾分布)(#712)。完整表格见下文 `/eco` 章节与 [`eval/eco/results/`](../../eval/eco/results/results.md)。 - **2026-07-12(v1.1.0):** **ClawCodex v1.1.0 —— 用订阅方案运行 OpenAI 和 Claude 模型,而非按 API 计费** —— 1.1.0 的重头戏是**为两大模型家族提供订阅认证**,让你可以把 ClawCodex 接到你已经付费的方案上。**用 ChatGPT 登录(#698):** `clawcodex login → openai → subscription`(浏览器、设备码,或从已有的 Codex CLI 登录导入)通过 ChatGPT Codex 后端的 Responses API 路由请求 —— 在你的 Plus/Pro 额度内使用 `gpt-5.5`、`gpt-5.4`、`gpt-5.4-mini` 与 `gpt-5.3-codex-spark`,跨轮次重放加密推理,计费为 **$0**。**Claude Pro/Max(#697):** `clawcodex login → anthropic → subscription` 通过 OAuth(PKCE)连接 Claude 订阅,自动刷新 token,同样按 $0 计账;后续修复了 Anthropic 将 OAuth 端点迁移到 `platform.claude.com` 后的登录(#702),并停止向不支持自适应思考(adaptive thinking)的模型发送该参数(#699)。已配置的 API key 始终优先,订阅用量报告为 `billing_mode: subscription`。**更多模型:** 新增 Meta(`api.meta.ai`)provider 及 1M 上下文的 `muse-spark-1.1` 推理模型(#692),并刷新 MiniMax 参数(#696)。**工作流与 TUI:** `/plan` 模式及隐式 plan 模式进入/退出(#676)、`--worktree/-w` 会话隔离(在独立 git worktree 中并行运行,#672)、`/memory` 选择器 + `$EDITOR` 打开(#693)、配置/状态目录从 `.claude` 更名为 `.clawcodex` 并一次性迁移(#678)、`/logo` 启动配色(#677),以及 TUI 打磨 —— Tab 接受建议占位符(#690)、历史输入显示 Claude Code 高亮条(#691)、可点击的 agent URL(#694)、按终端适配的链接打开提示(#701)。**质量:** 语义化工具输入强制转换与对齐的校验错误信息(#700),以及更宽松、忠于 Claude Code 的权限授予(#673)。 - **2026-07-07:** **`/loop` 定时任务现在真正触发 —— 完整移植 Claude Code 的会话级调度器(#680)** —— 内置的 `/loop` 技能终于有了真正的引擎:新的 `src/scheduled_tasks` 模块解析标准 5 字段 cron 表达式,并在 agent-server 空闲轮询时**在轮次之间**触发到期的提示。`CronCreate`/`CronList`/`CronDelete` 注册真正触发的任务(8 字符 ID、50 个任务上限、确定性抖动、7 天循环到期并最后触发一次),新的 **`ScheduleWakeup`** 工具驱动自定节奏的 `/loop` 模式 —— 模型自行挑选每次的下一个延迟(1 分钟–1 小时),`stop: true` 结束循环,约 20 分钟的回退唤醒兜底忘记重新调度的迭代。带类型的技能斜杠命令现在可达后端(新的 `skill_command` 控制),因此 `/loop 5m check ci` 可从 composer 键入运行,带补全与参数提示;TUI 显示实时倒计时指示(`⟳ loop wakeup in 2m 14s · ⏰ 1 scheduled`),且**空闲时按 Esc 停止等待中的循环**。`/clear` 丢弃会话任务,`--resume` 恢复未到期的任务,`CLAWCODEX_DISABLE_CRON=1` 禁用调度器。117 个新测试;已在 stdio NDJSON 与真实 PTY TUI 驱动下实测验证。 -- **2026-07-07:** **为 OpenAI 兼容流式传输的 ESC 取消块队列设置上限(#278)** —— `OpenAICompatibleProvider.chat_stream_response` 的工作线程队列(#148 引入)此前是无上限的 `queue.Queue`。当代理在 ESC 后仍持续发送字节(且从不关闭 SDK 迭代器)时,被孤立的工作线程会无限累积内存中的块。队列现在上限为 64 个块,队满后 `put()` 阻塞工作线程而非无限增长。 -- **2026-07-06(v1.0.0):** **ClawCodex v1.0.0 —— 1.0 正式版:目标驱动的自主性、hooks 与 MCP 的生产级接线、强化的权限系统** —— 自 v0.7.0 以来的 86 个提交(#580–#668)完成了各大子系统的端到端接线,ClawCodex 正式升级到 1.0。**目标驱动的自主性:** `/goal` + `/subgoal` 完成条件循环让 agent 持续工作,直到 LLM 评审确认目标真正达成(#664);新的 Monitor 工具以带背压的方式流式输出长时间运行的 shell 日志(#665);后台 bash 完成通知(#663);coordinator 模式在生产路径端到端接通(#634);`/advisor` 省 token 的 worker/reviewer 搭档模式在 Ink TUI 上恢复(#668)。**Hooks 进入生产:** 配置的 hooks 现在真正生效 —— bootstrap Hooks 抽象(#583)、UserPromptSubmit(#597)、多作用域 + 生命周期 hooks(#595)、`if` 预过滤(#643)、PreToolUse `permissionDecision`(#655)、权限询问点的 PermissionRequest hooks(#637)、MCP elicitation hooks(#659),以及 teammate TaskCompleted / TeammateIdle 停止 hooks(#642)。**MCP 补全:** 通过 `/mcp` 流程完成 OAuth 服务器认证(#662)、`tools/list_changed` 实时刷新(#598、#604)、服务器说明注入系统提示词(#654),`clawcodex mcp serve` 将 ClawCodex 工具重新暴露为 MCP stdio 服务器(#635)。**权限强化:** 可读的批准框与可扩展的持久会话授权(#608–#611)、复合命令权限对齐(#622)、Bash 归一化强化(#626)、`disableBypassPermissionsMode` 锁定(#660)、诚实的拒绝启动无沙箱守卫(#658)、子进程密钥擦除(#650),以及 auto 模式下由 flag 控制的 LLM 安全分类器通道(#589)。**TUI 成熟度:** 忠实还原 Claude Code 的观感 —— diff 渲染、工具调用记录、任务列表、composer + 权限模式徽章、busy 行(#612–#616)——外加精简的 vim 编辑引擎(#667)、Esc 中断 + 去武装的 Ctrl+C(#625)、完全可编辑的多行输入(#621)、斜杠命令参数提示(#631)、常驻会话统计行(#657),以及恢复的 `/cost`、`/skills` 与 `/model`(#627、#629、#630)。**可靠性:** 生产压缩管线接通、auto-compact 真正应用其结果(#587、#607),完整重试通道 + 模型回退 + 消息历史缓存(#586),并行 Agent 扇出并修复并发上限死锁(#590),杀死后台 agent 会真正停止运行(#606),输出样式端到端可用(#640)。代码库统计:Python 文件 1,170 个,**256,909 行**(高于 2026-06-11 的 233,520 行)。 -- **2026-06-30(v0.7.0):** **ClawCodex v0.7.0 —— TUI 自动主题、忠实的内联渲染与 Claude Code 风格的工具轨迹** —— Ink TUI 启动时会探测终端背景色(OSC 11)并自动匹配明/暗主题,任何终端上文字都清晰可读、无需环境变量(#577)。内联模式像 Claude Code 一样*真正*内联渲染:启动不清屏,启动时不与之前的终端输出重叠、退出时不与返回的 shell 提示符重叠(#573、#575)。工具轨迹采用 Claude 风格 —— 工作区相对路径(`Read(src/foo.ts)`)、`Grep(pattern)` 标签与 `Read N lines` 结果折叠(#574)——横幅新增 🦞 吉祥物,暗色主题下的次要文字更亮(#576)。 -- **2026-06-24(v0.6.0):** **ClawCodex v0.6.0 —— 交互式 TUI REPL 对齐** —— 一批输入侧移植让 Python REPL 与 ink 参考实现对齐:可用的斜杠命令菜单(像 ink REPL 一样执行 / 补全 / 过滤)、带实时 token 数 + 已用时长忙碌行的星光 spinner、上下文感知的提示符底部提示(中断 / bash / 语法)、`?` 快捷键帮助面板、`@` 文件提及下拉框(原位拼接)、双击 Ctrl+C / Ctrl+D 退出、Ctrl+R 历史搜索 + 双击 Esc 清空草稿、`[Pasted text #N +K lines]` 大段粘贴占位符,以及完成的命令队列(排空排队的提示 + 暗色预览)。登录文档现在列出全部 25 个 provider(#383)。 -- **2026-06-23:** **一键安装器** —— `curl -fsSL https://clawcodex.app/install.sh | bash` 自动安装 uv(无需 sudo)、准备 Python 3.10+、克隆到 `~/.clawcodex`、创建锁定版本的 venv,并把 `clawcodex` 注册到 PATH;附带 status / doctor / verify / update / uninstall 子命令,可安全重复运行,支持 macOS / Linux / WSL。 📚 更早的条目已移至完整的 **[News 归档](../NEWS.md)**。 @@ -94,6 +94,14 @@ clawcodex --dangerously-skip-permissions # 启动 REPL *** +## 🤝 多智能体团队 —— 持久化队友,一个监管者统管所有 agent + +v1.7.0 新增:`TeamCreate` 让当前会话成为领导者。具名队友在多次任务之间保持存活,可以彼此之间以及与领导者互发消息,并从共享的、感知依赖关系的任务板领取工作;同时,一个会话级监管者负责准入、跟踪和中断每一个子 agent(无论前台还是后台),并可在整个会话范围内暂停新的派生。 + +后台 worker 会以同一 ID 带着历史真正续跑,worker 被中断时其待处理的权限请求会被撤回,worktree 隔离要么运行在真实的 Git 检出中、要么在模型运行前就失败。团队运行在进程内、每个工作区一个 —— 已通过以脚本化 provider 驱动真实查询循环、工具与 WebSocket 连接的端到端测试,以及一次使用真实 DeepSeek provider 的冒烟测试验证。详见[验证说明](../multi-agent-runtime-verification.md)。 + +*** + ## 🏆 Terminal-Bench 2.1 —— Opus 5 上 **80.9%**,顶尖开源成绩 在经过验证的 89 任务 **Terminal-Bench 2.1** 套件上,ClawCodex 以 `claude-opus-5`(`effort=xhigh`)无头运行,解决了 **72 / 89** 个任务(**80.9% pass@1**,单次运行)。在[公开排行榜](https://www.tbench.ai/leaderboard/terminal-bench/2.1)(k=5 平均)上这大约排在**第 3 左右**: diff --git a/docs/nano.md b/docs/nano.md index b523926fe..2c3673f4a 100644 --- a/docs/nano.md +++ b/docs/nano.md @@ -46,7 +46,9 @@ modes (reward 1.0)** — nano at **$0.0042 total vs $0.0186 (4.4× cheaper)**, e.g. fix-git: 54.8K vs 282.6K input tokens, $0.00165 vs $0.01015, 75 s vs 128 s. On the full 89-task suite head-to-head with the pi harness (both with vision+websearch, deepseek-v4-flash at max -thinking, k=1): **nano 64/89 ($1.31) vs pi 63/89 ($2.01)**. Full +thinking, k=1): **nano 64/89 ($1.31) vs pi 63/89 ($2.01)**; a later full +run of the next round's build (#900, before its last commit) scored 63/89 +— read the pair as parity with pi on score. Full tables: `eval/harbor/RUN_NANO_TB21.md`. ## What nano does differently diff --git a/eval/harbor/RUN_NANO_TB21.md b/eval/harbor/RUN_NANO_TB21.md index 9c2d96e64..94e3d6165 100644 --- a/eval/harbor/RUN_NANO_TB21.md +++ b/eval/harbor/RUN_NANO_TB21.md @@ -82,7 +82,18 @@ timeout-race mechanics ×3); seven v1 passes flipped red on k=1 variance (build/training-time jitter). At this k, treat 64-vs-63 as parity-to-slight-edge on score with a durable ~35% cost advantage. -Nano sends six tools and a ~2K-token fixed payload (vs ~16K default), no +Round 3 (#900 — no Bash timeout ceiling + anti-poll guidance, the +constraint-checklist guideline, `run_in_background` removed from the nano +Bash schema, stuck-command detection), `tb21-nano-flash-max-3`, same model, +vision+websearch: **63/89 (70.8%)**. The draw was 61/87 at $1.04, with two +trials lost to a pre-agent infra failure (curl SSL error fetching the uv +installer, no agent involvement); both passed when re-run, and their cost is +not in the $1.04. It was measured before #900's last commit, the nano-only +tail-keeping truncation with a full-output spill file. Across the three draws +nano's pass@1 is 60 / 64 / 63 (three different builds, the first without +vision or web search); pass@2 70/89, pass@3 73/89 across those builds. + +Nano sends six tools and a ~2K-token fixed payload (vs ~17K default), no per-turn injections, /eco on. A trivial live A/B outside Harbor (deepseek-v4-pro, write-and-verify, both solved in 4 turns) showed the same shape: 3,439 vs 27,173 fresh input tokens. diff --git a/install.ps1 b/install.ps1 index a917304fc..769a7f9d2 100644 --- a/install.ps1 +++ b/install.ps1 @@ -77,7 +77,7 @@ $ErrorActionPreference = 'Stop' # ============================================================================ # Config (defaults; env vars override like install.sh) # ============================================================================ -$INSTALLER_VERSION = '1.6.0' +$INSTALLER_VERSION = '1.7.0' # CLAWCODEX_REPO_URL override: install from a fork/mirror (or a local # checkout when testing the installer itself). $REPO_URL = if ($env:CLAWCODEX_REPO_URL) { $env:CLAWCODEX_REPO_URL } diff --git a/install.sh b/install.sh index 439b23d0f..56d1b9f95 100755 --- a/install.sh +++ b/install.sh @@ -48,7 +48,7 @@ trap 'log_err "Installer crash at line $LINENO: $BASH_COMMAND"' ERR # ============================================================================ # Config (read-only defaults) # ============================================================================ -readonly INSTALLER_VERSION="1.6.0" +readonly INSTALLER_VERSION="1.7.0" # REPO_REF is intentionally NOT readonly — it gets reassigned when the user # passes --ref. We have no version tags, so the default is the main branch; # --ref is the escape hatch for installing a specific commit/tag/branch. diff --git a/pyproject.toml b/pyproject.toml index 5958f31c3..4aed2ff4c 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "clawcodex-cli" -version = "1.6.0" +version = "1.7.0" description = "A production-oriented Python rebuild of Claude Code — real architecture, reliable CLI agent" readme = "README.md" license = "MIT" diff --git a/src/__init__.py b/src/__init__.py index 0f9fd8866..de5e16665 100644 --- a/src/__init__.py +++ b/src/__init__.py @@ -5,7 +5,7 @@ try: __version__ = version("clawcodex-cli") except PackageNotFoundError: # Running directly from an unpackaged checkout. - __version__ = "1.6.0" + __version__ = "1.7.0" __author__ = "Claw Codex Team" from .config import load_config, get_provider_config diff --git a/uv.lock b/uv.lock index 196ebbb23..82be71e1e 100644 --- a/uv.lock +++ b/uv.lock @@ -298,7 +298,7 @@ wheels = [ [[package]] name = "clawcodex-cli" -version = "1.6.0" +version = "1.7.0" source = { editable = "." } dependencies = [ { name = "anthropic" }, @@ -335,7 +335,7 @@ dev = [ [package.metadata] requires-dist = [ - { name = "anthropic", specifier = ">=0.116.0" }, + { name = "anthropic", specifier = ">=0.116.0,<1" }, { name = "build", marker = "extra == 'dev'", specifier = ">=1.0.0" }, { name = "httpx-sse", specifier = ">=0.4" }, { name = "markdownify", specifier = ">=0.11" },