Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
311 changes: 241 additions & 70 deletions CHANGELOG.md

Large diffs are not rendered by default.

31 changes: 24 additions & 7 deletions README.md

Large diffs are not rendered by default.

4 changes: 4 additions & 0 deletions docs/NEWS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@

Full news history for ClawCodex. The [README News section](../README.md#-news) keeps only the 10 most recent items.

- **2026-09-23 (v1.7.0):** **The multi-agent system, rebuilt — persistent teams, workers that really resume, and one supervisor for every agent (#915, #950)** — ClawCodex used to spawn subagents down two paths that never met, and only one of them was observable. A foreground delegation registered nowhere, so nothing could list or stop it; the TUI's agents overlay called three RPCs that had no backend; a follow-up sent to a finished background worker flipped it back to `running` without ever starting a model loop; and the team tools existed with no production path that actually ran a teammate. v1.7.0 rebuilds that layer. **One supervisor (#915, #950):** a session-scoped supervisor admits every worker — foreground, background, team and workflow — behind two backstops (`CLAWCODEX_MAX_CONCURRENT_AGENTS`, default 32; `CLAWCODEX_MAX_AGENT_DEPTH`, default 3), with live status, per-agent interrupt and a session-wide pause on new spawns in the TUI's agents overlay and the web client's subagent list (#933); a refused spawn comes back as a tool error the model can act on. **Persistent teams (#950):** `TeamCreate` makes your session the leader; each named `Agent` call adds a teammate that stays alive between assignments with its context intact; `SendMessage` routes findings to a named peer or to the leader; a shared, locked task board hands out work in dependency order; plan approvals and shutdowns follow a matched request/response protocol; and interrupting a worker withdraws its pending permission prompt. **Lifecycles that hold (#950):** resume reloads the worker's history under the same ID, a correction accepted while it was finishing is no longer left unread, notifications reach the session and parent that own them, isolation runs in a real Git worktree or fails before the model runs, and ending a session interrupts every worker it owns and waits a bounded time for them to stop. Verified by end-to-end suites that drive real query loops, tools, worktrees and WebSocket connections against a scripted provider, plus a live DeepSeek smoke run — and scoped plainly: teams are in-process (no tmux/iTerm pane or remote backends), one per workspace, a worker resumes only within the session that started it, and there is no automatic crash recovery ([verification notes](multi-agent-runtime-verification.md)). **Also in v1.7.0:** `clawcodex --nano`, the pi-style minimal harness — a ≈2K-token fixed payload instead of ≈17K, and level with the pi harness on the full Terminal-Bench 2.1 suite (`deepseek-v4-flash`, k=1): 64/89 and 63/89 in nano's two latest runs, against pi's 63/89 (#879–#885, #889–#891, #894, #896–#902); a Bash fix for every mode, where a command printing more than ~64 KB stalled until the timeout and came back cut off (#900); DeepSeek-V4.1-Flash as the DeepSeek default, with `/cost` following its new peak/off-peak card (#905, #925); model and effort picks saved as your default on every interface (#930); cost-aware auto-compaction (#903); ChatGPT-subscription model discovery (#913, #917); and a round of web-client work — subagents in the header with a child view per run (#922), a sidebar of tabs that reads the workspace (#918), attachments of any file type (#949), and saved sessions that open in milliseconds instead of ~45 s (#947). **Upgrade notes:** `clawcodex serve` now refuses a non-loopback bind without `--allow-remote` (#921), and `clawcodex agent-server` refuses one without `--token` (#924).
- **2026-08-15 (v1.6.0):** **ClawCodex Web — the whole agent in a browser tab (`ui-web/`, `clawcodex web`)** — one command serves the full agent at a stable, typeable address: **`http://127.0.0.1:8081`**. It is deliberately *not* a second server: the browser drives the **same in-process agent, JSON-RPC gateway, and durable session store as the TUI and the Desktop app** — start a session in one surface, resume it in another. The UI is a three-column shell (session tree | conversation | details) with streaming replies and collapsible reasoning, live tool cards (ANSI-colored terminal transcripts, unified diffs, line-numbered file reads), permission approvals as a composer takeover with once/session/always grants, a prompt queue for follow-ups typed mid-turn, slash-command completion, a context meter with per-category breakdown, a To-dos panel folded live from the transcript, and light/dark/system themes. Two views per session: **Chat** is the conversation; **Trajectory** is the same run as a metered ledger — every model request and tool call on a three-lane timeline with per-step tokens and timing (TTFT, throughput, cache-hit rate), model time and tool time reported separately. A **Settings** page covers what used to need slash commands or hand-editing `config.json`: provider **API keys** (add / replace / disconnect, stored where `clawcodex login` stores them), the **default provider** new sessions start on, the approvals default (with a Full-access confirmation that means it), response language, output style, and the end-of-turn recap toggle. `GET /` inlines the session token so the bare URL works after launch — and because of that, `clawcodex web` binds loopback only unless you explicitly pass `--allow-remote`.
- **2026-08-09 (later that day):** **ClawCodex Desktop runs and packages on Windows** — `npm run dist:win:nsis` in `ui-desktop/` now produces a working unsigned NSIS installer (`ClawCodex-<version>-win-x64.exe`, crab icon and version info stamped via rcedit, node-pty **conpty** binaries staged), and `clawcodex desktop` launches the dev app natively. The packaged app boots the same `%USERPROFILE%\.clawcodex\clawcodex` backend that `install.ps1` creates — **one shared install, config, and session store across CLI, TUI, and Desktop**. Verified live on Windows 11: silent NSIS install → app boot → backend spawned from the shared venv → healthy loopback gateway → real DeepSeek turns on the shared keys. Also fixed along the way: the in-app updater's relaunch resolver pointed at the upstream `apps/desktop/` layout (never matching `ui-desktop/`), its bash relaunch handoff now honestly reports manual-restart on Windows, and a hashbang broke vitest collection of the packaging tests.
- **2026-08-09:** **Native Windows support for the CLI** — ClawCodex now runs first-class on Windows 10/11 (PowerShell / cmd / Windows Terminal, no WSL required), installed with one line: `irm https://clawcodex.app/install.ps1 | iex`. The new `install.ps1` mirrors `install.sh` end to end — uv install, Python provisioning, lock-pinned deps, PATH registration, TUI build, plus the same `doctor` / `verify` / `update` / `uninstall` lifecycle. Under the hood the port is structural, not cosmetic: a shell platform layer resolves **Git Bash** for the Bash tool (explicitly refusing the WSL `System32\bash.exe` shim), so shell commands keep their POSIX semantics everywhere; process-tree kills go through `taskkill /T`; the mailbox/transcript/lockfile gain real `msvcrt` locking (the Windows CRT's `O_APPEND` emulation can silently *lose* concurrent writes); the permission layer folds NTFS case-insensitivity and refuses whole-drive grants and drive-relative escapes (deny-side only); and `@`-mentions, `CLAWCODEX.md` `@includes`, and persistent-`cd` tracking all round-trip real `C:\` paths. The full test suite now runs on `windows-latest` alongside Ubuntu in CI.
- **2026-08-08 (v1.5.0):** **ClawCodex Desktop — the whole agent in a native app (#802–#808)** — ClawCodex now ships a real desktop application (`ui-desktop/`): streaming chat with a live tool trail and reasoning, permission approvals with once/session/always grants, a session sidebar that lists and resumes the **same durable sessions as the TUI**, side-by-side previews, settings, and the official pixel-art crab as the dock icon and in-app brand mark. `clawcodex desktop` launches it from a checkout (first run installs the UI deps; `--no-dev` builds once and launches Electron directly). The architecture is the interesting part: the app spawns **`clawcodex serve`** — one loopback port serving `/api/*` REST plus a JSON-RPC WebSocket gateway at `/api/ws` — and sessions run on the **same in-process agent core the TUI uses**, so both surfaces share one config, one session store, one skills set, and one permission system; the wire contract is the TUI's own gateway vocabulary, adapted server-side. The port itself is one of the largest single features ClawCodex has landed: ~310K lines of TypeScript across ~1,500 files brought over from the reference desktop implementation, rebranded end to end, with every quality gate held to the reference's own baseline (typecheck green, lint identical, unit-test failure set byte-identical) and the whole loop verified live — boot → real chat turn → streamed reply rendered in the window. macOS packaging works today (`npm run dist:mac` → DMG/zip via electron-builder, hardened-runtime config in place). Ship-week fixes landed the same day: the root `.gitignore`'s Python-oriented `lib/` pattern had silently kept 180 renderer source files out of the initial import — fresh clones failed at boot until #806; the default UI scale moved from the reference's dense 90% preset to Chromium's 100% actual size (#807); and a second `clawcodex desktop` now gets a friendly "already running" message instead of a vite stack trace (#808).
- **2026-08-02 (v1.4.0):** **Fusion models — give a text-only model vision (#771, #787)** — several strong reasoning models cannot see images at all: `deepseek-v4-pro` rejects an image content block outright (`400 unknown variant \`image_url\``), so pasting a screenshot, `@`-mentioning one, or letting `Read` return one ended the turn. A **fusion model** pairs that base model with a second, vision-capable one — every image is described by the vision model first, and the base model reads the description. `/fusion create <name> <base> <vision>` saves one; it then behaves like a normal model in the `/model` picker, as `--model <name>`, in `-p`, and across restarts. Ported from [claude-code-router](https://ccrdesk.top/en/configuration/fusion-models/)'s Fusion Model concept, with one deliberate difference: CCR is a proxy, so it can only offer vision as a *tool* the model may choose to call — which cannot help a pasted image, already on the wire before the model gets a turn. ClawCodex owns the agent loop, so it substitutes images in place, covering paste, `@file.png`, `Read`, and Bash image output at once. Verified end to end on Terminal-Bench 2.1's `code-from-image` task — transcribing handwritten pseudocode from a PNG and reproducing its output — with `deepseek-v4-flash` + `openai:gpt-5.6-luna` (#787); the base model alone returns a 400 on the same image. **Also in v1.4.0:** GPT-5.6 Sol/Terra/Luna (#773); four more OpenAI-compatible providers — groq, cerebras, baseten, xai — taking the registry to 30 (#784); `/mode` becomes `/permissions` with a three-level picker and Full Access by default (#768); `AskUserQuestion` finally renders a real picker instead of returning JSON to the model (#774); the OpenAI provider now picks its wire protocol from the model rather than the auth mode, which is what makes `gpt-5.6-luna` usable on an API key (#783); cached prompt tokens are billed at the cache rate instead of the full input rate, and OpenRouter's streamed reasoning is no longer discarded (#785, #786); and headless runs stop reporting a cut-short run as a success (#777–#782).
- **2026-07-29 (v1.3.0):** **ClawCodex scores 80.9% on Terminal-Bench 2.1 — a top-tier open-source result on Opus 5 (#720–#725, #747–#754)** — running headless on `claude-opus-5` at `effort=xhigh`, ClawCodex solved **72 of 89** Terminal-Bench 2.1 tasks: **80.9% pass@1** on a single run. On the [public 2.1 leaderboard](https://www.tbench.ai/leaderboard/terminal-bench/2.1) (k=5 averages) that would slot **around third** — behind Claude Code / Fable 5 (83.8%) and Codex / GPT-5.5 (83.1%), statistically level with the 79–80% cluster, and **ahead of Claude Code on Opus 4.8 (78.9%) and Sonnet 5 (74.6%)**. Getting there was open, unglamorous parity work: a Harbor eval adapter (`eval/harbor/`) for three-way ClawCodex-vs-openclaude-vs-Claude-Code runs (#720, #724, #725), then a run of prompt- and reliability-parity fixes — restored task-tool skip conditions and parallel-tool guidance, deferred nonessential initial tools, and recovery of trials lost to empty turns and transport drops (#747–#754). **Also in v1.3.0:** `claude-opus-5` support with an interactive `/effort` fix (#746), bounded persistent memory with a background self-improvement review (#731), a VS Code extension driving the agent-server over stdio (#727), image-paste input with an `[Image #N]` un-attach chip (#761, #762), the `CLAUDE.md → CLAWCODEX.md` context-file rebrand (#732), and transport-retry hardening (#757, #760). Stated plainly: this is a single k=1 pass (binomial 1σ ±4.2pp) against the board's k=5 ± ~1.2pp averages, benchmarked on `main` at #756 (before the v1.3.0 tag), so read it as directional rather than a ranked submission.
Expand Down
Loading
Loading