Repository navigation
feat(review): the local review leaves its findings for an agent and reviews again what changed - #17
Conversation
…eviews again what changed local-review.sh writes its verdict to .git/dx-review/ (never committed): findings.md, a list to fix with the file and line of each finding, and last.json. The next run on the same branch is incremental, as on a pull request: the brief carries the earlier findings and only what changed since, and the review says which are fixed; with nothing new since, it answers from the file without spending a review; --full reviews everything again. It now says Not ready on everything that would stop the pull request on GitHub: a broken convention, a blocker or a major, a description the review says does not match the code, and a title of another type than the change (GitHub relabels it and the conventions check then fails). The self-test covers each with a fake reviewer (LOCAL_REVIEW_CLAUDE), so it spends nothing. AGENTS.md tells an agent to fix what findings.md lists and run it again until it is ready. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r can read them A pull request that changed only languages/*.mo counted as low risk and could merge on its own, though the brief leaves binaries out, so neither the review nor a person had seen what changed. .po and .pot, which are text, stay low risk. policy.py --test covers both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…licy's --model <id> picks the reviewer's model for one run, for a second opinion from a stronger or a newer model; it must be a model id. The self-test covers it with the fake reviewer, and a value that is not an id. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…me commit is never incremental The local review answered from its last verdict whenever the commit had not changed, so fixing only the title or the description returned the same Not ready, and an agent following 'fix, run again' would loop. The last verdict now records the title and description it saw, and a run with a different one reviews again. On the commit already reviewed, --no-claude builds the whole brief instead of an incremental one with an empty diff. The comment and CONTRIBUTING.md give the real reason a title of the wrong type is not ready (on GitHub the review retitles the pull request), the self-test keeps its description file inside the scratch repository, and CONTRIBUTING.md mentions --model. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude review · risk high · complexity high · type featIncremental review, 90a05c7..9984338. The new commits fix both earlier findings. The cache key is now No findings. Policy floor: high (touches high-risk paths: .github/workflows/pull-request.yml, AGENTS.md, CONTRIBUTING.md, README.md, policy/review-policy.default.yml …). Reviewed 9984338 (since 90a05c7; review 2 of 5 automatic). Author trusted for auto-merge: true. 🤖 AI review · claude-opus-5-5 (Anthropic) · $0.30, 7 turns |
…base, with a hash macOS has The local review hashed the title and the description with sha256sum, which stock macOS does not ship, so every run there stopped before the brief under set -e; it uses git hash-object now. The key also takes the model, the profile and the base, so --model on a commit already reviewed asks the other model instead of returning the cached verdict. The self-test covers --model on a reviewed commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A new job runs the self-tests of conventions.sh, review-brief.sh, local-review.sh and policy.py on macos-latest, with the stock bash 3.2 as every bash the scripts call and BSD tools: what works only with GNU tools, like the sha256sum the review just caught, fails there instead of on a contributor's laptop. CONTRIBUTING.md lists what the local review needs, and how to get it on macOS. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cache key already took the profile and the base; two cases now show it, so dropping either would fail the self-test. CONTRIBUTING.md names the profile among what keys the cache. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
📝 What changes
scripts/local-review.shnow leaves its findings where an agent can act on them, and reviews again the way the pull request review does:.git/dx-review/(never committed):findings.md, a list to fix with the file and line of each finding, andlast.json.git hash-object, which macOS has), it answers from the file without spending a review; a new title or description is reviewed again;--fullreviews everything again.--model <id>picks the reviewer's model for one run instead of the policy's (a second opinion from a stronger or newer model).The self-test covers the verdict path with a fake reviewer (
LOCAL_REVIEW_CLAUDE), so it spends nothing: a clean verdict, a major, a title of the wrong type, a description that does not match, a second run with nothing new, a retitle on the same commit, an incremental run and the earlier findings reaching its brief,--full,--model. AGENTS.md tells an agent to fix whatfindings.mdlists and run it again until it is ready; CONTRIBUTING.md and the README say the same.A new job, Scripts on macOS, runs the self-tests of the four scripts a contributor runs (conventions, review-brief, local-review, policy) on macos-latest with the stock bash 3.2 and BSD tools, so what only works with GNU tools fails in CI instead of on a contributor's laptop. CONTRIBUTING.md lists what the local review needs (git, bash, jq, python3 with yq or PyYAML, the Claude Code CLI) and how to get it on macOS.
Separately, compiled translations (
languages/*.mo) are no longer low-risk eligible: the brief leaves binaries out, so a pull request that changed only them could merge on its own with nobody, reviewer or person, having seen the change..poand.pot, which are text, stay low risk.policy.py --testcovers both.💡 Why
The local review printed its findings and forgot them: an agent had nothing to work from, and every run reviewed the whole branch again. Development here is meant to happen locally, with the same checks and the same review as the pull request, and to reach GitHub only when it is clean; the agent doing the work needs the findings as a list it can fix and check off.
🧪 How I tested it
--test: conventions, review-brief, local-review (26 cases), policy (25), release-markers, release-ready, stamp-version, next-version.local-review.shitself, with--model claude-fable-5-1. The first run found that a retitle on the same commit returned the old verdict (major) and four minor points; all fixed here, and the second, incremental run checked them. The review on GitHub then found thatsha256sumis not on macOS and that the cache ignored the model: fixed, with a test, and the macOS job added so that class of bug is caught in CI.🤖 AI-generated · Claude Opus 5.5 (Anthropic)