Repository navigation
feat(bench): cgis-instructed and cgis-forced arms, four multi-hop impact tasks - #549
Conversation
The first pilot session named connect and was scored as a miss; the key was wrong, the agent was right (sqlite_store.py:75-84 at the pinned sha). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
The pilot's cgis arm made no cgis call in 12 of 12 sessions, and every task scored recall 1.0 in both arms. Add an arm that appends one system-prompt line telling the agent to query cgis first (a stand-in for #542's server instructions), and two impact tasks a single grep does not answer: a five-hop chain up to the entry points, and a chain whose first hop shares the target's method name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
The cgis-instructed arm called cgis in 2 of 18 sessions. The new arm runs the guard with --cgis-first MARKER: Read, Grep, Glob and Bash are refused until an mcp__cgis__ call creates the marker, so the benchmark can measure what the graph adds once used, apart from whether the agent picks it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
There was a problem hiding this comment.
Code Review
This pull request introduces two new evaluation arms to the agent A/B benchmark: cgis-instructed (which appends a system prompt instructing the agent to use cgis first) and cgis-forced (which uses a guard hook to block direct source-reading tools until a cgis call is made). It also adds new impact-analysis benchmark tasks, updates the guard hook implementation to support the --cgis-first option, and adds corresponding unit tests. Feedback on the changes suggests improving the robustness of the --cgis-first argument parsing in the guard hook and addressing potential cross-platform quoting issues on Windows when using shlex.quote.
A transitive chain from finance.service.credit_wallet to its HTTP routes and worker tasks, through a same-named method and an injected service, and a direct-caller question where another module defines a function with the same name. Keys were read at the pinned sha with an AST call index, not with cgis; the graph misses both routes of the first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
….exe Sonar S8707 flagged the hook writing to a path taken from its argv. The guard now takes a bare --cgis-first flag and keeps its marker in its working directory, the run's fresh worktree, so a later run never sees an earlier run's marker and nothing the session passes picks the path. The hook command uses list2cmdline quoting on Windows, where POSIX single quotes mean nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
|



Before: the agent A/B benchmark (#543) had two arms, and in the 12-session pilot the
cgisarm never called a cgis tool, so it measured nothing about the graph. The four shipped tasks were all solved by grep (recall 1.00 in every arm), and the SQLite pragma task's answer key named the wrong method.After: two more arms and four harder tasks.
cgis-instructedappends one system-prompt line telling the agent to query cgis first (a stand-in for feat(mcp): server-level instructions + CGIS_MCP_TOOLS allowlist #542's server instructions).cgis-forcedalso runs the guard with--cgis-first, which refuses Read/Grep/Glob/Bash until the session has made onemcp__cgis__call. This measures what the graph adds once used, separately from whether the agent picks it.cgis-impact-transitive-tsconfig(a five-hop chain to the entry points) andcgis-impact-drift-same-name(the first hop shares the target's method name).owner-impact-credit-wallet(a same-name method plus DI, up to routes and workers) andowner-impact-occupied-window-same-name(another module defines a function with the same name). These keys were read with an AST call index, not with cgis.SQLiteStore.connect, not__init__(sqlite_store.py:75-84).Results from 96 sessions (Sonnet 5.5, about $12.7 API-equivalent):
cgis-instructedcalled cgis in 2 of 18 sessions on cgis.cgis-forcedcut median cost by 52% on the same-name impact task and by 27% on the five-hop chain. It cost 14–60% more on the simple tasks.cgis_analyze_impactat depth 8 returned 51k characters, and the graph misses calls that FastAPI DI makes into routes.How:
cgis.bench.guard.maintakes a bare--cgis-firstflag. A cgis call creates a marker at a fixed name in the hook's working directory, which is the run's fresh worktree, so no path comes from the hook's arguments (Sonar S8707).list2cmdlineon Windows.🤖 Generated with Claude Code
https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa