Repository navigation
feat(bench): agent A/B harness for cgis vs no cgis - #547
Conversation
Runner, scoring and guard for a headless `claude -p` benchmark that asks the same code questions with and without the cgis MCP server. - cgis.bench.agent_task: task YAML, hand-checked answer keys, deterministic recall/precision from the answer's closing JSON block. - cgis.bench.transcript: stream-json parser; tool calls, reads, cost, residual context, plus sufficiency and allocation for #220. - cgis.bench.guard: PreToolUse hook blocking the cgis CLI, uv, sqlite3 and graph files in both arms; the same predicate flags contaminated runs. - scripts/agent_ab.py: fresh worktree per run, `run` (with --dry-run) and `report`. Not in CI: every real run is a paid session. - Four cgis tasks (impact, flow, orientation, negative control). Re-measured the self.* placeholder band (66 on main, 69 here: three str/dict receivers) and raised the test-count baseline to 2943. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
There was a problem hiding this comment.
Code Review
This pull request introduces an Agent A/B benchmark harness to evaluate agent performance with and without the cgis MCP server, including a runner script, task definitions, parsing and scoring logic, a PreToolUse guard hook, and unit tests. The review feedback provides several valuable improvements to enhance robustness and cross-platform compatibility, such as handling subprocess timeouts gracefully, resolving executable paths on Windows, stripping whitespace from repository arguments, validating the tasks directory, improving JSON block extraction, and deduplicating files in the allocation metric calculation.
Addresses review on #547: a timed-out session keeps its partial transcript instead of ending the batch; entry points resolve with .exe on Windows; --repo strips whitespace; a missing tasks directory fails loudly; the answer is the last JSON block that carries symbols or files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request introduces an Agent A/B benchmark framework to evaluate agent performance with and without the cgis MCP server. It includes a runner script, task definitions, deterministic scoring, a PreToolUse guard hook to prevent direct CLI/graph access, and transcript parsing for behavioral metrics. The review feedback suggests several improvements: expanding symbol normalization to cover JavaScript/JSX extensions, adding a leading word boundary to the graph file regex to avoid false positives, and robustly normalizing file paths before calculating the files_read metric to prevent duplicate counts.
… read count Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request introduces an Agent A/B benchmark harness to evaluate agent performance with and without the cgis MCP server. It includes a runner script, task definitions, a PreToolUse guard hook to prevent direct graph access, and deterministic scoring and metric extraction logic from transcripts, along with comprehensive unit tests. The review feedback suggests improving robustness by safely quoting the Python executable path in the guard hook command to handle spaces, and making the markdown code block language check case-insensitive to avoid scoring failures when agents output uppercase JSON blocks.
… fence Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
|



Requested by Andrei · project thread
Before: there was no way to measure whether an agent answers a code question better or cheaper with the cgis MCP server connected.
After:
uv run python scripts/agent_ab.py runasks the same questions of a headlessclaude -pwith and without cgis, scores each answer against a hand-checked key, and appends one line per run tobenchmarks/agent_ab/results.jsonl;reportprints medians per task and arm.--dry-runprints the commands and spends nothing. This PR runs no paid sessions.Part of #543. Unlike codegraph's benchmark, the primary metric is answer correctness (recall and precision against a key written from the code, not from the graph) alongside cost, so "fewer file reads" cannot win on its own.
How:
cgis.bench.agent_task: task YAML, answer keys with alternative spellings andallowed_*items, deterministic scoring of the answer's closing JSON block.cgis.bench.transcript: stream-json parser; tool calls, distinct files and bytes read, cost, tokens, residual context, plussufficiencyandallocationfor feat(context): adaptive depth + centrality budgeting + L3 domain-graph for cgis context #220.cgis.bench.guard: PreToolUse hook in both arms that refuses the cgis CLI, uv, sqlite3 andgraph.db/graph.json. The same predicate marks a runcontaminatedif a blocked call ever returned output.scripts/agent_ab.py: fresh detached worktree per run at the task's pinned commit; control gets an empty--strict-mcp-configand all graph files deleted (this repo commitsui/public/graph.json); the cgis arm gets this checkout'scgis-mcp, the plugin's skills without its uvx.mcp.json, and a graph built before the clock starts. PATH has no cgis or uv; edits, web and sub-agents are refused.language_forcallers), flow (ingest → SQLite), orientation (call resolution), and a negative control (SQLite pragmas).str/dictreceivers, declined by D1) and raised the test-count baseline to 2943, ascheck_test_count.pyasked.Checks: format, lint, mypy and interrogate pass; the 90 new tests pass with 99% coverage of
cgis.benchandscripts/agent_ab.py. Locally 17 history-dependent tests (backfill fingerprints, recordings from corpus) fail because the clone is shallow; they fail the same way without this change.🤖 Generated with Claude Code
https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
Generated by Claude Code