Skip to content

feat(bench): agent A/B harness for cgis vs no cgis - #547

Merged
zaebee merged 5 commits into
mainfrom
claude/project-thread-1y82x2
Oct 2, 2026
Merged

zaebee merged 5 commits into
mainfrom
claude/project-thread-1y82x2

Conversation

@zaebee

@zaebee zaebee commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Requested by Andrei · project thread

Before: there was no way to measure whether an agent answers a code question better or cheaper with the cgis MCP server connected.

After: uv run python scripts/agent_ab.py run asks the same questions of a headless claude -p with and without cgis, scores each answer against a hand-checked key, and appends one line per run to benchmarks/agent_ab/results.jsonl; report prints medians per task and arm. --dry-run prints the commands and spends nothing. This PR runs no paid sessions.

Part of #543. Unlike codegraph's benchmark, the primary metric is answer correctness (recall and precision against a key written from the code, not from the graph) alongside cost, so "fewer file reads" cannot win on its own.

How:

  • cgis.bench.agent_task: task YAML, answer keys with alternative spellings and allowed_* items, deterministic scoring of the answer's closing JSON block.
  • cgis.bench.transcript: stream-json parser; tool calls, distinct files and bytes read, cost, tokens, residual context, plus sufficiency and allocation for feat(context): adaptive depth + centrality budgeting + L3 domain-graph for cgis context #220.
  • cgis.bench.guard: PreToolUse hook in both arms that refuses the cgis CLI, uv, sqlite3 and graph.db/graph.json. The same predicate marks a run contaminated if a blocked call ever returned output.
  • scripts/agent_ab.py: fresh detached worktree per run at the task's pinned commit; control gets an empty --strict-mcp-config and all graph files deleted (this repo commits ui/public/graph.json); the cgis arm gets this checkout's cgis-mcp, the plugin's skills without its uvx .mcp.json, and a graph built before the clock starts. PATH has no cgis or uv; edits, web and sub-agents are refused.
  • Four cgis tasks: impact (language_for callers), flow (ingest → SQLite), orientation (call resolution), and a negative control (SQLite pragmas).
  • Re-measured the self.* placeholder band (66 on main, 69 here: three str/dict receivers, declined by D1) and raised the test-count baseline to 2943, as check_test_count.py asked.

Checks: format, lint, mypy and interrogate pass; the 90 new tests pass with 99% coverage of cgis.bench and scripts/agent_ab.py. Locally 17 history-dependent tests (backfill fingerprints, recordings from corpus) fail because the clone is shallow; they fail the same way without this change.

🤖 Generated with Claude Code

https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa


Generated by Claude Code

Runner, scoring and guard for a headless `claude -p` benchmark that asks the
same code questions with and without the cgis MCP server.

- cgis.bench.agent_task: task YAML, hand-checked answer keys, deterministic
  recall/precision from the answer's closing JSON block.
- cgis.bench.transcript: stream-json parser; tool calls, reads, cost, residual
  context, plus sufficiency and allocation for #220.
- cgis.bench.guard: PreToolUse hook blocking the cgis CLI, uv, sqlite3 and
  graph files in both arms; the same predicate flags contaminated runs.
- scripts/agent_ab.py: fresh worktree per run, `run` (with --dry-run) and
  `report`. Not in CI: every real run is a paid session.
- Four cgis tasks (impact, flow, orientation, negative control).

Re-measured the self.* placeholder band (66 on main, 69 here: three str/dict
receivers) and raised the test-count baseline to 2943.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
@zaebee zaebee self-assigned this Oct 2, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an Agent A/B benchmark harness to evaluate agent performance with and without the cgis MCP server, including a runner script, task definitions, parsing and scoring logic, a PreToolUse guard hook, and unit tests. The review feedback provides several valuable improvements to enhance robustness and cross-platform compatibility, such as handling subprocess timeouts gracefully, resolving executable paths on Windows, stripping whitespace from repository arguments, validating the tasks directory, improving JSON block extraction, and deduplicating files in the allocation metric calculation.

Comment thread scripts/agent_ab.py Outdated
Comment thread scripts/agent_ab.py
Comment thread scripts/agent_ab.py Outdated
Comment thread scripts/agent_ab.py Outdated
Comment thread scripts/agent_ab.py
Comment thread src/cgis/bench/agent_task.py
Comment thread src/cgis/bench/agent_task.py Outdated
Comment thread src/cgis/bench/transcript.py
claude added 2 commits October 2, 2026 13:23
Addresses review on #547: a timed-out session keeps its partial transcript
instead of ending the batch; entry points resolve with .exe on Windows;
--repo strips whitespace; a missing tasks directory fails loudly; the answer
is the last JSON block that carries symbols or files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
@zaebee
zaebee marked this pull request as ready for review October 2, 2026 13:28
@zaebee

zaebee commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an Agent A/B benchmark framework to evaluate agent performance with and without the cgis MCP server. It includes a runner script, task definitions, deterministic scoring, a PreToolUse guard hook to prevent direct CLI/graph access, and transcript parsing for behavioral metrics. The review feedback suggests several improvements: expanding symbol normalization to cover JavaScript/JSX extensions, adding a leading word boundary to the graph file regex to avoid false positives, and robustly normalizing file paths before calculating the files_read metric to prevent duplicate counts.

Comment thread src/cgis/bench/agent_task.py Outdated
Comment thread src/cgis/bench/guard.py Outdated
Comment thread src/cgis/bench/transcript.py Outdated
Comment thread src/cgis/bench/transcript.py Outdated
… read count

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L7RD9DDX3xidUCfgBEtRDa
@zaebee

zaebee commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an Agent A/B benchmark harness to evaluate agent performance with and without the cgis MCP server. It includes a runner script, task definitions, a PreToolUse guard hook to prevent direct graph access, and deterministic scoring and metric extraction logic from transcripts, along with comprehensive unit tests. The review feedback suggests improving robustness by safely quoting the Python executable path in the guard hook command to handle spaces, and making the markdown code block language check case-insensitive to avoid scoring failures when agents output uppercase JSON blocks.

Comment thread scripts/agent_ab.py Outdated
Comment thread src/cgis/bench/agent_task.py Outdated
@sonarqubecloud

sonarqubecloud Bot commented Oct 2, 2026

Copy link
Copy Markdown

@zaebee
zaebee merged commit 788d67f into main Oct 2, 2026
3 checks passed
@zaebee
zaebee deleted the claude/project-thread-1y82x2 branch October 2, 2026 14:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants