Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/test-count-baseline
Original file line number Diff line number Diff line change
@@ -1 +1 @@
2775
2943
3 changes: 3 additions & 0 deletions benchmarks/agent_ab/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Raw stream-json sessions: large, and they quote repository source verbatim
# (owner-api is private). results.jsonl keeps the scored line for each.
transcripts/
84 changes: 84 additions & 0 deletions benchmarks/agent_ab/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# Agent A/B: the same questions with and without cgis (#543)

Does an agent answer a code question better, or cheaper, when the cgis MCP
server is connected? This directory holds the questions, their answer keys and
the results. The runner is `scripts/agent_ab.py`; parsing and scoring live in
`src/cgis/bench/` and are unit-tested on recorded transcripts, so the harness
can change without spending anything.

Design notes and cost estimates: `research/543.md` in the project files.

## What is measured

**Primary, fixed before the first paid run:**

- `recall`: the share of the answer key the answer names (symbols, files and
literal facts, pooled). `precision`: the share of what it names that is in
the key or in its `allowed_*` lists.
- `cost_usd`: `total_cost_usd` from the session's closing event.

The hypothesis is that the `cgis` arm reaches at least the control arm's recall
for no more cost on `impact` and `flow` questions, and that `control`-type
questions show what cgis costs when it cannot help.

**Secondary:** tool calls by name, distinct files read and bytes read, turns,
wall time, token usage, and the context the session ends holding
(`residual_context`).

**For #220, not for ranking arms:** `sufficiency` (what the agent did right
after each cgis answer) and `allocation` (the share of files cgis named that
the final answer relied on). Definitions are in the docstring of
`src/cgis/bench/transcript.py`.

The answer key is written from the code and checked by hand, never taken from
the cgis graph: a key derived from the graph would agree with every edge the
resolver gets wrong. Each task's `notes` say how it was checked.

## Arms

| arm | MCP | plugin skills | graph files | cgis / uv on PATH |
|---|---|---|---|---|
| `control` | none (`--strict-mcp-config`, empty config) | no | all deleted | no |
| `cgis` | this checkout's `cgis-mcp` | yes (the plugin minus its `.mcp.json`) | `graph.db` built before the clock starts | no |

Both arms run under `cgis.bench.guard` as a PreToolUse hook, which refuses the
cgis CLI, uv, sqlite3 and any read of `graph.db`/`graph.json` through Bash or
the file tools. Without it the control arm is not a control: codegraph's own
benchmark caught its control agent calling their CLI through Bash in 26 of 28
runs. The same predicate marks a finished run `contaminated` if a blocked call
ever returned output; such runs are counted in the report and left out of the
medians.

Both arms allow `Read`, `Grep`, `Glob` and `Bash` (plus `mcp__cgis` in the
treatment arm, which has no server in control) and refuse edits, web access and
sub-agents, under `--permission-mode dontAsk`. Each session starts in a fresh
detached worktree at the task's pinned commit, with `--no-session-persistence`
and `--setting-sources project`, so user settings and earlier sessions do not
reach it.

## Running

```bash
uv run python scripts/agent_ab.py run --dry-run # commands only, spends nothing
uv run python scripts/agent_ab.py run --task cgis-impact-language-for --runs 2
uv run python scripts/agent_ab.py run --repo owner-api=../ownima-backend
uv run python scripts/agent_ab.py report
```

`--model` defaults to `claude-sonnet-5-5`; `--max-budget-usd` (default 2.00)
caps each session. Every non-dry run is a paid session, so this never runs in
CI. Transcripts are written to `transcripts/` (git-ignored); `results.jsonl`
gets one line per (task, arm, run) and is meant to be committed with the write-up.

## Tasks

| id | type | repo |
|---|---|---|
| `cgis-impact-language-for` | impact | cgis |
| `cgis-flow-ingest-to-sqlite` | flow | cgis |
| `cgis-orientation-call-resolution` | orientation | cgis |
| `cgis-control-sqlite-pragmas` | control | cgis |

This repository's `CLAUDE.md` describes its architecture in some detail. It is
the same in both arms, but it lowers what a graph can add here, which is one
reason owner-api and httpx come next.
20 changes: 20 additions & 0 deletions benchmarks/agent_ab/tasks/cgis-control-sqlite-pragmas.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
id: cgis-control-sqlite-pragmas
repo: cgis
sha: 715cd9a2ecee47949651344757edb4c7d09f64d7
src_root: src
type: control
question: |
Which SQLite journal mode and busy timeout does `SQLiteStore` configure when it
opens a database, and in which method?
gold:
symbols:
- SQLiteStore.__init__
files:
- src/cgis/storage/sqlite_store.py
facts:
- WAL
- "5000"
notes: |
Negative control: one file, one method, no cross-file structure. A graph should
not help here; this measures what cgis costs when it is not needed.
sqlite_store.py:83-84.
50 changes: 50 additions & 0 deletions benchmarks/agent_ab/tasks/cgis-flow-ingest-to-sqlite.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
id: cgis-flow-ingest-to-sqlite
repo: cgis
sha: 715cd9a2ecee47949651344757edb4c7d09f64d7
src_root: src
type: flow
question: |
Trace what happens when `cgis ingest <path> --output graph.db` runs: starting
from the CLI command, list in order the functions and methods on the main path
that parse source files, resolve calls, and write nodes and edges to SQLite.
gold:
symbols:
- cli.ingest
- IngestionPipeline.run
- IngestionPipeline._process_file
- id: extractor.parse
any_of: [BaseExtractor.parse, PythonExtractor.parse, TypeScriptExtractor.parse]
- ResolverEngine.resolve
- IngestionPipeline._persist_incremental
- SQLiteStore.save_incremental_batch
files:
- src/cgis/cli.py
- src/cgis/pipeline.py
- src/cgis/resolver/engine.py
- src/cgis/storage/sqlite_store.py
allowed_symbols:
- IngestionPipeline.workspace_root
- SQLiteStore.record_ingest
- SQLiteStore.record_ingest_options
- build_extractors
- IngestionPipeline._get_extractor
- ImportNameCollector.collect
- SemanticUpliftEngine.execute_uplift
- SymbolResolver
- IndexBuilder.build
allowed_files:
- src/cgis/extractors/base.py
- src/cgis/extractors/python_extractor.py
- src/cgis/extractors/typescript_extractor.py
- src/cgis/extractors/registry.py
- src/cgis/resolver/symbols.py
- src/cgis/resolver/indices.py
- src/cgis/resolver/uplift.py
- src/cgis/import_names.py
notes: |
Read 2026-10-02: cli.ingest (cli.py:174) opens SQLiteStore and calls
IngestionPipeline.run(store=..., rebuild=True); run calls _process_file per file,
which calls the extractor's parse; then ResolverEngine.resolve (pipeline.py:213);
then _persist_incremental (pipeline.py:237), which calls
SQLiteStore.save_incremental_batch (pipeline.py:411). Order is not scored, only
presence.
27 changes: 27 additions & 0 deletions benchmarks/agent_ab/tasks/cgis-impact-language-for.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
id: cgis-impact-language-for
repo: cgis
sha: 715cd9a2ecee47949651344757edb4c7d09f64d7
src_root: src
type: impact
question: |
In this repository, the function `language_for` in `src/cgis/extractors/registry.py`
is about to change its signature. List every function or method that calls it
directly, so each call site can be updated.
gold:
symbols:
- cli.structure
- registry.is_supported
- SourceCollector.collect_full_files
- GraphContextCollector.sections
- mermaid._node_slug
files:
- src/cgis/cli.py
- src/cgis/extractors/registry.py
- src/cgis/guardian/collector.py
- src/cgis/query/render/mermaid.py
allowed_symbols:
- registry.language_for
notes: |
Checked 2026-10-02 by grep (`language_for(` outside imports) and by reading each
enclosing def: cli.py:843, registry.py:88, collector.py:175 and :214,
mermaid.py:77. The cgis graph agrees, but the key was not taken from it.
36 changes: 36 additions & 0 deletions benchmarks/agent_ab/tasks/cgis-orientation-call-resolution.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
id: cgis-orientation-call-resolution
repo: cgis
sha: 715cd9a2ecee47949651344757edb4c7d09f64d7
src_root: src
type: orientation
question: |
Which files implement turning unresolved call targets (`raw_call:<name>` edges
emitted by the extractors) into fully qualified names, and what does each of
them contribute?
gold:
symbols:
- ResolverEngine
- SymbolResolver
- id: indices
any_of: [IndexBuilder, SymbolIndex]
files:
- src/cgis/resolver/engine.py
- src/cgis/resolver/symbols.py
- src/cgis/resolver/indices.py
allowed_symbols:
- ResolverEngine.resolve
- ResolverEngine._resolved_call_edge
- js_builtin_target
- IngestionPipeline.run
- Edge
allowed_files:
- src/cgis/resolver/js_builtins.py
- src/cgis/core/models.py
- src/cgis/pipeline.py
- src/cgis/extractors/python_extractor.py
- src/cgis/extractors/typescript_extractor.py
- src/cgis/extractors/base.py
notes: |
Read 2026-10-02: engine.py drives resolution, symbols.py holds the strategies
over a SymbolIndex, indices.py builds the indices (module docstrings say so).
js_builtins.py and models.py are correct but optional.
Loading
Loading