Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,3 +188,42 @@ the signal does not catch confidently wrong rankings.
- Sandbox repo; each case's diffs are small and fit the 3000-char truncation.
- The baseline is already 90% top-1, leaving little room for runbook context to
help; retrieval only delivers the right runbook in about half the cases.

## Model comparison and re-run (2026-10-07)

Same eval set v2, no runbooks, 24 cases × 3 trials per model. Raw results:
`eval/results/baseline_v2_sonnet5_rerun_2026-10-07.json` and
`eval/results/baseline_v2_haiku45_2026-10-07.json`; reproduce with
`python -m eval.report <sonnet file> <haiku file>`.

| Metric | Claude Sonnet 5 (re-run) | Claude Haiku 4.5 |
|---|---|---|
| Top-1 | **91.7% (66/72)** | 80.6% (58/72) |
| Top-3 | 100% | 100% |
| MRR | 0.958 | 0.903 |
| Null-runbook cases top-1 | 20/21 | 18/21 |
| Ranking latency, median | 6.5s | 4.5s |
| Tokens per call, mean (in / out) | 3066 / 665 | 2698 / 536 |
| Cost per ranking call, list price | ~$0.013 | ~$0.005 |

- **Stability:** the Sonnet 5 re-run (91.7%) matches the Phase 9 baseline
(90.3%) within one trial, so the 90% figure is not a lucky run.
- **Haiku 4.5 is ~2.4x cheaper and ~2s faster per ranking but 11 points less
accurate** (paired: 2 case wins, 5 losses, 17 ties). Its losses include
`config_reformat_hidden_string` and `worker_queue_renamed` (0/3 each), the
cases where the decoy is most surface-plausible. At ~$0.013 per incident
the accuracy is worth more than the saving, so **Sonnet 5 stays the
default**. `SENTINEL_MODEL` makes the choice configurable.
- Haiku 4.5 runs without extended thinking by default, while Sonnet 5 thinks
adaptively; this compares the models as shipped by default, not at matched
reasoning effort.
- Costs use list prices ($2/$10 per MTok for Sonnet 5, $1/$5 for Haiku 4.5)
times the metered tokens, for the ranking call only (the postmortem is a
separate call).

**Close-call rule, out of sample.** The 0.1 rule was chosen on the Phase 9
trials. On the fresh Sonnet 5 re-run it flags **4/6 wrong picks with 7/66
false alarms** (in-sample it was 9/13 and 9/131), so it holds. On Haiku 4.5 it
flags 7/14 wrong but 22/58 right picks, and Haiku's raw confidence is identical
for right and wrong picks (0.93 / 0.93): the rule is calibrated for Sonnet 5
and should be re-checked if the model changes.
2 changes: 1 addition & 1 deletion core/services/llm.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
from dotenv import load_dotenv

load_dotenv()
MODEL = "claude-sonnet-5"
MODEL = os.getenv("SENTINEL_MODEL", "claude-sonnet-5")


@cache
Expand Down
30 changes: 24 additions & 6 deletions eval/report.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,9 +40,24 @@ def summarize(records: list[dict]) -> dict:
"n_right": len(right),
"n_wrong": len(wrong),
"latency_median_s": statistics.median(r["rank_latency_s"] for r in records) if records else None,
# runs before token metering have no usage; report None rather than a fake zero
"tokens_in_mean": _mean_usage(records, "input_tokens"),
"tokens_out_mean": _mean_usage(records, "output_tokens"),
}


def _mean_usage(records: list[dict], key: str):
vals = [r["usage"][key] for r in records if r.get("usage")]
return round(statistics.mean(vals)) if vals else None


def _label(a: dict, b: dict) -> tuple[str, str]:
"""Name the two runs by whatever differs: condition for an A/B, model for a model comparison."""
if a["condition"] != b["condition"]:
return a["condition"], b["condition"]
return a["model"], b["model"]


def retrieval(records: list[dict]) -> dict:
"""Retrieval is deterministic per case, so score one record per case."""
per_case = {}
Expand Down Expand Up @@ -114,6 +129,8 @@ def print_summary(name: str, res: dict):
f"wrong {_fmt_conf(s['conf_wrong'])} (n={s['n_wrong']})"
)
print(f" rank latency median {s['latency_median_s']}s (n={s['n']})")
if s["tokens_in_mean"] is not None:
print(f" tokens/call in {s['tokens_in_mean']}, out {s['tokens_out_mean']} (mean)")
if res.get("aborted"):
print(f" WARNING run ABORTED, not a valid measurement: {res['aborted']}")
truncated = sorted({r["case_id"] for r in res["records"] if r["truncated_diffs"]})
Expand All @@ -137,20 +154,21 @@ def print_retrieval(res: dict):

def print_comparison(a: dict, b: dict):
rows = paired(a["records"], b["records"])
la, lb = _label(a, b)
wins = sum(hb / tb > ha / ta for _, ha, ta, hb, tb in rows)
losses = sum(hb / tb < ha / ta for _, ha, ta, hb, tb in rows)
print(f"\n== Paired per-case top-1 hits ({a['condition']} vs {b['condition']})")
print(f" {'case':<34} {a['condition']:>9} {b['condition']:>9}")
print(f"\n== Paired per-case top-1 hits ({la} vs {lb})")
print(f" {'case':<34} {la:>15} {lb:>15}")
for cid, ha, ta, hb, tb in rows:
mark = "+" if hb / tb > ha / ta else "-" if hb / tb < ha / ta else " "
print(f" {cid:<34} {ha:>5}/{ta:<3} {hb:>5}/{tb:<3} {mark}")
print(f" {cid:<34} {ha:>11}/{ta:<3} {hb:>11}/{tb:<3} {mark}")
ties = len(rows) - wins - losses
print(f" {b['condition']} wins {wins}, losses {losses}, ties {ties} (n={len(rows)} cases)")
print(f" {lb} wins {wins}, losses {losses}, ties {ties} (n={len(rows)} cases)")

print("\n== Null-runbook cases only (does irrelevant context hurt?)")
for res in (a, b):
for label, res in ((la, a), (lb, b)):
s = summarize([r for r in res["records"] if not r["expected_runbook_id"]])
print(f" {res['condition']:<9} top-1 {_pct(s['top1'], s['n'])}, MRR {s['mrr']:.3f}")
print(f" {label:<15} top-1 {_pct(s['top1'], s['n'])}, MRR {s['mrr']:.3f}")


def main(argv=None):
Expand Down
2 changes: 2 additions & 0 deletions eval/results/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,4 +6,6 @@ Each file records `sentinel_git_sha`, the commit the run used. Phase 9 was rebas
those SHAs (`7b4dc90`, `2e1eb8d`, `530b73f`) live on the preserved branch `feat/phase9-rag-eval`, not on
`main`. Do not delete that branch.

The 2026-10-07 model-comparison runs record `dfd2797`, preserved on branch `feat/model-comparison`; keep it too.

Runs under `aborted/` are kept for the record and are never scored as measurements.
Loading
Loading