Skip to content

Phase 9: diagnostic-accuracy eval + runbook context in ranking (measured, left off) - #8

Merged
Lushenwar merged 20 commits into
mainfrom
feat/phase9-rag-eval
Oct 6, 2026
Merged

Lushenwar merged 20 commits into
mainfrom
feat/phase9-rag-eval

Conversation

@Lushenwar

Copy link
Copy Markdown
Owner

Summary

Answers one question with a measurement: does giving Claude the matched runbooks improve its ability to find the commit that caused an incident? At this corpus size, no measurable benefit, so SENTINEL_RAG_RANKING ships off. Full write-up: METRICS.md → Phase 9, ADR.md → ADR-7.

Eval set v2, 24 cases × 3 trials Baseline RAG (floor 0.35)
Top-1 90.3% (65/72) 91.7% (66/72)
MRR 0.947 0.958
Paired per case 2 wins / 1 loss / 21 ties

On the 10 cases where runbook text actually reached the prompt, both conditions scored 27/30; the single extra RAG hit came from a case whose prompt was byte-identical, i.e. noise.

What changed

  • Eval harness (eval/): builds each case as a throwaway git repo and calls the production get_recent_diffs / find_matching_runbooks / rank_suspect_commits. Scorer (top-1/top-3/MRR, paired table, retrieval + floor table), aborts after 5 consecutive LLM errors, results committed under eval/results/.
  • Cases: v1 (29 cases) hit 87/87 at baseline (ceiling) and was replaced by v2 (24 cases, a decoy per case) before any RAG run. Lint test bans answer-revealing words; leakage test asserts no label/id/category reaches the prompt in either condition.
  • Runbook corpus: 2 → 10 generic runbooks, written before any case, frozen.
  • Core: optional runbook context in ranking behind SENTINEL_RAG_RANKING (default 0), floor SENTINEL_RUNBOOK_FLOOR = 0.35 chosen from baseline retrieval scores only. Golden-string test pins the no-runbook prompt byte-for-byte. Retrieval failure now logs and ranks without runbooks instead of reverting the incident. Runbook text is prompt-only, never persisted.
  • LLM client: OpenRouter → Anthropic SDK direct (core/services/llm.py), same model claude-sonnet-5, max_tokens 16000 (Sonnet 5 thinks by default). Env var is now ANTHROPIC_API_KEY; optional ANTHROPIC_WORKSPACE_ID.
  • Docs: Phase 5 caveat (those runs used commit messages naming the bug — pipeline verification, not accuracy), METRICS Phase 9, ADR-7, one README line citing baseline accuracy only.

Merge note

Please merge with a merge commit, not squash: the results files record the Sentinel commit each run used (7b4dc90, 2e1eb8d, 530b73f), and squashing would orphan those SHAs. 1dd55ae is the already-squash-merged PR #7 commit; it is a no-op against main.

Test plan

  • black --check --line-length 110 core sandbox eval, flake8 core sandbox eval
  • pytest core/tests eval/tests -q — 83 passed, all offline (vector store and Claude mocked)
  • Dashboard lint + tests
  • python -m eval.report eval/results/baseline_v2_2026-10-04.json eval/results/rag_v2_2026-10-04.json reproduces every number in METRICS.md
  • CI green on this PR (CI now also covers eval/)

🤖 Generated with Claude Code

Lushenwar and others added 20 commits July 14, 2026 12:53
The alert-dedup merge (#6) landed with two files that failed the
`black --check --line-length 110` backend gate. Formatting only — no
logic change; 20 core tests still pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… as v1

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2-failed baseline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nRouter

Same model (claude-sonnet-5) and same ranking prompt; forced rank_commits tool
call now uses the native tool_use block. max_tokens raised to 16000 because
Sonnet 5 runs adaptive thinking by default and thinking counts toward the cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…CE_ID is set

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1 after ceiling

Every case now carries a decoy that matches the alert as well as the culprit at
first glance; label.note records why the label is right (scorer-only).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…L_RAG_RANKING

- find_matching_runbooks(include_content=True) adds Symptoms + Root Causes (capped)
- rank_suspect_commits(runbooks=) inserts a <runbooks> block only above the floor (0.35, frozen)
- flag defaults off; prompt without runbooks is byte-identical to the golden baseline
- retrieval failure logs and ranks without runbooks instead of reverting the incident
- runbook text is prompt-only; persisted matched_runbooks shape unchanged

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RAG in ranking measured with no benefit at this corpus size; SENTINEL_RAG_RANKING
stays off. README cites only the baseline accuracy from METRICS.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Lushenwar
Lushenwar merged commit c3f0e29 into main Oct 6, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant