Repository navigation
Phase 9: diagnostic-accuracy eval + runbook context in ranking (measured, left off) - #8
Merged
Merged
Conversation
The alert-dedup merge (#6) landed with two files that failed the `black --check --line-length 110` backend gate. Formatting only — no logic change; 20 core tests still pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… as v1 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2-failed baseline Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nRouter Same model (claude-sonnet-5) and same ranking prompt; forced rank_commits tool call now uses the native tool_use block. max_tokens raised to 16000 because Sonnet 5 runs adaptive thinking by default and thinking counts toward the cap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…CE_ID is set Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1 after ceiling Every case now carries a decoy that matches the alert as well as the culprit at first glance; label.note records why the label is right (scorer-only). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…L_RAG_RANKING - find_matching_runbooks(include_content=True) adds Symptoms + Root Causes (capped) - rank_suspect_commits(runbooks=) inserts a <runbooks> block only above the floor (0.35, frozen) - flag defaults off; prompt without runbooks is byte-identical to the golden baseline - retrieval failure logs and ranks without runbooks instead of reverting the incident - runbook text is prompt-only; persisted matched_runbooks shape unchanged Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RAG in ranking measured with no benefit at this corpus size; SENTINEL_RAG_RANKING stays off. README cites only the baseline accuracy from METRICS.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # core/orchestrator.py
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Answers one question with a measurement: does giving Claude the matched runbooks improve its ability to find the commit that caused an incident? At this corpus size, no measurable benefit, so
SENTINEL_RAG_RANKINGships off. Full write-up:METRICS.md→ Phase 9,ADR.md→ ADR-7.On the 10 cases where runbook text actually reached the prompt, both conditions scored 27/30; the single extra RAG hit came from a case whose prompt was byte-identical, i.e. noise.
What changed
eval/): builds each case as a throwaway git repo and calls the productionget_recent_diffs/find_matching_runbooks/rank_suspect_commits. Scorer (top-1/top-3/MRR, paired table, retrieval + floor table), aborts after 5 consecutive LLM errors, results committed undereval/results/.SENTINEL_RAG_RANKING(default0), floorSENTINEL_RUNBOOK_FLOOR= 0.35 chosen from baseline retrieval scores only. Golden-string test pins the no-runbook prompt byte-for-byte. Retrieval failure now logs and ranks without runbooks instead of reverting the incident. Runbook text is prompt-only, never persisted.core/services/llm.py), same modelclaude-sonnet-5,max_tokens16000 (Sonnet 5 thinks by default). Env var is nowANTHROPIC_API_KEY; optionalANTHROPIC_WORKSPACE_ID.Merge note
Please merge with a merge commit, not squash: the results files record the Sentinel commit each run used (
7b4dc90,2e1eb8d,530b73f), and squashing would orphan those SHAs.1dd55aeis the already-squash-merged PR #7 commit; it is a no-op againstmain.Test plan
black --check --line-length 110 core sandbox eval,flake8 core sandbox evalpytest core/tests eval/tests -q— 83 passed, all offline (vector store and Claude mocked)python -m eval.report eval/results/baseline_v2_2026-10-04.json eval/results/rag_v2_2026-10-04.jsonreproduces every number in METRICS.mdeval/)🤖 Generated with Claude Code