Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
1dd55ae
style: black-format dedup changes to fix CI
Lushenwar Jul 14, 2026
24c34e0
feat(core): make Chroma path configurable via SENTINEL_CHROMA_PATH
Lushenwar Oct 4, 2026
1eed3bf
feat(runbooks): expand corpus to 10 generic runbooks, freeze as v1
Lushenwar Oct 4, 2026
8d593b0
feat(eval): offline-buildable eval harness, scorer, and leakage tests
Lushenwar Oct 4, 2026
d6f6bd1
feat(eval): 29-case fault set on a shared orders-service base, frozen…
Lushenwar Oct 4, 2026
8f07d04
fix(eval): ignore untracked files in results git-sha dirty check
Lushenwar Oct 4, 2026
41c8527
style(eval): black-format run_eval
Lushenwar Oct 4, 2026
fd3faba
feat(eval): abort runs after 5 consecutive LLM errors; archive the 40…
Lushenwar Oct 4, 2026
c2d93bd
chore(eval): move 402-failed baseline under results/aborted
Lushenwar Oct 4, 2026
5f93cb7
feat(core): call Claude directly via the Anthropic SDK instead of Ope…
Lushenwar Oct 4, 2026
7b4dc90
feat(core): send anthropic-workspace-id header when ANTHROPIC_WORKSPA…
Lushenwar Oct 4, 2026
b677ac1
feat(eval): baseline run on eval set v1 (87/87 top-1, ceiling)
Lushenwar Oct 4, 2026
2e1eb8d
feat(eval): harder eval set v2 (24 cases, decoy per case), replaces v…
Lushenwar Oct 4, 2026
150d678
feat(eval): baseline run on eval set v2 (65/72 top-1)
Lushenwar Oct 4, 2026
530b73f
feat(core): optional runbook context in commit ranking behind SENTINE…
Lushenwar Oct 4, 2026
4dc3f1f
feat(eval): RAG run on eval set v2 (66/72 top-1, floor 0.35)
Lushenwar Oct 4, 2026
b7b7329
docs: Phase 9 eval results, ADR-7, Phase 5 accuracy caveat
Lushenwar Oct 6, 2026
f02dbb8
Merge remote-tracking branch 'origin/main' into feat/phase9-rag-eval
Lushenwar Oct 6, 2026
e5cb969
docs(core): note the Phase 9 outcome on the RAG ranking flag
Lushenwar Oct 6, 2026
4b5f31d
docs(eval): note where recorded run SHAs live after rebase-merge
Lushenwar Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
@@ -1,8 +1,9 @@
# Required
OPENROUTER_API_KEY=sk-or-...
ANTHROPIC_API_KEY=sk-ant-...
# Required for local (non-Docker) runs; docker-compose provides its own
DATABASE_URL=postgresql://user:password@localhost:5432/sentinel
# Optional — Slack diagnostic cards are skipped if unset
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/...
# Optional — docker-compose Postgres password (defaults to "sentinel")
POSTGRES_PASSWORD=sentinel
ANTHROPIC_WORKSPACE_ID= # only needed if the key is not workspace-scoped
8 changes: 4 additions & 4 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,17 +10,17 @@ jobs:
env:
# tests mock all external boundaries; these only satisfy import-time env lookups
DATABASE_URL: postgresql://ci:ci@localhost:5432/ci
OPENROUTER_API_KEY: ci-dummy
ANTHROPIC_API_KEY: ci-dummy
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
cache: pip
- run: pip install -e core/ pytest jsonschema black flake8
- run: black --check --line-length 110 core sandbox
- run: flake8 core sandbox
- run: pytest core/tests -q
- run: black --check --line-length 110 core sandbox eval
- run: flake8 core sandbox eval
- run: pytest core/tests eval/tests -q

dashboard:
runs-on: ubuntu-latest
Expand Down
41 changes: 41 additions & 0 deletions ADR.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,3 +104,44 @@ Phase 5 verified 3/3 correct #1 rankings *because* ground truth exists.
Trade-off: no exposure to messy production alert noise; the mitigation is that
the simulation's log and alert shapes mirror production profiles (mock Sentry
payloads, structured JSON logs).

---

## ADR-7: Runbook context in commit ranking — built, measured, left off

**Context.** Sentinel already retrieved matching runbooks (Chroma, cosine
similarity) but only used them in the postmortem and the Slack card. The open
question was whether passing them into `rank_suspect_commits` would help
Claude find the faulty commit. Two other retrieval targets were considered and
rejected: **past incidents**, because a sandbox has too little history for
retrieval over it to matter, and **diff embeddings**, because "what changed
recently" is already answered exactly by the git time window.

**Decision.** Runbook text (Symptoms + Root Causes, capped at 1500 chars) can
be passed into the ranking prompt behind `SENTINEL_RAG_RANKING`, **default
off**. Only runbooks scoring at or above `SENTINEL_RUNBOOK_FLOOR` (0.35) are
sent; if none clear it, the block is omitted and the prompt is byte-identical
to the no-runbook prompt (guarded by a golden-string test). The prompt tells
the model runbooks may be irrelevant and that diffs are the evidence. Runbook
text is never persisted. A retrieval failure is logged and ranking proceeds
without runbooks.

The floor was picked from the baseline's retrieval scores before any RAG run
(maximise correct runbooks kept minus null cases given a runbook).

**Measured result** (METRICS.md, Phase 9; 24 cases × 3 trials): baseline
90.3% top-1, RAG 91.7%. On the 10 cases where runbook text actually reached the
prompt, both scored 27/30; the single extra RAG hit came from a case whose
prompt was unchanged. **No measurable benefit**, and no measurable harm on
null-runbook cases.

**Consequences.** The flag stays off, so the shipped ranking path is the
measured baseline. The plumbing and the eval harness stay, because the answer
depends on two things that could change:
- **Retrieval quality.** The correct runbook is top-1 in only 10/17 cases, and
at the floor only 8/17 cases receive it. Better retrieval (richer alert text
than a one-line error, or a stronger embedding model) would raise the
ceiling on what runbooks can contribute.
- **Baseline headroom.** At 90% top-1 there is little left to win. A larger,
harder or real incident corpus — especially real runbooks written by the
team that wrote the code — would be the test that could change this decision.
96 changes: 96 additions & 0 deletions METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,11 @@ Instrumented directly in `core/orchestrator.py` (`[metrics]` log lines):
**Cite the medians: ~13s to surface the faulty commit, ~14s to a complete
postmortem draft.**

> **Caveat (added in Phase 9):** these runs used chaos-CLI commits whose
> messages named the injected bug, with typically one commit in the window;
> they verify the pipeline end-to-end but are not a measure of diagnostic
> accuracy. See [Phase 9](#phase-9--diagnostic-accuracy-and-runbook-context-in-ranking).

## Resume Reconciliation

- The prior `<10s` suspect-commit claim is **not supported** — measured median
Expand Down Expand Up @@ -76,3 +81,94 @@ python -m sandbox.chaos_cli trigger-bug --type db_failure
python -m sandbox.chaos_cli resolve-bug --incident <incident_id>
# wait for "[metrics] ... resolve_to_postmortem_s=" in terminal 1
```

## Phase 9 — Diagnostic accuracy and runbook context in ranking

**Question:** does giving Claude the matched runbooks improve its ability to
find the commit that caused an incident?

**Answer at this corpus size: no measurable benefit.** `SENTINEL_RAG_RANKING`
stays off by default. Raw results:
[`eval/results/`](eval/results/); every table below is reproduced by
`python -m eval.report eval/results/baseline_v2_2026-10-04.json eval/results/rag_v2_2026-10-04.json`.

### Methodology

- **Harness:** `eval/` builds each case as a throwaway git repo (never the
Sentinel repo) and calls the production `get_recent_diffs`,
`find_matching_runbooks` and `rank_suspect_commits`; it never rebuilds the
prompt itself.
- **Model:** `claude-sonnet-5`, Anthropic API direct, provider-default
sampling (non-deterministic), `max_tokens` 16000. **3 trials per case.**
- **Runbook corpus v1:** 10 generic runbooks in `sandbox/runbooks/`, written
before any eval case and frozen at Phase 9B.
- **Eval set v2:** 24 cases, 17 fault categories, 5 commits per case (2 cases
have 4). Culprit position from newest: 0→5 cases, 1→6, 2→6, 3→4, 4→3.
7/24 cases (29%) deliberately have no matching runbook.
- Every case carries a **decoy** that matches the alert as well as the
culprit at first glance: same config key, same subsystem, or the very
line that raises.
- Commit messages are neutral; a lint test bans words that announce the
answer. A leakage test asserts that no case id, category or label value
reaches the prompt in either condition.
- **Eval set v1 (superseded):** the first 29-case set scored **87/87 top-1**
at baseline (`eval/results/baseline_2026-10-04.json`). It was too easy to
show any difference, so it was replaced by v2 before any RAG run (the one
revision the protocol allows).
- **Similarity floor 0.35:** chosen from the v2 baseline retrieval table alone,
before the RAG run, as the candidate maximising (correct runbooks kept −
null cases given a runbook): 0.15→5, 0.25→3, 0.30→6, **0.35→7**, 0.40→5.

### Results (eval set v2, n = 24 cases × 3 trials = 72 per condition)

| Metric | Baseline | RAG (floor 0.35) |
|---|---|---|
| Top-1 (culprit ranked #1) | 90.3% (65/72) | 91.7% (66/72) |
| Top-3 | 100% (72/72) | 100% (72/72) |
| MRR | 0.947 | 0.958 |
| LLM failures (counted as misses) | 0/72 | 0/72 |
| Mean #1 confidence, right / wrong | 0.90 (n=65) / 0.89 (n=7) | 0.89 (n=66) / 0.83 (n=6) |
| Ranking latency, median | 5.8s | 6.0s |

**Paired per case (top-1 hits out of 3):** RAG wins 2, loses 1, ties 21.

| Case | Baseline | RAG | Runbook text in RAG prompt? |
|---|---|---|---|
| `export_tempfile_never_closed` | 0/3 | 1/3 | yes (`file_handle_exhaustion`) |
| `page_size_zero_division` | 2/3 | 3/3 | **no** — prompt byte-identical to baseline |
| `payments_retry_loop` | 3/3 | 2/3 | yes (`rate_limiting`) |
| `httpx_028_drops_proxies` | 0/3 | 0/3 | no (best match 0.29) |

**Where runbooks actually reached the prompt** (10 of 24 cases had a runbook
above the floor): baseline 27/30, RAG 27/30 — identical. The one extra RAG hit
overall came from a case whose prompt was unchanged, i.e. run-to-run noise.

**Null-runbook cases** (7 cases, 21 trials): 21/21 in both conditions. One of
them (`worker_queue_renamed`) received three irrelevant runbooks above the
floor and was still ranked correctly 3/3, so wrong context did not visibly
hurt either.

**How large a difference would matter:** a 1/3 swing on a single case occurred
with a byte-identical prompt, so per-case noise is at least that large. With
only 10 cases where the prompt differs, a real effect would need to show up as
several (roughly 3+) net case wins concentrated in those cases. This is a
judgement from the observed noise, not a significance test; none was computed.

### Retrieval

- Runbook top-1 on cases with an expected runbook: **10/17**; correct runbook
anywhere in top-3: 13/17.
- At floor 0.35: 1 of 7 null cases still receives a runbook, and only 8 of 17
expected cases receive the correct one.
- Indexing only title + Symptoms (11/17) or title + Symptoms + Root Causes
(12/17) was measured offline and did not separate correct matches from null
cases either; retrieval was left unchanged. The limiting factor is short
error strings against the default embedding model.

### Limitations

- Synthetic cases written by the same author as the system and the runbooks.
- Small n: 24 cases × 3 trials; single model (`claude-sonnet-5`).
- Sandbox repo; each case's diffs are small and fit the 3000-char truncation.
- The baseline is already 90% top-1, leaving little room for runbook context to
help; retrieval only delivers the right runbook in about half the cases.
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ Slack, and drafts the postmortem on resolution.

```bash
git clone https://github.com/Lushenwar/Sentinel.git && cd Sentinel
cp .env.example .env # set OPENROUTER_API_KEY (and optionally SLACK_WEBHOOK_URL)
cp .env.example .env # set ANTHROPIC_API_KEY (and optionally SLACK_WEBHOOK_URL)
docker-compose up # dashboard :3000, core API :8000, sandbox app :8001, Postgres :5432

# trigger an incident (in another terminal)
Expand All @@ -30,7 +30,7 @@ docker-compose exec core python -m sandbox.chaos_cli resolve-bug --incident <inc

```bash
pip install -e core/
cp .env.example .env # fill in DATABASE_URL, OPENROUTER_API_KEY
cp .env.example .env # fill in DATABASE_URL, ANTHROPIC_API_KEY
uvicorn core.main:app --port 8000 # terminal 1: core engine
uvicorn sandbox.app.main:app --port 8001 # terminal 2: sandbox toy app
python -m sandbox.chaos_cli trigger-bug --type db_failure
Expand Down Expand Up @@ -84,6 +84,7 @@ Every hop reads and writes one auditable incident row conforming to
## Docs

- [METRICS.md](METRICS.md) — measured end-to-end timings, methodology, and per-run results
- Measured diagnostic accuracy: faulty commit ranked #1 in 90.3% of 72 trials across 24 injected-fault cases with decoy commits ([METRICS.md, Phase 9](METRICS.md#phase-9--diagnostic-accuracy-and-runbook-context-in-ranking))
- [ADR.md](ADR.md) — why explicit state loops over LangChain, polling over WebSockets, Chroma over Pinecone, FastAPI over Flask, tool-call JSON over free-form parsing, sandbox over live infra

## Sandbox reference
Expand Down
20 changes: 16 additions & 4 deletions core/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,8 @@
_DEDUP_WINDOW_S = float(os.getenv("SENTINEL_DEDUP_WINDOW_S", "300"))
_DEDUP_THRESHOLD = int(os.getenv("SENTINEL_DEDUP_THRESHOLD", "1")) # Nth hit in window fires
_recent_alerts: dict[str, list[float]] = defaultdict(list) # signature -> hit times
# Pass matched runbook text into commit ranking. Off: Phase 9 A/B found no measurable benefit (ADR-7).
_RAG_RANKING = os.getenv("SENTINEL_RAG_RANKING", "0") == "1"
_open_signatures: dict[str, str] = {} # signature -> live incident_id


Expand Down Expand Up @@ -69,20 +71,30 @@ def run_diagnostics(incident_id: str, alert_data: dict):
diffs = git_client.get_recent_diffs(REPO_PATH, alert_data["timestamp"])
print(f"[orchestrator] {incident_id} -- {len(diffs)} commits in window")

runbooks = vector_store.find_matching_runbooks(alert_data.get("error_signature", ""))
print(f"[orchestrator] {incident_id} -- {len(runbooks)} runbooks matched")
try:
runbooks = vector_store.find_matching_runbooks(
alert_data.get("error_signature", ""), include_content=_RAG_RANKING
)
print(f"[orchestrator] {incident_id} -- {len(runbooks)} runbooks matched")
except Exception as e: # retrieval is optional context; never a reason to abort triage
runbooks = []
print(f"[orchestrator] {incident_id} -- runbook retrieval failed, ranking without: {e}")

degraded_reason = None
try:
ranked = llm_analyzer.rank_suspect_commits(diffs, alert_data) if diffs else []
rank_runbooks = runbooks if _RAG_RANKING else None
ranked = (
llm_analyzer.rank_suspect_commits(diffs, alert_data, runbooks=rank_runbooks) if diffs else []
)
print(f"[orchestrator] {incident_id} -- LLM ranked {len(ranked)} suspects")
except llm_analyzer.LLMUnavailable as e:
ranked, degraded_reason = [], str(e)
print(f"[orchestrator] {incident_id} -- LLM unavailable, degrading: {e}")

diagnostics = {
"suspect_commits": ranked,
"matched_runbooks": runbooks,
# runbook text is prompt-only; never persisted or sent to Slack
"matched_runbooks": [{k: v for k, v in r.items() if k != "content"} for r in runbooks],
"impact_assessment": {
"error_rate_delta_pct": alert_data.get("error_rate_pct"),
"estimated_affected_users": None,
Expand Down
1 change: 1 addition & 0 deletions core/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ dependencies = [
"psycopg2-binary>=2.9",
"python-dotenv>=1.0",
"httpx>=0.27",
"anthropic>=0.116",
"chromadb>=0.5",
]

Expand Down
26 changes: 26 additions & 0 deletions core/services/llm.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
import os
from functools import cache

import anthropic
from dotenv import load_dotenv

load_dotenv()
MODEL = "claude-sonnet-5"


@cache
def _client() -> anthropic.Anthropic:
# ANTHROPIC_API_KEY from env / .env; org-level keys also need the workspace header
workspace = os.getenv("ANTHROPIC_WORKSPACE_ID")
return anthropic.Anthropic(default_headers={"anthropic-workspace-id": workspace} if workspace else None)


def chat(
messages: list[dict],
tools: list[dict] | None = None,
tool_choice: dict | None = None,
max_tokens: int = 16000,
) -> anthropic.types.Message:
"""One Messages API call. Raises anthropic.APIError on transport or API failure."""
extra = {"tools": tools, "tool_choice": tool_choice} if tools else {}
return _client().messages.create(model=MODEL, max_tokens=max_tokens, messages=messages, **extra)
Loading
Loading