Skip to content

feat(eval): answer-level evaluation on a private corpus, with a first baseline (P2-3) - #410

Open
jwvanderstam wants to merge 5 commits into
mainfrom
feat/p2-3-answer-eval
Open

jwvanderstam wants to merge 5 commits into
mainfrom
feat/p2-3-answer-eval

Conversation

@jwvanderstam

@jwvanderstam jwvanderstam commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Implements ROADMAP P2-3, except the one acceptance item that needs a human: see Still open below.

What it adds

  • scripts/eval_answers.py:
    • draft: proposes question, reference answer and verbatim-proof cases from a corpus. It skips tables of contents and fill-in templates, and the drafting model can skip boilerplate.
    • run: answers through the chat's own path (get_rag_context → _build_context_prompt → OllamaClient), so the numbers describe the chat, not a reimplementation of it. Then a judge model grades correctness and faithfulness.
    • rejudge: re-grades existing answers under edited cases. A rejected case is dropped, and a changed question is flagged for a fresh run.
    • calibrate: measures judge agreement and records who graded.
    • check: enforces the committed baseline, and refuses a run measured with a different answer model, judge model or judge prompt.
  • scripts/eval_review.html: an offline, single-file viewer. Case review shows the proof highlighted in its passage; calibration grading is blind. It saves back to the JSON file in Edge or Chrome, and nothing leaves the machine.
  • Privacy: outputs that quote the corpus are refused inside the repository. The committed baseline carries only aggregates, models and caveats; I scanned it and the diff for names from the corpus.

Baseline (tests/eval/answer_baseline.json)

The corpus is a private customer RFP and contract set of 208 documents. There are 105 cases: drafted by mistral, reviewed and corrected by an AI assistant, and approved by the maintainer. ``llama3.2answers andmistral` judges, with prompt `v1`. Retrieval runs on the repository's default settings, which the baseline records.

source@1 source@5 MRR proof in context citation correct answer correct* faithful*
0.41 0.56 0.48 0.47 0.57 0.66 0.93

*Judge-scored and lenient: against an assistant's grades of 20 answers, the judge scored +0.21 on correct and +0.16 on faithful. Treat these two numbers as a regression tripwire, not a quality claim.

What it found

  • Retrieval recall is the weak link. The source document is never retrieved for 43% of questions; once retrieved it usually ranks first (41 of 57). The misses concentrate in office formats: not retrieved for Excel 76%, PowerPoint 67%, Word 52%, against PDF 20%. It isn't a threshold or pool-size effect (doubling TOP_K_RESULTS changes nothing), so structured office ingest is the next lever.
  • The first baseline's rank metrics were wrong, and are corrected here (4917e08):
    • Rank was read from position, but _rank_and_finalize returns documents alphabetically (a reading-order sort), so "source@1 = 0.14" measured the alphabet. It's now taken from relevance scores: 0.41.
    • A local app_state.json persisted retrieval overrides that beat the defaults and every env var. run now pins the defaults, records the settings, and check refuses a run whose settings differ.
  • The model also receives its context alphabetically, and truncation would cut by name. That's latent at current settings; a follow-up PR will fix it.
  • The judge's blind spot is a correct core with an invented specific: a clause number, a section or a template name. A stricter judge prompt, v2, catches those but over-corrects (−0.26 on faithful) and fails to parse 6 of 105 verdicts. Both prompts are kept, and check won't let them be confused.
  • eval_retrieval.ingest_corpus read only the top level of a folder, so a real corpus in subfolders ingested as 4 documents of 274, and its cases would have been scored against an empty database. It now recurses, with a regression test.

Still open

  • Human calibration. The 20 calibration answers were graded by an AI assistant, not a person. The baseline says so, and the ROADMAP keeps P2-3 at ◐. Grading them by hand in the viewer and re-running calibrate --score closes it.
  • DEL-2, the GraphRAG comparison on this corpus.
  • Nightly runs, which need a machine with a model.

Verification

  • ruff, mypy, bandit: clean.
  • Fast suite: 3209 passed, 0 failed.
  • 26 plus 10 unit tests cover the scoring logic.
  • Both scripts had been ignored by /scripts/* in .gitignore. They're now allow-listed; otherwise this PR would have shipped a baseline for a script that wasn't in it.

Merge note

It targets main's get_rag_context(workspace_id=...). After #406, run needs scope=ALL_WORKSPACES, and I'll adjust whichever lands second.

🤖 Generated with Claude Code

jwvanderstam and others added 4 commits September 30, 2026 12:24
… baseline (P2-3)

scripts/eval_answers.py drafts question/reference/verbatim-proof cases from a
document corpus, answers them through the chat's own path (get_rag_context,
_build_context_prompt, OllamaClient), has a local judge grade correctness and
faithfulness, re-grades under edited cases (rejudge), measures judge
agreement (calibrate) and enforces a committed baseline (check).
scripts/eval_review.html is an offline viewer for case review and blind
calibration grading. Corpus-quoting outputs are refused inside the repo.

Baseline: 105 cases on a 208-document customer RFP/contract corpus;
llama3.2 answers, mistral judges (prompt v1). Retrieval is the weak link:
source@1 0.14, source@5 0.51. The judge is lenient against an AI assistant's
grades (+0.21 correct, +0.16 faithful), so those two numbers are a
regression tripwire, not a quality claim; human calibration remains open.
check refuses to compare runs measured with a different model or judge.

Also: eval_retrieval.ingest_corpus read only the top level of a folder, so a
real corpus in subfolders ingested as 4 documents of 274. It now recurses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…de its agreement

calibrate --dims correct writes a sheet with no context: grading "correct"
needs only the reference answer and the assistant's, about ten seconds each,
so twenty answers is a few minutes of a human's time rather than an hour.
The viewer follows the sheet's dims (one keypress per answer). Scoring now
reports judge_bias - mean judge score minus mean human score - because
agreement says how often the judge matches, not which way it leans, and the
lean is what says how much a judged metric is overstated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hat is right about half the time

A second calibration sheet, graded on the corrected reference answers
(18 by an AI assistant, 2 by the maintainer), confirmed the first: judge v1
overstates 'correct' by +0.30 (first sheet +0.21), while v2 is near-unbiased
(-0.03) and scores the same answers 0.52. The baseline keeps v1 as a
regression tripwire and now says so beside the number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two errors in the first baseline's rank metrics, found by a settings sweep
that changed nothing:
- Rank was read from position, and retrieval returns documents
  alphabetically (_rank_and_finalize sorts by filename for reading order),
  so source@1 measured the alphabet: 0.14. By best chunk score it is 0.41.
- A local app_state.json persisted RAG overrides (TOP_K 40, RERANK_TOP_K 10,
  DIVERSITY 0.8) that beat the defaults and every env var. run now drops them
  in memory, records the settings in force, and check refuses a run whose
  settings differ from the baseline's.

Re-baselined on repository defaults: source@1 0.41, @5 0.56, MRR 0.48,
retrieved 0.57; answer_correct 0.66 (v1) / 0.51 (v2, near-unbiased).
Adds run --retrieval-only (41 s instead of an hour). The misses concentrate in
office formats: not retrieved for 76% of Excel, 67% PowerPoint, 52% Word
questions, 20% PDF.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 07:07

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Several issues currently invalidate evaluation comparability, judging context, and the draft-to-review workflow.

Review effort: Balanced
Findings: 1 High severity · 9 Medium severity

Open (10)
What changed in this PR

Adds private-corpus answer evaluation, calibration tooling, regression baselines, and recursive corpus ingestion.

Changes:

  • Adds answer drafting, execution, judging, calibration, and regression checks.
  • Adds an offline review interface and baseline metrics.
  • Fixes ingestion of documents in nested folders.
File Description
scripts/​eval_answers.py Implements the evaluation workflow.
scripts/​eval_review.html Adds the offline review and calibration UI.
scripts/​eval_retrieval.py Recursively discovers corpus documents.
tests/​unit/​test_eval_answers.py Tests evaluation scoring and calibration helpers.
tests/​unit/​test_eval_retrieval_corpus.py Tests nested corpus ingestion.
tests/​eval/​answer_baseline.json Records the initial aggregate baseline.
docs/​ROADMAP.md Documents P2-3 progress and findings.
CHANGELOG.md Announces the evaluation tooling.
.gitignore Allows the new scripts to be committed.
.claude/​rules/​file-map.md Catalogues the new files.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/eval_answers.py
Comment on lines +384 to +385
args.out.write_text(
header + yaml.safe_dump({"cases": cases}, allow_unicode=True, sort_keys=False, width=100),
Comment thread scripts/eval_answers.py
Comment on lines +76 to +90
def parse_json_object(raw: str) -> dict[str, Any] | None:
"""The first JSON object in a model reply, or None — models wrap JSON in prose."""
start = raw.find("{")
while start != -1:
depth = 0
for i in range(start, len(raw)):
depth += {"{": 1, "}": -1}.get(raw[i], 0)
if depth == 0:
try:
value = json.loads(raw[start:i + 1])
except json.JSONDecodeError:
break
return value if isinstance(value, dict) else None
start = raw.find("{", start + 1)
return None
Comment thread scripts/eval_answers.py
Comment on lines +178 to +181
return [
f"{key} {run_info.get(key)!r} vs baseline {baseline.get(key)!r}"
for key in ("answer_model", "judge_model", "judge_prompt", "retrieval")
if run_info.get(key) != baseline.get(key)
Comment thread scripts/eval_answers.py
Comment on lines +207 to +211
keys = ("TOP_K_RESULTS", "RERANK_TOP_K", "DIVERSITY_THRESHOLD", "SEMANTIC_WEIGHT")
settings = {k: config.app_state.get_rag_param(k) for k in keys}
settings.update({
k: getattr(config, k)
for k in ("RERANKER_ENABLED", "RERANKER_WEIGHT", "CHUNK_SIZE", "CHUNK_OVERLAP", "MAX_CONTEXT_LENGTH")
Comment thread scripts/eval_answers.py
Comment on lines +426 to +428
# What chat.get_rag_context does with MCP off, unrolled to keep each chunk's score.
retrieved = doc_processor.retrieve_context(case["question"])
context = doc_processor.format_context_for_llm(retrieved, max_length=config.MAX_CONTEXT_LENGTH)
Comment thread scripts/eval_answers.py
Comment on lines +429 to +433
scored = [
(r.filename, r.metadata["rerank_score"] if r.metadata.get("rerank_score") is not None
else r.metadata.get("combined_score", r.similarity))
for r in retrieved
]
Comment thread scripts/eval_answers.py
Comment on lines +437 to +438
# Exactly what the judge saw, so a human calibrating it can see the same.
"context": context[:6000] or "(no documents were retrieved)",
Comment thread scripts/eval_answers.py
Comment on lines +446 to +447
reply = loop.run_until_complete(client.generate_chat_completion(args.model, messages))
result["answer"] = reply["message"]["content"]
Comment thread scripts/eval_answers.py
Comment on lines +530 to +534
sheet = [{
"id": r["id"], "status": "pending", "question": r["question"],
"reference_answer": r["reference_answer"], "answer": r["answer"],
**({"context": r["context"]} if "faithful" in dims else {}),
**{f"human_{d}": "" for d in dims},
Comment thread scripts/eval_answers.py
r.add_argument("--ingest", action="store_true")
r.add_argument("--model", default="llama3.2")
r.add_argument("--judge-model", default="mistral")
r.add_argument("--judge-prompt", choices=sorted(JUDGE_PROMPTS), default="v2")
@sonarqubecloud

sonarqubecloud Bot commented Oct 1, 2026

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants