feat(eval): answer-level evaluation on a private corpus, with a first baseline (P2-3) - #410
Open
jwvanderstam wants to merge 5 commits into
Open
jwvanderstam wants to merge 5 commits into
jwvanderstam wants to merge 5 commits into
Conversation
… baseline (P2-3) scripts/eval_answers.py drafts question/reference/verbatim-proof cases from a document corpus, answers them through the chat's own path (get_rag_context, _build_context_prompt, OllamaClient), has a local judge grade correctness and faithfulness, re-grades under edited cases (rejudge), measures judge agreement (calibrate) and enforces a committed baseline (check). scripts/eval_review.html is an offline viewer for case review and blind calibration grading. Corpus-quoting outputs are refused inside the repo. Baseline: 105 cases on a 208-document customer RFP/contract corpus; llama3.2 answers, mistral judges (prompt v1). Retrieval is the weak link: source@1 0.14, source@5 0.51. The judge is lenient against an AI assistant's grades (+0.21 correct, +0.16 faithful), so those two numbers are a regression tripwire, not a quality claim; human calibration remains open. check refuses to compare runs measured with a different model or judge. Also: eval_retrieval.ingest_corpus read only the top level of a folder, so a real corpus in subfolders ingested as 4 documents of 274. It now recurses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…de its agreement calibrate --dims correct writes a sheet with no context: grading "correct" needs only the reference answer and the assistant's, about ten seconds each, so twenty answers is a few minutes of a human's time rather than an hour. The viewer follows the sheet's dims (one keypress per answer). Scoring now reports judge_bias - mean judge score minus mean human score - because agreement says how often the judge matches, not which way it leans, and the lean is what says how much a judged metric is overstated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hat is right about half the time A second calibration sheet, graded on the corrected reference answers (18 by an AI assistant, 2 by the maintainer), confirmed the first: judge v1 overstates 'correct' by +0.30 (first sheet +0.21), while v2 is near-unbiased (-0.03) and scores the same answers 0.52. The baseline keeps v1 as a regression tripwire and now says so beside the number. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two errors in the first baseline's rank metrics, found by a settings sweep that changed nothing: - Rank was read from position, and retrieval returns documents alphabetically (_rank_and_finalize sorts by filename for reading order), so source@1 measured the alphabet: 0.14. By best chunk score it is 0.41. - A local app_state.json persisted RAG overrides (TOP_K 40, RERANK_TOP_K 10, DIVERSITY 0.8) that beat the defaults and every env var. run now drops them in memory, records the settings in force, and check refuses a run whose settings differ from the baseline's. Re-baselined on repository defaults: source@1 0.41, @5 0.56, MRR 0.48, retrieved 0.57; answer_correct 0.66 (v1) / 0.51 (v2, near-unbiased). Adds run --retrieval-only (41 s instead of an hour). The misses concentrate in office formats: not retrieved for 76% of Excel, 67% PowerPoint, 52% Word questions, 20% PDF. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Several issues currently invalidate evaluation comparability, judging context, and the draft-to-review workflow.
Review effort: Balanced
Findings: 1
Open (10)
Drafted review files are invalid JSON · New Brace counter rejects valid JSON strings · New Comparability ignores evaluated case set · New Retrieval fingerprint omits result-affecting settings · New Evaluation bypasses the configured MCP retrieval path · New Retrieval metrics use inconsistent ranking scores · New Judge and calibration receive truncated context · New Evaluation omits production temperature setting · New Missing grader metadata is misclassified as human · New Default run and baseline use mismatched judge prompts · New
What changed in this PR
Adds private-corpus answer evaluation, calibration tooling, regression baselines, and recursive corpus ingestion.
Changes:
- Adds answer drafting, execution, judging, calibration, and regression checks.
- Adds an offline review interface and baseline metrics.
- Fixes ingestion of documents in nested folders.
| File | Description |
|---|---|
scripts/eval_answers.py |
Implements the evaluation workflow. |
scripts/eval_review.html |
Adds the offline review and calibration UI. |
scripts/eval_retrieval.py |
Recursively discovers corpus documents. |
tests/unit/test_eval_answers.py |
Tests evaluation scoring and calibration helpers. |
tests/unit/test_eval_retrieval_corpus.py |
Tests nested corpus ingestion. |
tests/eval/answer_baseline.json |
Records the initial aggregate baseline. |
docs/ROADMAP.md |
Documents P2-3 progress and findings. |
CHANGELOG.md |
Announces the evaluation tooling. |
.gitignore |
Allows the new scripts to be committed. |
.claude/rules/file-map.md |
Catalogues the new files. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+384
to
+385
| args.out.write_text( | ||
| header + yaml.safe_dump({"cases": cases}, allow_unicode=True, sort_keys=False, width=100), |
Comment on lines
+76
to
+90
| def parse_json_object(raw: str) -> dict[str, Any] | None: | ||
| """The first JSON object in a model reply, or None — models wrap JSON in prose.""" | ||
| start = raw.find("{") | ||
| while start != -1: | ||
| depth = 0 | ||
| for i in range(start, len(raw)): | ||
| depth += {"{": 1, "}": -1}.get(raw[i], 0) | ||
| if depth == 0: | ||
| try: | ||
| value = json.loads(raw[start:i + 1]) | ||
| except json.JSONDecodeError: | ||
| break | ||
| return value if isinstance(value, dict) else None | ||
| start = raw.find("{", start + 1) | ||
| return None |
Comment on lines
+178
to
+181
| return [ | ||
| f"{key} {run_info.get(key)!r} vs baseline {baseline.get(key)!r}" | ||
| for key in ("answer_model", "judge_model", "judge_prompt", "retrieval") | ||
| if run_info.get(key) != baseline.get(key) |
Comment on lines
+207
to
+211
| keys = ("TOP_K_RESULTS", "RERANK_TOP_K", "DIVERSITY_THRESHOLD", "SEMANTIC_WEIGHT") | ||
| settings = {k: config.app_state.get_rag_param(k) for k in keys} | ||
| settings.update({ | ||
| k: getattr(config, k) | ||
| for k in ("RERANKER_ENABLED", "RERANKER_WEIGHT", "CHUNK_SIZE", "CHUNK_OVERLAP", "MAX_CONTEXT_LENGTH") |
Comment on lines
+426
to
+428
| # What chat.get_rag_context does with MCP off, unrolled to keep each chunk's score. | ||
| retrieved = doc_processor.retrieve_context(case["question"]) | ||
| context = doc_processor.format_context_for_llm(retrieved, max_length=config.MAX_CONTEXT_LENGTH) |
Comment on lines
+429
to
+433
| scored = [ | ||
| (r.filename, r.metadata["rerank_score"] if r.metadata.get("rerank_score") is not None | ||
| else r.metadata.get("combined_score", r.similarity)) | ||
| for r in retrieved | ||
| ] |
Comment on lines
+437
to
+438
| # Exactly what the judge saw, so a human calibrating it can see the same. | ||
| "context": context[:6000] or "(no documents were retrieved)", |
Comment on lines
+446
to
+447
| reply = loop.run_until_complete(client.generate_chat_completion(args.model, messages)) | ||
| result["answer"] = reply["message"]["content"] |
Comment on lines
+530
to
+534
| sheet = [{ | ||
| "id": r["id"], "status": "pending", "question": r["question"], | ||
| "reference_answer": r["reference_answer"], "answer": r["answer"], | ||
| **({"context": r["context"]} if "faithful" in dims else {}), | ||
| **{f"human_{d}": "" for d in dims}, |
| r.add_argument("--ingest", action="store_true") | ||
| r.add_argument("--model", default="llama3.2") | ||
| r.add_argument("--judge-model", default="mistral") | ||
| r.add_argument("--judge-prompt", choices=sorted(JUDGE_PROMPTS), default="v2") |
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.





Implements ROADMAP P2-3, except the one acceptance item that needs a human: see Still open below.
What it adds
scripts/eval_answers.py:draft: proposes question, reference answer and verbatim-proof cases from a corpus. It skips tables of contents and fill-in templates, and the drafting model can skip boilerplate.run: answers through the chat's own path (get_rag_context→_build_context_prompt→OllamaClient), so the numbers describe the chat, not a reimplementation of it. Then a judge model grades correctness and faithfulness.rejudge: re-grades existing answers under edited cases. A rejected case is dropped, and a changed question is flagged for a freshrun.calibrate: measures judge agreement and records who graded.check: enforces the committed baseline, and refuses a run measured with a different answer model, judge model or judge prompt.scripts/eval_review.html: an offline, single-file viewer. Case review shows the proof highlighted in its passage; calibration grading is blind. It saves back to the JSON file in Edge or Chrome, and nothing leaves the machine.Baseline (
tests/eval/answer_baseline.json)The corpus is a private customer RFP and contract set of 208 documents. There are 105 cases: drafted by
mistral, reviewed and corrected by an AI assistant, and approved by the maintainer. ``llama3.2answers andmistral` judges, with prompt `v1`. Retrieval runs on the repository's default settings, which the baseline records.*Judge-scored and lenient: against an assistant's grades of 20 answers, the judge scored +0.21 on correct and +0.16 on faithful. Treat these two numbers as a regression tripwire, not a quality claim.
What it found
TOP_K_RESULTSchanges nothing), so structured office ingest is the next lever.4917e08):_rank_and_finalizereturns documents alphabetically (a reading-order sort), so "source@1 = 0.14" measured the alphabet. It's now taken from relevance scores: 0.41.app_state.jsonpersisted retrieval overrides that beat the defaults and every env var.runnow pins the defaults, records the settings, andcheckrefuses a run whose settings differ.v2, catches those but over-corrects (−0.26 on faithful) and fails to parse 6 of 105 verdicts. Both prompts are kept, andcheckwon't let them be confused.eval_retrieval.ingest_corpusread only the top level of a folder, so a real corpus in subfolders ingested as 4 documents of 274, and its cases would have been scored against an empty database. It now recurses, with a regression test.Still open
calibrate --scorecloses it.Verification
ruff,mypy,bandit: clean./scripts/*in.gitignore. They're now allow-listed; otherwise this PR would have shipped a baseline for a script that wasn't in it.Merge note
It targets
main'sget_rag_context(workspace_id=...). After #406,runneedsscope=ALL_WORKSPACES, and I'll adjust whichever lands second.🤖 Generated with Claude Code