Skip to content

docs(adr): ADR-6: generation path is a deterministic pipeline; orchestration is eval-gated - #412

Merged
jwvanderstam merged 4 commits into
mainfrom
docs/adr-5-generation-path
Oct 2, 2026
Merged

jwvanderstam merged 4 commits into
mainfrom
docs/adr-5-generation-path

Conversation

@jwvanderstam

Copy link
Copy Markdown
Owner

What

Adds ADR-5 to docs/ADR.md and its row in .claude/rules/file-map.md. Documentation only; no code changes.

Why

The code carries two answers to who decides what gets retrieved. The vocabulary says agent; the wiring says pipeline (planner tools never read, multi-hop sub-questions retrieved in parallel, api_routes.py:346 disabling tool calling whenever retrieval returned anything). ADR-5 records the pipeline as the product and licenses a bounded re-query as an experiment, gated on P2-3.

Both alternatives are argued at full strength in the record and rejected on the same test: each acts before the evidence exists (the DEL-2 rule).

Owner decisions before merging

  • Status line: it says Proposed 2026-10-01; change to Accepted on merge, matching ADR-1..4.
  • The 2026-12-31 date in the second revisit trigger is a proposal. Move it if P2-3's place in the sprint order says otherwise, but keep a date.
  • Em dashes: ADR-5 deliberately uses none, unlike ADR-1..4. Headline format is ADR-5: rather than ADR-5 —. Normalise either way.

Found while writing it

get_rag_context_multi_hop (src/services/chat.py) takes neither source_ids nor additional_workspace_ids. On the default path (planner on, RAG on, query of 7 or more words, plan judged multi-hop) a source-narrowed chat searches the whole workspace, and extra workspaces are silently dropped. Workspace isolation holds. Listed first among the ADR's consequences; deserves its own fix PR rather than waiting on this one.

🤖 Generated with Claude Code

https://claude.ai/code/session_018CuUHpys6UBronLzBdtAkS


Generated by Claude Code

…tration is eval-gated

The code carries two answers to "who decides what gets retrieved". The vocabulary says
agent (a "ReAct-style" aggregator, a planner that emits tools and hops, a multi-round
tool loop); the wiring says pipeline (the planner's tools field is never read, multi-hop
sub-questions run in parallel, and api_routes.py:346 disables tool calling whenever
retrieval returned anything). ADR-1 ended the same kind of split in the deployment layer.

ADR-5 records the pipeline as the product and licenses a bounded re-query as an
experiment, off by default, built after P2-3 and shipped only if it beats the P2-3
baseline on multi-hop questions by a margin fixed before the run. Both alternatives were
argued at full strength and rejected on the same test: each acts before the evidence
exists. Deleting the planner and aggregator now would break the DEL-2 rule (defer, do not
delete, when the measurement cannot yet be made); building a harness now has no evidence
of lift, runs on the models ADR-4 leaves us, and its strongest stated motive is the one
ADR-1's revisit clause excludes.

A second revisit trigger (2026-12-31 if P2-3 has not started) keeps "let the eval
decide" from becoming indefinite deferral.

Found while writing it and listed first among the consequences: get_rag_context_multi_hop
drops source_ids and additional_workspace_ids, so on the default path a source-narrowed
chat asking a multi-hop question searches the whole workspace. Workspace isolation holds.
No code changes in this commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CuUHpys6UBronLzBdtAkS
Copilot AI balanced review requested due to automatic review settings October 1, 2026 06:12

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

ADR-5 contains a stale evaluation result and contradicts the default model-driven retrieval fallback.

Review effort: Balanced
Findings: 2 Low severity

Open (2)
What changed in this PR

Adds ADR-5 to establish deterministic retrieval as the product architecture and gate agentic orchestration on answer-level evaluation.

Changes:

  • Documents the decision, alternatives, consequences, and revisit criteria.
  • Adds ADR-5 to the repository file map.
File Description
docs/​ADR.md Adds ADR-5.
.claude/​rules/​file-map.md Registers ADR-5 in the documentation index.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docs/ADR.md
Comment on lines +288 to +289
exists, and the one comparable measurement, DEL-2's GraphRAG run, saw clever retrieval
fire on 3 of 20 questions. ADR-4 confines inference to local or EU OpenAI-compatible
Comment thread docs/ADR.md
Comment on lines +311 to +314
**What this rules out:**
- The model choosing tools or retrieval targets on the default chat path.
- Unbounded tool rounds. Any experimental loop has a hard round cap and a per-request token
budget, both enforced in code rather than by `TOOL_MAX_ROUNDS` alone.
Copilot AI balanced review requested due to automatic review settings October 1, 2026 07:07

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The ADR conflicts with the existing model-directed retrieval fallback and remains marked Proposed.

Review effort: Balanced
Findings: 3 Low severity

Open (3)

Comment thread docs/ADR.md

## ADR-5: The generation path is a deterministic pipeline; agentic orchestration is an experiment until an answer-level eval says otherwise

**Proposed 2026-10-01.** Supersedes nothing. Records a choice the code has been making
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 08:15
@jwvanderstam jwvanderstam changed the title docs(adr): ADR-5: generation path is a deterministic pipeline; orchestration is eval-gated docs(adr): ADR-6: generation path is a deterministic pipeline; orchestration is eval-gated Oct 2, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The ADR numbering conflicts with the PR description, and its fixed-pipeline decision contradicts the retained model-directed tool fallback.

Review effort: Balanced
Findings: 4 Low severity

Open (4)

Comment thread docs/ADR.md

---

## ADR-6: The generation path is a deterministic pipeline; agentic orchestration is an experiment until an answer-level eval says otherwise
ADR-5 (row-level security, #404) sits before ADR-6; the file-map row lists both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 11:51

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@sonarqubecloud

sonarqubecloud Bot commented Oct 2, 2026

Copy link
Copy Markdown

@jwvanderstam
jwvanderstam merged commit 776b863 into main Oct 2, 2026
15 checks passed
@jwvanderstam
jwvanderstam deleted the docs/adr-5-generation-path branch October 2, 2026 12:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants