Skip to content

[FIX] [BREAKING]: Evaluate attack verdicts over final traces - #150

Draft
Spencer Schoenberg (spencrr) wants to merge 12 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-xpia-stopping
Draft

Spencer Schoenberg (spencrr) wants to merge 12 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-xpia-stopping

Conversation

@spencrr

@spencrr Spencer Schoenberg (spencrr) commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

XPIA now evaluates the attack objective once over the final trace, separately from online stopping. stop_when accepts an evaluator, None, or the default StopWhen.AUTO. StopWhen.AUTO reuses the verdict evaluator as the stop condition only when detection stays true as turns are appended: ToolCalled, SideEffectOccurred, and ResponseContains with ResponseScope.ANY_TURN. StopWhen is exported from rampart.attacks and rampart.

When the stop and verdict evaluators are the same object, the latest online evaluation is reused if it already covers the final trace, so no duplicate judge call is made. Final verdict evidence is stored on Result.final_trace_evaluation, trace completion uses TraceEndReason, and online stop evidence uses EvaluationPurpose.STOP_CHECK. Final evaluation runs inside the active session and injection stack. Observability downgrades are recorded without mutating response metadata, and cleanup failures discard otherwise successful verdict evidence and return ERROR.

The list-based resolve_as_attack and per-turn evaluate_turn_async helpers are removed because no built-in strategy uses them. The trace-contract declaration records a compatible change against main.

Depends on #149.

Breaking changes

  • XPIA verdicts are computed once over the final trace. Single-trigger attacks with deterministic evaluators keep their verdicts.
  • With the default StopWhen.AUTO, other evaluators, including LLM judges, no longer stop early. They are called once on the final trace, and adaptive drivers can run up to max_turns. Pass the same evaluator as stop_when to restore per-turn early stopping without a duplicate final call.
  • With the default, stochastic evaluators are sampled once per run instead of once per turn, so trial pass rates can shift.
  • Auto-stopped attacks expose online evaluator feedback to adaptive drivers.
  • Removed helpers: replace resolve_as_attack(eval_results=...) with resolve_attack_verdict(evaluation=...), and replace evaluate_turn_async with run_trace_async plus evaluate_final_trace_async.

The upgrade note in docs/attacks/xpia.md covers these changes.

Checklist

  • pre-commit run --all-files passes
  • Tests added or updated for changes — automatic, explicit, and disabled stopping; StopWhen values; exact evaluator call counts; stable-evaluator classification and composition; cleanup ordering; zero and max turns; observability; metadata isolation; and summaries
  • Documentation updated

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch from d80c55b to 84d0197 Compare August 8, 2026 02:31
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 2 times, most recently from cdcad35 to bedbf6b Compare August 27, 2026 17:30
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 2 times, most recently from 3051684 to 4074ae9 Compare September 8, 2026 22:46
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch 3 times, most recently from bb6279f to 3ee9e25 Compare September 29, 2026 22:20
Require the same observability and manifest before reusing an online judgment. Copy optional evidence through the shared tolerant renderer so malformed supporting text cannot discard an established verdict. Keep terminal evidence and operand lists independent from online records.
Rename evaluate_terminal_async to evaluate_final_trace_async so the public runner helper matches Result.final_trace_evaluation. Document the trace execution helpers where they are introduced.
Retire the list reducer after probe execution adopts terminal evaluation. Keep explicit response scopes, online evidence separation, zero-turn errors and xdist v3 semantics consistent across tests, exports and extension guidance.
Call evaluate_final_trace_async and describe probe verdict evidence as final-trace evaluation. Record a compatible trace-contract decision against the merged v2 base; Result fields and schemas are unchanged.
Explain the per-turn to final-trace verdict change, its single-prompt blast radius, trial sampling impact, and resolver replacement. Remove private xdist envelope details already covered by schema-drift rejection and the mixed-version limitation.
Remove the old reducers and helper now that both built-in strategies evaluate terminal traces. Keep explicit-scope stopping classification, provenance and extension documentation aligned with the shared runner.
Call evaluate_final_trace_async and describe XPIA verdict evidence as final-trace evaluation in code, tests, and extension guidance.
Explain the per-turn to final-trace verdict change, its single-trigger blast radius, automatic stopping defaults, LLM-judge cost trade-off, trial sampling impact, and replacements for removed helpers.
Type Attacks.xpia(stop_when=...) as Evaluator | StopWhen | None and default to StopWhen.AUTO, following the enum-over-Literal standard. StopWhen is a string enum, so its equal string value remains accepted at runtime. Export it from rampart.attacks and the top-level package.
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-xpia-stopping branch from 3ee9e25 to bcc3fd6 Compare October 1, 2026 00:59
@spencrr Spencer Schoenberg (spencrr) changed the title [FIX]: Evaluate attack verdicts over final traces [FIX] [BREAKING]: Evaluate attack verdicts over final traces Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant