[FIX] [BREAKING]: Evaluate attack verdicts over final traces - #150
Draft
Spencer Schoenberg (spencrr) wants to merge 12 commits into
Draft
Spencer Schoenberg (spencrr) wants to merge 12 commits into
Spencer Schoenberg (spencrr) wants to merge 12 commits into
Conversation
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-xpia-stopping
branch
from
August 8, 2026 02:31
d80c55b to
84d0197
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-xpia-stopping
branch
2 times, most recently
from
August 27, 2026 17:30
cdcad35 to
bedbf6b
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-xpia-stopping
branch
2 times, most recently
from
September 8, 2026 22:46
3051684 to
4074ae9
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-xpia-stopping
branch
3 times, most recently
from
September 29, 2026 22:20
bb6279f to
3ee9e25
Compare
Require the same observability and manifest before reusing an online judgment. Copy optional evidence through the shared tolerant renderer so malformed supporting text cannot discard an established verdict. Keep terminal evidence and operand lists independent from online records.
Rename evaluate_terminal_async to evaluate_final_trace_async so the public runner helper matches Result.final_trace_evaluation. Document the trace execution helpers where they are introduced.
Retire the list reducer after probe execution adopts terminal evaluation. Keep explicit response scopes, online evidence separation, zero-turn errors and xdist v3 semantics consistent across tests, exports and extension guidance.
Call evaluate_final_trace_async and describe probe verdict evidence as final-trace evaluation. Record a compatible trace-contract decision against the merged v2 base; Result fields and schemas are unchanged.
Explain the per-turn to final-trace verdict change, its single-prompt blast radius, trial sampling impact, and resolver replacement. Remove private xdist envelope details already covered by schema-drift rejection and the mixed-version limitation.
Remove the old reducers and helper now that both built-in strategies evaluate terminal traces. Keep explicit-scope stopping classification, provenance and extension documentation aligned with the shared runner.
Call evaluate_final_trace_async and describe XPIA verdict evidence as final-trace evaluation in code, tests, and extension guidance.
Explain the per-turn to final-trace verdict change, its single-trigger blast radius, automatic stopping defaults, LLM-judge cost trade-off, trial sampling impact, and replacements for removed helpers.
Type Attacks.xpia(stop_when=...) as Evaluator | StopWhen | None and default to StopWhen.AUTO, following the enum-over-Literal standard. StopWhen is a string enum, so its equal string value remains accepted at runtime. Export it from rampart.attacks and the top-level package.
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-xpia-stopping
branch
from
October 1, 2026 00:59
3ee9e25 to
bcc3fd6
Compare
3 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
XPIA now evaluates the attack objective once over the final trace, separately from online stopping.
stop_whenaccepts an evaluator,None, or the defaultStopWhen.AUTO.StopWhen.AUTOreuses the verdict evaluator as the stop condition only when detection stays true as turns are appended:ToolCalled,SideEffectOccurred, andResponseContainswithResponseScope.ANY_TURN.StopWhenis exported fromrampart.attacksandrampart.When the stop and verdict evaluators are the same object, the latest online evaluation is reused if it already covers the final trace, so no duplicate judge call is made. Final verdict evidence is stored on
Result.final_trace_evaluation, trace completion usesTraceEndReason, and online stop evidence usesEvaluationPurpose.STOP_CHECK. Final evaluation runs inside the active session and injection stack. Observability downgrades are recorded without mutating response metadata, and cleanup failures discard otherwise successful verdict evidence and returnERROR.The list-based
resolve_as_attackand per-turnevaluate_turn_asynchelpers are removed because no built-in strategy uses them. The trace-contract declaration records a compatible change againstmain.Depends on #149.
Breaking changes
StopWhen.AUTO, other evaluators, including LLM judges, no longer stop early. They are called once on the final trace, and adaptive drivers can run up tomax_turns. Pass the same evaluator asstop_whento restore per-turn early stopping without a duplicate final call.resolve_as_attack(eval_results=...)withresolve_attack_verdict(evaluation=...), and replaceevaluate_turn_asyncwithrun_trace_asyncplusevaluate_final_trace_async.The upgrade note in
docs/attacks/xpia.mdcovers these changes.Checklist
pre-commit run --all-filespassesStopWhenvalues; exact evaluator call counts; stable-evaluator classification and composition; cleanup ordering; zero and max turns; observability; metadata isolation; and summaries