Summary
PyRIT already supports exposing tools to a target via Tool / ToolProvider
(see openai_response_target.py), but I don't see an attack strategy that
specifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.
This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.
Proposed approach (open to feedback before implementing)
- A new attack, e.g.
ExcessiveAgencyAttack, that:
- Takes a target configured with a defined set of "allowed" tools/scope
(via the existing Tool/ToolProvider system).
- Attempts to elicit a tool call outside that scope — either a tool the
agent has access to but shouldn't use for the stated task, or a chained
sequence of legitimate calls that together exceed the intended
permission boundary.
- Scores success based on the tool call actually made (inspecting the
target's tool-call output), not on the text response — this is the
distinct piece existing text/output scorers don't cover.
- Could reuse the single-turn or multi-turn structure depending on whether
the elicitation needs conversation history (multi-turn is likely more
realistic here, similar to how Crescendo works for text escalation).
Questions for maintainers
- Is there existing tooling for asserting on tool-call output specifically
(as opposed to text output) that I should reuse for scoring?
- Any prior discussion/PR on agentic risk testing I should be aware of
before designing this?
Happy to take this on once the approach is validated.
Summary
PyRIT already supports exposing tools to a target via
Tool/ToolProvider(see
openai_response_target.py), but I don't see an attack strategy thatspecifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.
This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.
Proposed approach (open to feedback before implementing)
ExcessiveAgencyAttack, that:(via the existing
Tool/ToolProvidersystem).agent has access to but shouldn't use for the stated task, or a chained
sequence of legitimate calls that together exceed the intended
permission boundary.
target's tool-call output), not on the text response — this is the
distinct piece existing text/output scorers don't cover.
the elicitation needs conversation history (multi-turn is likely more
realistic here, similar to how Crescendo works for text escalation).
Questions for maintainers
(as opposed to text output) that I should reuse for scoring?
before designing this?
Happy to take this on once the approach is validated.