AI Agent Tool Reliability Lab. ToolTrust reviews tool definitions before an AI agent is allowed to use them. It focuses on schema clarity, risk, permissions, side effects, confirmation, idempotency, structured errors, and execution replay.
- deterministic risk classification: read-only, reversible write, irreversible write, external side effect
- JSON Schema quality checks for names, types, required fields, descriptions, enums, and argument ambiguity
- explicit permission and side-effect checks
- mandatory human confirmation policy for high-impact tools
- idempotency and structured-error checks
- five category scores + overall reliability score
- confirmation preview
- replay viewer with schema validation → permission check → confirmation → simulated execution
- polished dashboard, FastAPI API, tests, Docker, read-only CI
pip install -r requirements.txt
uvicorn app.main:app --reloadAn agent can only be as safe and predictable as the tools it is allowed to call. ToolTrust treats the tool interface itself as a reliability boundary.
“ToolTrust is about Agent Experience, not just prompts. I evaluate whether a tool is understandable, permissioned, recoverable, and safe to execute, then replay the decision path so a reviewer can see exactly where confirmation or validation stopped the action.”
All execution is simulated. ToolTrust never calls external systems. A production system would add signed tool registries, policy-as-code, identity-aware permissions, durable audit logs, sandbox execution, secret handling, versioned schemas, and model-specific evaluation suites.
The Verify workflow runs the test suite and Python compilation on Python 3.12 for pushes to main and pull requests. It has read-only repository permissions. Run the tests locally with python -m pytest.