A task-based evaluation suite for canvas-editing AI agents — passing ≠ getting it right
-
Updated
Jul 4, 2026 - Python
A task-based evaluation suite for canvas-editing AI agents — passing ≠ getting it right
Open-source agent assurance: turn a PRD and agent URL into frozen test suites with evidence-backed PASS/FAIL/UNVERIFIABLE verdicts. Web console, REST API, SQLite. Domain packs for regulated teams (fintech, insurance, health).
To associate your repository with the agent-evaluation-eval topic, visit your repo's landing page and select "manage topics."