Evidence-aware benchmark and PR advisory prototype for AI coding-agent context safety.
benchmark pull-request software-engineering github-actions llm-evaluation ai-coding coding-agent agent-safety evalops pr-advisory
-
Updated
Jul 13, 2026 - Python