Skip to content

Latest commit

 

History

History
21 lines (14 loc) · 1.07 KB

File metadata and controls

21 lines (14 loc) · 1.07 KB

Local Terminal-Bench corpus audit

Generated by scripts/audit_terminal_bench.py. This proves aggregate read coverage of every local trial result and available signal stream; it does not claim every transcript was manually interpreted.

Coverage

  • Per-trial results read: 445
  • Run/corpus groups: 1
  • Signal streams read: 445 (100.00% coverage)
  • Aggregate Harbor result files excluded: 1
  • Unreadable trial results: 0

Groups

Corpus Run Model Trials Pass Partial Fail Errors Signals Parse errors Command timeouts
2026-07-15__18-08-50-submission 2026-07-15__18-08-50 gemini/gemini-3.1-pro-preview 445 296 0 147 2 445 77 112

Interpretation boundary

The generated JSON contains outcome, terminal-reason, protocol-error, timeout, and selected loop-event aggregates for every group. Semantic claims about mechanisms still require the paired gates, experiment log, and representative trajectory inspection; aggregate coverage alone cannot establish causality.