Generated by
scripts/audit_terminal_bench.py. This proves aggregate read coverage of every local trial result and available signal stream; it does not claim every transcript was manually interpreted.
- Per-trial results read: 445
- Run/corpus groups: 1
- Signal streams read: 445 (100.00% coverage)
- Aggregate Harbor result files excluded: 1
- Unreadable trial results: 0
| Corpus | Run | Model | Trials | Pass | Partial | Fail | Errors | Signals | Parse errors | Command timeouts |
|---|---|---|---|---|---|---|---|---|---|---|
| 2026-07-15__18-08-50-submission | 2026-07-15__18-08-50 | gemini/gemini-3.1-pro-preview | 445 | 296 | 0 | 147 | 2 | 445 | 77 | 112 |
The generated JSON contains outcome, terminal-reason, protocol-error, timeout, and selected loop-event aggregates for every group. Semantic claims about mechanisms still require the paired gates, experiment log, and representative trajectory inspection; aggregate coverage alone cannot establish causality.