Generated by
scripts/audit_terminal_bench.py. This proves aggregate read coverage of every local trial result and available signal stream; it does not claim every transcript was manually interpreted.
- Per-trial results read: 3571
- Run/corpus groups: 19
- Signal streams read: 3540 (99.13% coverage)
- Aggregate Harbor result files excluded: 18
- Unreadable trial results: 0
| Corpus | Run | Model | Trials | Pass | Partial | Fail | Errors | Signals | Parse errors | Command timeouts |
|---|---|---|---|---|---|---|---|---|---|---|
| jobs-cloud | 2026-06-09__16-53-01 | gemini/gemini-3.1-pro-preview | 445 | 300 | 0 | 136 | 9 | 445 | 27 | 154 |
| jobs-fullrun-2.1 | 2026-06-10__15-35-20 | gemini/gemini-3.1-pro-preview | 445 | 291 | 0 | 144 | 10 | 444 | 55 | 159 |
| jobs-gate | 2026-06-10__01-44-32 | gemini/gemini-3.1-pro-preview | 225 | 84 | 0 | 132 | 9 | 225 | 19 | 93 |
| jobs-gate | 2026-06-10__06-50-59 | gemini/gemini-3.1-pro-preview | 225 | 79 | 0 | 136 | 10 | 224 | 19 | 250 |
| jobs-gate | 2026-06-10__06-51-19 | gemini/gemini-3.1-pro-preview | 225 | 67 | 0 | 151 | 7 | 224 | 12 | 200 |
| jobs-gate | 2026-06-10__09-38-07 | gemini/gemini-3.1-pro-preview | 221 | 85 | 0 | 127 | 9 | 221 | 29 | 130 |
| jobs-gate | 2026-06-10__11-15-41 | gemini/gemini-3.1-pro-preview | 225 | 84 | 0 | 131 | 10 | 225 | 20 | 119 |
| jobs-gate | 2026-06-10__13-36-04 | gemini/gemini-3.1-pro-preview | 225 | 83 | 0 | 81 | 61 | 198 | 13 | 91 |
| jobs-rerun | 2026-06-09__22-01-02 | gemini/gemini-3.1-pro-preview | 163 | 57 | 0 | 98 | 8 | 162 | 17 | 67 |
| qwen-floor | 2026-06-11__16-36-32 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 100 | 18 | 0 | 80 | 2 | 100 | 0 | 32 |
| qwen-floor-board | 2026-06-12__07-56-37 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 90 | 24 | 0 | 63 | 3 | 90 | 5 | 18 |
| qwen-mutation-trial | 2026-06-12__02-23-37 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 8 | 3 | 0 | 5 | 0 | 8 | 0 | 0 |
| qwen-raise40 | 2026-06-11__19-12-16 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 99 | 17 | 0 | 79 | 3 | 99 | 1 | 24 |
| qwen-spec-first | 2026-06-11__23-54-32 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 99 | 25 | 0 | 69 | 5 | 99 | 1 | 22 |
| qwen-spec-first-board | 2026-06-12__05-38-49 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 194 | 46 | 0 | 142 | 6 | 194 | 2 | 62 |
| qwen-spec-first-s30 | 2026-06-11__23-07-24 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 39 | 5 | 0 | 34 | 0 | 39 | 0 | 4 |
| qwen-verify-repair | 2026-06-11__22-05-07 | vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maas | 99 | 18 | 0 | 80 | 1 | 99 | 1 | 27 |
| submission-2.0 | 2026-06-10__18-19-50 | gemini/gemini-3.1-pro-preview | 432 | 264 | 0 | 151 | 17 | 432 | 35 | 234 |
| submission-2.0 | 2026-06-10__21-03-33 | gemini/gemini-3.1-pro-preview | 12 | 5 | 0 | 7 | 0 | 12 | 1 | 4 |
The generated JSON contains outcome, terminal-reason, protocol-error, timeout, and selected loop-event aggregates for every group. Semantic claims about mechanisms still require the paired gates, experiment log, and representative trajectory inspection; aggregate coverage alone cannot establish causality.