Skip to content

Record the post-#906 typed-graph bench: two groups, thresholds read, tokens by phase - #918

Merged
WaylandYang merged 1 commit into
devfrom
docs/bench-fourth-run
Sep 25, 2026
Merged

WaylandYang merged 1 commit into
devfrom
docs/bench-fourth-run

Conversation

@WaylandYang

Copy link
Copy Markdown
Contributor

Two full groups (100 documents, judge 200, errata) on dev at b9e3aae, with #906 and #916, same corpus, model and judge as the pre-#906 baseline recorded in #917. Written into the bench README as the formal cut-2 measurement.

  • Judged precision 81.0% / 79.0% before errata, 83.0% / 80.5% after: over the 75.3% threshold, about ten points under the baseline, and the difference is misworded facts from minimal reasoning on extraction.
  • Same-sentence gold recall 14.2% / 14.3% before errata, 17.8% / 17.1% after: under the prototype's 15.3% before errata, over it after, and errata now adds without retracting wrongly (2 and 1 retractions, none judged stated).
  • 31k and 34k tokens per document: a fifth of the baseline's 158k / 183k, still five times 0044's ceiling; seven tenths of it in alignment and rule proposals, which are paid per signature and are at their most expensive on a fresh 100-document base. The README says the bench still lacks the marginal-tokens-on-a-warm-base reading that would measure the product's per-document cost.
  • 47 and 48 minutes a group, against three to five hours.

🤖 Generated with Claude Code

…tokens by phase

Same corpus, model and judge as the pre-#906 baseline, on dev with #906
and #916. Judged precision 81.0% / 79.0% before errata and 83.0% / 80.5%
after (the loss against the baseline is misworded facts from minimal
reasoning); same-sentence gold recall 14.2% / 14.3% before errata and
17.8% / 17.1% after, with errata now adding without retracting wrongly
(2 and 1 retractions, none judged stated); 31k and 34k tokens per
document, a fifth of the baseline and still five times 0044's ceiling,
seven tenths of it in alignment and rule proposals. 47 and 48 minutes a
group against three to five hours.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Wayland Yang <wayland0916@gmail.com>
@WaylandYang
WaylandYang merged commit e47c9b0 into dev Sep 25, 2026
7 checks passed
@WaylandYang
WaylandYang deleted the docs/bench-fourth-run branch September 25, 2026 13:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant