Skip to content

v0.2.2: add honest paired benchmark framework - #3

Merged
ctdaniel merged 14 commits into
mainfrom
v0.2.2-benchmark-framework
Sep 21, 2026
Merged

ctdaniel merged 14 commits into
mainfrom
v0.2.2-benchmark-framework

Conversation

@ctdaniel

Copy link
Copy Markdown
Owner

Summary

Adds the v0.2.2 benchmark framework so CQO can be evaluated with reproducible, acceptance-gated evidence instead of guessed token-savings claims.

Benchmark design

  • Paired Baseline vs CQO runs
  • Shared case_id + pair_id
  • Same repository commit / prompt / acceptance criteria
  • Two modes:
    • controlled: same starting model/reasoning
    • full-policy: CQO may use its own routing policy
  • Efficiency deltas are only calculated when both runs pass acceptance

Observable metrics

  • files inspected / changed
  • searches
  • focused / broad checks
  • subagents
  • model escalations
  • repeated reads
  • wall time

Unknown values are omitted rather than guessed.

Added

  • benchmarks/README.md
  • Chinese benchmark methodology
  • result JSON schema
  • case-study template
  • benchmarks/report.py with text / Markdown / JSON output
  • benchmark tests and CI smoke coverage

Principle

Ship the measurement method before publishing savings claims.

Version bumped to 0.2.2.

@ctdaniel
ctdaniel merged commit a40185d into main Sep 21, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant