Skip to content

vouchbench season 1 — submit ranking code, dethrone the champion, ship as the default #573

Description

@plind-junior

vouchbench season 1 is open. submit ranking code, get scored by a benchmark anyone can reproduce, and if you win: payout, changelog credit, and your code ships as the default in every vouch install.

start here: docs/mining-on-vouch.md — fork to shipped in five commands.

the reigning champion

vouch.strategies.provenance (promoted from #567): provenance-aware ranking — hearsay demoted, stored instructions buried, change-of-state phrasing boosted. contrib/strategies/baseline.py delegates to it; your submission is scored against it, paired on identical seeds.

the map of where the money is

champion composite: 0.575 ± 0.022 (seeds 1-12)

category mean read
single-session-recall 0.75 headroom
multi-session 0.67 headroom
knowledge-update 0.00 lever — lifecycle, not ranking
point-in-time 0.92 near-won
decoy-discrimination 0.25 lever — ranking can now exclude
injection-resistance 0.75 headroom
abstention 0.42 lever — ranking can now exclude
citation-correctness 1.00 guard: do not lose it
receipt-coverage 1.00 guard: do not lose it
supersede-hygiene 0.00 lever — lifecycle, not ranking

the zeros are the levers. reproduce this table yourself:

pip install -e '.[dev]'
vouch bench run --seeds 1,2,3,4,5,6,7,8,9,10,11,12 \
    --strategy contrib/strategies/baseline.py --json

how to enter

  • engine lane (this is the competition): one new file contrib/strategies/<you>.py against the koth-ladder branch. the gate scores it in a sandbox and posts the scorecard; engine code never auto-merges — the highest verified score merges after human review for benchmark-keyed logic (lookup tables, category-pattern dispatch, generator matching = disqualified).
  • kit lane (warm-up): competition/kits/current/kit.yaml only; auto-merges on a dethrone. its ceiling is low by design — learn the loop here.
  • practice locally with the exact ci loop before pushing:
    vouch bench run --seeds ... --strategy contrib/strategies/<you>.py --against contrib/strategies/baseline.py

season rules

  • daily throne is provisional (public seeds, derived from champion sha + utc date — printed in every scorecard).
  • the season closes with a sealed commit-reveal scored run on seeds that do not exist until the cutoff (docs/vouchbench-seasons.md).
  • cutoff: 2026-08-31 23:59 utc. rank shares 65 / 14 / 10 / 7 / 4, paid via a pr-native bounty platform or github sponsors within a week of close.
  • pool: announced by the maintainer in this thread before the first scored run.

ladder b — flat bounties (benchmark + engine hardening)

paid on merge after review, labeled bounty:$X:

  1. compile working-set selector (headline). vouch compile currently inlines every live claim into one prompt — it chokes past a few hundred claims. wire the ranking-strategy hook into compile's claim selection (rank the pool per planned topic, compile from the top-n under a budget). this makes a merged strategy improve capture, recall, and page compilation at once.
  2. supersession detection at capture — the knowledge-update and supersede-hygiene zeros are lifecycle levers: emit supersede proposals from detected value-updates (the review gate stays the write boundary).
  3. abstention hardening — cross-person leak under lexical match; ranking can now exclude, so there is headroom.
  4. bench categories + generators — new verifiability axes, better decoys, anti-overfit work.

every scorecard's inputs (seeds, commit, command) are public; every score can be recomputed offline. reproducibility is the trust model.

Metadata

Metadata

Assignees

No one assigned

    Labels

    season-1vouchbench season 1 entries and announcements

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions