Skip to content

feat(search): Weigh entry age against depth in transposition table replacement - #375

Merged
aywrite merged 1 commit into
masterfrom
claude/tt-aging-spec-mosqqn
Oct 5, 2026
Merged

aywrite merged 1 commit into
masterfrom
claude/tt-aging-spec-mosqqn

Conversation

@aywrite

@aywrite aywrite commented Oct 5, 2026

Copy link
Copy Markdown
Owner

The transposition table's replacement policy now values an entry at its depth less eight plies for each search since it was stored. A full bucket gives up the entry worth least rather than the shallowest, and a store is refused only by an entry worth more than its depth. Before this, age counted only at the twelve search cliff, so an entry from an earlier move held its slot against a shallower one from the current search until then.

What a reviewer should know

  • Within one search every entry is worth its depth, so the policy is the old one there. arche bench counts 5,965,973 on the base and the branch. bench games runs +0.19% instructions with identical node counts.
  • Two behaviours change beyond the victim choice, and both have tests. A position's own entry from an earlier search gives way to a shallower store of the same position once the store reaches its worth. The clause that protects an exact entry of equal depth now compares worth, so an older exact entry no longer turns away a fresh bound of the same depth.
  • Played in two stages. At 10+0.1 with a 16 MB table, where the table is full and the policy decides most stores: +19 ±13 over 1,500 games, sprt [0, 10] passed. At 30+0.3 with the default 256 MB: +4 ±9 over 2,500 games, sprt [-10, 0] passed, after one extension past a 2,000 game cap that the test reached at LLR 2.93 against 2.94.
  • In the 30+0.3 games the candidate counted 0.968 to 0.985 of the baseline's nodes a second. That is the tree, not the code. Replayed on the same positions at fixed nodes with four processes at once, the two builds run level (4.24M against 4.25M). In the games the gap sits past ply 120, where old entries answer repeated endgame positions almost for nothing and the candidate searches two plies deeper in the same time.
  • Rebased onto d9379f3. The one conflict was a comment in a block that master had rewritten; the comments now say "replacement contest" where they said "depth contest". Release (780) and debug (782) tests, clippy with and without machine-test, and the tactical and strategic suites pass.

The workflow change that let the 16 MB match be dispatched is a separate pull request.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WmzY2g6SZGYszT5dnuQ6tw


Generated by Claude Code

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Speed against d9379f35

Measured on AMD EPYC 9V74 80-Core Processor.
Both sides built with rustc 1.98.1 (48a229cea 2026-09-01).

Both sides built and run on this runner in this job, the way
scripts/speed.sh measures a perf commit. Each round runs
both sides on a layout of its own, the same compiled code with
its code and data shuffled and moved, so the 95% interval carries
where the code landed as well as the run. The verdict holds the
interval against a 1% threshold. The default layout, the one a
release ships, is measured after as a diagnostic. The node
counts are the search's: they move when the search does, and a
speed change leaves them alone.

round     base nps  candidate nps  change
    1      5049456        5104733   +1.1%
    2      5052898        5119033   +1.3%
    3      5217807        5358664   +2.7%
    4      5373607        5494060   +2.2%
    5      5481627        5501229   +0.4%
    6      5604367        5734014   +2.3%
    7      5722053        5189485   -9.3%
    8      5340546        5419847   +1.5%
    9      5554004        5655256   +1.8%
   10      5721899        5518944   -3.5%
   11      5427751        5513450   +1.6%
   12      5381523        5330392   -1.0%
   13      5267567        5554625   +5.4%
   14      5512391        5479059   -0.6%
   15      5476600        5387739   -1.6%
   16      5381329        5251870   -2.4%
   17      5413012        5514480   +1.9%
   18      5495421        5473987   -0.4%
   19      5609210        5386960   -4.0%
   20      5381096        5446208   +1.2%
   21      5365966        5727464   +6.7%
   22      5715453        5743553   +0.5%
   23      5708917        5369961   -5.9%
   24      5493494        5243299   -4.6%
   25      5481516        5536197   +1.0%
   26      5547100        5368274   -3.2%
   27      5830199        5842720   +0.2%
   28      5572051        5796568   +4.0%
   29      5684163        5502903   -3.2%
   30      5617444        5601011   -0.3%
   31      5456749        5478591   +0.4%
   32      5795554        5806186   +0.2%
   33      5771089        5511907   -4.5%
   34      5683573        5529378   -2.7%
   35      5723826        5604157   -2.1%
   36      5580035        5318162   -4.7%
   37      5504944        5463820   -0.7%
   38      5750485        5309113   -7.7%
   39      5637802        5609748   -0.5%
   40      5687989        5423296   -4.7%

             nodes    time  median nps  faster half
base       5965973  1.08 s     5529746      5675861
candidate  5965973  1.09 s     5486560      5614892
change                           -0.8%        -1.1%

paired change -0.7%, 95% interval -1.8% to +0.4%
diagnostic, on the default layout alone +1.8%, 95% interval -0.3% to +4.0%, 13 rounds

not resolved: the interval reaches past ±1.0% without clearing it, so
more rounds are needed to say either way

Speed: -0.7% (bench nps, 95% interval -1.8% to +0.4%, 40 interleaved rounds over shuffled layouts vs d9379f35)

The instructions each side's bench executed, counted under
cachegrind. The count repeats to within a few hundred
instructions, so a small change here is a real one, but it
prices instructions only: cache misses, mispredicted branches
and where the code lands are the speed's to show.

instructions against d9379f35, one cachegrind run a side, bench

            instructions    nodes  per node
base       9,011,280,137  5965973    1510.4
candidate  9,024,467,921  5965973    1512.7
change            +0.15%             +0.15%

within ±0.7%, as far as edits not made for speed have moved the count

…placement

An entry's worth to the replacement policy is now its depth less eight
plies for each search since it was stored. In a full bucket the victim is
the entry worth least rather than the shallowest. A store is refused by
an entry whose worth is above its depth, or by an exact entry of the same
position whose worth equals it where the store is not exact. Before this,
age counted only at the twelve search cliff, so an entry stored for an
earlier move held its slot against a shallower one from this search until
then. Stockfish chooses its victim with the same weight.

A position's own entry from an earlier search now gives way to a
shallower store of the same position once the store reaches its worth,
where the deeper old entry used to turn it away.

Within one search every entry is worth its depth, so the policy is the old
one and the bench is identical. The change shows only across searches.
Eight self-play games put the share of stores it can reach at under 0.2%
at 256 MB and 10+0.1, and at 11% to 22% at 16 MB, where that many stores
are turned away by a deeper entry for another position in a bucket that
also holds an older one.

It was played in two stages. At 10+0.1 with a 16 MB table, where the
policy decides most, it read +19 ±13 over 1,500 games and passed sprt
[0, 10]. At 30+0.3 with the default 256 MB it read +4 ±9 over 2,500 games
and passed sprt [-10, 0], after one extension past a 2,000 game cap the
test reached at an LLR of 2.93 against 2.94. In those games the candidate
counted 0.968 to 0.985 of the baseline's nodes a second. That is the tree
and not the code: replayed on the same positions at fixed nodes the two
run level, and the gap sits in long endgames, where the baseline's old
entries answer nodes almost for nothing and the candidate searches two
plies deeper in the same time.

The age is read in the victim loop: bench games +0.19% instructions,
node counts identical.

Elo: +4 ±9 (sprt [-10, 0] passed, 2500 games, 30+0.3, vs 31eb33b)
Bench: 5965973
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WmzY2g6SZGYszT5dnuQ6tw
@aywrite
aywrite force-pushed the claude/tt-aging-spec-mosqqn branch from df2851f to 195ccf1 Compare October 5, 2026 08:25
@aywrite
aywrite merged commit 813803e into master Oct 5, 2026
21 checks passed
@aywrite
aywrite deleted the claude/tt-aging-spec-mosqqn branch October 5, 2026 09:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants