Skip to content

Store the static evaluation in the transposition table and merge the pawn and shelter table entries - #364

Merged
aywrite merged 4 commits into
masterfrom
claude/eval-pair-sum-table-eval
Oct 4, 2026
Merged

aywrite merged 4 commits into
masterfrom
claude/eval-pair-sum-table-eval

Conversation

@aywrite

@aywrite aywrite commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Four exact changes to the evaluation and the table, one commit each. The search visits the same tree: every per position line of the bench and the games suite is unchanged (5,965,973 and 9,409,127 nodes).

  1. The pair term is kept as the sum and difference of the two perspectives, so the leaf reads one dot product where it read two sums of squares.
  2. The shelter table's entry holds the shelter and the pawn structure added, and the pawn table is asked only behind a shelter miss (about 11% of evaluations). Both tables are fixed length arrays and are probed before the walk.
  3. Material is one white relative number read from the piece's row.
  4. The table entry's two reserved bytes hold the position's static evaluation, and a full width node reads it from the entry its own probe found instead of evaluating again.

Rebased onto adfd966. Commit 4's stores go through master's new Node, and a test master added that calls search_table_move passes NO_EVAL.

Instructions. Callgrind, the tree search, against master adfd966: the games suite 11,748,841,555 to 11,398,942,426 (2.978% fewer), the bench 9,096,745,929 to 8,830,624,901 (2.925% fewer). By commit, in points of master's games suite: 0.893, 0.521, 0.663 and 0.901. On the old base b6e82bb the series saved 3.077%.

Clock. scripts/speed.sh on the rebased head against adfd966: +0.8% (95% interval −0.4% to +1.9%, 120 interleaved rounds over shuffled layouts), not resolved. The same comparison before the rebase read +2.2% (+0.8% to +3.8%) against b6e82bb. Each commit's Speed: trailer is 60 rounds against its own parent, on the rebased commits: +0.0%, +1.6%, +1.0% and +0.7%. A step of half a point to a point is below what 60 rounds resolve. An earlier clock of the same code over shuffled layouts read the whole series at +1.81% (+0.92% to +2.73%) on the games suite over 450 rounds, and +2.52% (+0.40% to +4.63%) on the bench over 80.

Games. Four Strength runs of 500 games at 10+0.1 on 8moves_v3 against b6e82bb, before the rebase, the decision fixed before the first game: the mean rate ratio of the four and its 95% interval. The runs (#258 to #261) read 1.0024, 1.0112, 1.0117 and 1.0115, so r = 1.0092 [1.0020, 1.0164], wholly above 1.000. Every game ended normally. The games returned about a third of the instruction saving, less than the search node change did, most likely because most of what this removes is arithmetic that overlaps with the walk.

Non-regression. An sprt [-10, 0] at 10+0.1 on 8moves_v3, the rebased head against adfd966, in batches of 500 up to 2,000 games: Strength #263 read +0 ±10 over 2,000 games, inconclusive at the cap (LLR 1.88 against ±2.94). The batches read −24, +8, +10 and +6, and the test moved toward the upper bound in each batch after the first. Every game ended normally.

Checked, on the rebased head. The release, debug, machine-test and baseline target test runs pass, as do the tactical and strategic suites. Master and the head search the 1,800 positions of tactics.epd and strategy.epd to depth 7 with identical nodes, scores and moves (101,372,856 nodes each), and 150 play-outs of up to 60 plies with the table kept agree at every ply (8,460 plies, 171,842,739 nodes each).

Constraints this leaves:

  • The pair term's identity needs the term to stay the plain difference of the two perspectives' squared norms. A lane of the sum or difference is up to twice a perspective's, which in_range checks at compile time. On the shipped factors the i32 limit on the sums of squares still binds first (a uniform scale of 1.275 against 1.314 for the lanes).
  • The shelter entry holds the pawn structure because the shelter's key holds the pawn key. A term that reads anything else cannot share it.
  • The stored evaluation is exact only while the evaluation reads nothing the table key does not cover (no fifty move counter, no line, no history). A correction to the evaluation belongs after the read, and the entry keeps the raw value.
  • A false accept on the table's 32 bit signature now also hands a node another position's evaluation for its pruning, and the node stores it back under its own key.
  • Machine::of keeps two index loops under an allow: the iterator form clippy asks for moved the search's code by 0.2% to 0.8% of its instructions.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae


Generated by Claude Code

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Speed against 587f1f68

Measured on AMD EPYC 9V74 80-Core Processor.
Both sides built with rustc 1.98.1 (48a229cea 2026-09-01).

Both sides built and run on this runner in this job, the way
scripts/speed.sh measures a perf commit. Each round runs
both sides on a layout of its own, the same compiled code with
its code and data shuffled and moved, so the 95% interval carries
where the code landed as well as the run. The verdict holds the
interval against a 1% threshold. The default layout, the one a
release ships, is measured after as a diagnostic. The node
counts are the search's: they move when the search does, and a
speed change leaves them alone.

round     base nps  candidate nps  change
    1      4298852        4408059   +2.5%
    2      4222489        4247411   +0.6%
    3      4074069        4370205   +7.3%
    4      4215422        4231360   +0.4%
    5      4282285        4440775   +3.7%
    6      4162290        4241848   +1.9%
    7      4158309        4332125   +4.2%
    8      4219918        4313154   +2.2%
    9      4125583        4311851   +4.5%
   10      4128081        4301976   +4.2%
   11      4208837        4182852   -0.6%
   12      4075427        4219608   +3.5%
   13      4050521        4176156   +3.1%
   14      4116707        4310107   +4.7%
   15      4118395        4415821   +7.2%
   16      4167474        4270829   +2.5%
   17      4284533        4261744   -0.5%
   18      4332184        4340095   +0.2%
   19      4313226        4306019   -0.2%
   20      4262965        4256298   -0.2%
   21      4314215        4284112   -0.7%
   22      4099108        4252221   +3.7%
   23      3923828        3999265   +1.9%
   24      3925524        4168927   +6.2%
   25      4014724        3993178   -0.5%
   26      3989483        3860124   -3.2%
   27      4086764        4121049   +0.8%
   28      3945342        4145619   +5.1%
   29      3895255        4074305   +4.6%
   30      4119484        4149843   +0.7%
   31      3940753        4379264  +11.1%
   32      4233155        4132464   -2.4%
   33      4118358        4243664   +3.0%
   34      4383441        4418846   +0.8%
   35      4342375        4489805   +3.4%
   36      4257367        4345038   +2.1%
   37      4321543        4367169   +1.1%
   38      4237533        4404342   +3.9%
   39      3961711        4125600   +4.1%
   40      4149110        4156053   +0.2%

             nodes    time  median nps  faster half
base       5965973  1.44 s     4153710      4260921
candidate  5965973  1.40 s     4259021      4353567
change                           +2.5%        +2.2%

paired change +2.2%, 95% interval +1.5% to +3.2%
diagnostic, on the default layout alone +1.4%, 95% interval +0.1% to +4.8%, 13 rounds

faster: the whole interval is above +1.0%

Speed: +2.2% (bench nps, 95% interval +1.5% to +3.2%, 40 interleaved rounds over shuffled layouts vs 587f1f68)

The instructions each side's bench executed, counted under
cachegrind. The count repeats to within a few hundred
instructions, so a small change here is a real one, but it
prices instructions only: cache misses, mispredicted branches
and where the code lands are the speed's to show.

instructions against 587f1f68, one cachegrind run a side, bench

            instructions    nodes  per node
base       9,367,189,087  5965973    1570.1
candidate  9,084,213,562  5965973    1522.7
change            -3.02%             -3.02%

fewer instructions, beyond the ±0.7% that edits not made for speed have moved the count

claude added 4 commits October 4, 2026 10:52
…erspectives

The pair term is white's perspective less black's, and each perspective
is a sum of squares over sixteen lanes less its diagonal. A difference of
two squared norms is the dot product of their sum and their difference,
so the accumulator now keeps the two perspectives' lanes added and
subtracted, and white's diagonal less black's. The leaf reads two
`pmaddwd` where it read four, and llvm now inlines the score into the
evaluation. The rows hold each piece's factors the same way, so make and
unmake add the same 64 bytes of lanes as before and one diagonal where
they added two. The accumulator is 84 bytes where it was 88.

A lane of the sum or the difference can reach twice a perspective's, so
`in_range` now asks that twice the largest lane bound fit an i16. The
shipped factors' largest bound is 12,467, so 24,934 against 32,767. The
i32 limit on a perspective's sum of squares still binds first: it allows
the table scaled up by 1.275, and the lanes now allow 1.314 where they
allowed 2.63. A higher rank adds to the sum of squares and to no lane's
bound. The product is the difference of two sums of squares that each
fit an i32, so it fits too and the wrapping sum lands on it.

`Machine::of`, the recompute behind the state check, builds the kept
sums with index loops under an `allow`. The iterator form clippy asks
for changed how llvm compiled the search: 0.19% more of the games
suite's instructions on this commit, and 0.85% on the series.

Callgrind, the tree search (`search_root` inclusive) on a build with
line tables, against this commit's parent: the games suite
11,748,841,555 instructions to 11,643,939,549 (0.893% fewer), the bench
9,096,745,929 to 9,022,590,898 (0.815% fewer). Every per position line
of both suites is unchanged.

Bench: 5965973
Speed: +0.0% (bench nps, 95% interval -1.4% to +1.6%, 60 interleaved rounds over shuffled layouts vs adfd966)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
… table's entry

The shelter's key is the pawn key with both kings folded in, so an entry
under it can hold the shelter and the pawn structure added. A hit is now
the one probe, and the pawn table is asked only behind a shelter miss:
10.4% of evaluations on the games suite, and 11.8% over eight games of
twelve moves played with one table kept. A king move still misses the
shelter's entry and finds the pawn structure in the pawn table, which is
why that table stays.

Both tables are arrays of their length, so the slot's mask bounds the
index and a probe carries no bounds check. The sum asks the tables
before the walk, where llvm had kept both tables' pointers across it.
The value is the same sum added in another order.

The empty entry (key zero holding zero) now stands for both terms. It is
read only at a shelter key of zero, which needs the two kings' keys to
cancel the pawn key: the same 64 bit coincidence as any wrong hit.

Callgrind, against this commit's parent: the games suite 11,643,939,549
instructions to 11,582,762,329 (0.521% of the tree search before the
series), the bench 9,022,590,898 to 8,978,986,652 (0.479%). Every per
position line of both suites is unchanged.

Bench: 5965973
Speed: +1.6% (bench nps, 95% interval +0.1% to +3.1%, 60 interleaved rounds over shuffled layouts vs d133ba1)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
The accumulator kept each side's material, a make looked the piece's
value up in a table of its own, and the score subtracted the two sides.
A row now carries the piece's material signed from white's side, as it
carries the piece square pair, so a make adds one field of the row it
already reads and the score reads one number. `count` no longer takes
the colour. The accumulator is 80 bytes.

The recount behind the state check (`Accumulator::recomputed`) and the
seeding from a fen compute the difference their own way, and the check
holds them to the field at every make and unmake in debug builds.

Callgrind, against this commit's parent: the games suite 11,582,762,329
instructions to 11,504,829,960 (0.663% of the tree search before the
series), the bench 8,978,986,652 to 8,927,303,119 (0.568%). Every per
position line of both suites is unchanged.

Bench: 5965973
Speed: +1.0% (bench nps, 95% interval -0.8% to +2.6%, 60 interleaved rounds over shuffled layouts vs 6b214ea)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
…entry

The entry's two reserved bytes now hold the static evaluation of its
position, or `NO_EVAL` where the node that stored it did not evaluate (a
node in check, one the shortcuts' gates turned away before they
evaluated, and the root's own stores). Every store below the root takes
the evaluation as an argument: quiescence its standing evaluation, a
full width node the one its shortcuts read or found. A full width node
reads the evaluation from the entry its own probe found, before the
shortcuts, so a position evaluated and stored before is not evaluated
again. The probe has already loaded the line, so the read is two bytes
of it. Quiescence still stands pat before it probes and reads nothing
new.

docs/ROADMAP.md counted this over the bench at `54d85b9` and put it at
1.5% of the evaluations (32,807 of 2,116,844), about 0.3% of the run,
counting only the evaluations the shortcuts made. Read after the probe a
full width node already makes, where it also serves the late move gate,
it spares 8.1% of the bench's evaluations (296,542 of 3,665,549), 9.8%
of the games suite's (505,451 of 5,155,451), and 15.9% over eight games
of twelve moves played with one table kept (1,860,914 of 11,697,747),
where the entries outlive the search that wrote them. The difference
from the roadmap's count has not been traced further than that. The
roadmap entry goes; the finding about the late move gate it also held
stays as an entry of its own.

The evaluation reads the placement and the side to move, which the key
covers, and not the fifty move counter, the castling rights, the en
passant square or the line, so an entry found for the position holds the
position's evaluation. The table's replacement and its bounds read none
of the new bytes, so every store, probe and cutoff is as before. A false
accept (another position under the same 32 bit signature in the same
bucket) hands the node that position's evaluation for its pruning
decisions, as it already hands over that position's move and, where the
bound allows, its score. The entry holds the raw evaluation, so a
correction that reads anything else belongs after the read. The read
takes the key, and a debug build asserts that the last probe was of it,
so a search added between the node's probe and its read fails there
rather than handing the node another position's evaluation. A new test
searches four positions and holds every evaluation stored within two
plies of the root to the evaluation computed afresh.

Callgrind, against this commit's parent: the games suite 11,504,829,960
instructions to 11,398,942,426 (0.901% of the tree search before the
series), the bench 8,927,303,119 to 8,830,624,901 (1.063%). Every per
position line of both suites is unchanged.

Bench: 5965973
Speed: +0.7% (bench nps, 95% interval -1.2% to +2.5%, 60 interleaved rounds over shuffled layouts vs a620e4b)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
@aywrite
aywrite force-pushed the claude/eval-pair-sum-table-eval branch from 1dd84a9 to dbe4207 Compare October 4, 2026 10:59
@aywrite
aywrite merged commit af1c8d4 into master Oct 4, 2026
21 checks passed
@aywrite
aywrite deleted the claude/eval-pair-sum-table-eval branch October 4, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants