Repository navigation
Store the static evaluation in the transposition table and merge the pawn and shelter table entries - #364
Conversation
Speed against
|
…erspectives The pair term is white's perspective less black's, and each perspective is a sum of squares over sixteen lanes less its diagonal. A difference of two squared norms is the dot product of their sum and their difference, so the accumulator now keeps the two perspectives' lanes added and subtracted, and white's diagonal less black's. The leaf reads two `pmaddwd` where it read four, and llvm now inlines the score into the evaluation. The rows hold each piece's factors the same way, so make and unmake add the same 64 bytes of lanes as before and one diagonal where they added two. The accumulator is 84 bytes where it was 88. A lane of the sum or the difference can reach twice a perspective's, so `in_range` now asks that twice the largest lane bound fit an i16. The shipped factors' largest bound is 12,467, so 24,934 against 32,767. The i32 limit on a perspective's sum of squares still binds first: it allows the table scaled up by 1.275, and the lanes now allow 1.314 where they allowed 2.63. A higher rank adds to the sum of squares and to no lane's bound. The product is the difference of two sums of squares that each fit an i32, so it fits too and the wrapping sum lands on it. `Machine::of`, the recompute behind the state check, builds the kept sums with index loops under an `allow`. The iterator form clippy asks for changed how llvm compiled the search: 0.19% more of the games suite's instructions on this commit, and 0.85% on the series. Callgrind, the tree search (`search_root` inclusive) on a build with line tables, against this commit's parent: the games suite 11,748,841,555 instructions to 11,643,939,549 (0.893% fewer), the bench 9,096,745,929 to 9,022,590,898 (0.815% fewer). Every per position line of both suites is unchanged. Bench: 5965973 Speed: +0.0% (bench nps, 95% interval -1.4% to +1.6%, 60 interleaved rounds over shuffled layouts vs adfd966) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
… table's entry The shelter's key is the pawn key with both kings folded in, so an entry under it can hold the shelter and the pawn structure added. A hit is now the one probe, and the pawn table is asked only behind a shelter miss: 10.4% of evaluations on the games suite, and 11.8% over eight games of twelve moves played with one table kept. A king move still misses the shelter's entry and finds the pawn structure in the pawn table, which is why that table stays. Both tables are arrays of their length, so the slot's mask bounds the index and a probe carries no bounds check. The sum asks the tables before the walk, where llvm had kept both tables' pointers across it. The value is the same sum added in another order. The empty entry (key zero holding zero) now stands for both terms. It is read only at a shelter key of zero, which needs the two kings' keys to cancel the pawn key: the same 64 bit coincidence as any wrong hit. Callgrind, against this commit's parent: the games suite 11,643,939,549 instructions to 11,582,762,329 (0.521% of the tree search before the series), the bench 9,022,590,898 to 8,978,986,652 (0.479%). Every per position line of both suites is unchanged. Bench: 5965973 Speed: +1.6% (bench nps, 95% interval +0.1% to +3.1%, 60 interleaved rounds over shuffled layouts vs d133ba1) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
The accumulator kept each side's material, a make looked the piece's value up in a table of its own, and the score subtracted the two sides. A row now carries the piece's material signed from white's side, as it carries the piece square pair, so a make adds one field of the row it already reads and the score reads one number. `count` no longer takes the colour. The accumulator is 80 bytes. The recount behind the state check (`Accumulator::recomputed`) and the seeding from a fen compute the difference their own way, and the check holds them to the field at every make and unmake in debug builds. Callgrind, against this commit's parent: the games suite 11,582,762,329 instructions to 11,504,829,960 (0.663% of the tree search before the series), the bench 8,978,986,652 to 8,927,303,119 (0.568%). Every per position line of both suites is unchanged. Bench: 5965973 Speed: +1.0% (bench nps, 95% interval -0.8% to +2.6%, 60 interleaved rounds over shuffled layouts vs 6b214ea) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
…entry The entry's two reserved bytes now hold the static evaluation of its position, or `NO_EVAL` where the node that stored it did not evaluate (a node in check, one the shortcuts' gates turned away before they evaluated, and the root's own stores). Every store below the root takes the evaluation as an argument: quiescence its standing evaluation, a full width node the one its shortcuts read or found. A full width node reads the evaluation from the entry its own probe found, before the shortcuts, so a position evaluated and stored before is not evaluated again. The probe has already loaded the line, so the read is two bytes of it. Quiescence still stands pat before it probes and reads nothing new. docs/ROADMAP.md counted this over the bench at `54d85b9` and put it at 1.5% of the evaluations (32,807 of 2,116,844), about 0.3% of the run, counting only the evaluations the shortcuts made. Read after the probe a full width node already makes, where it also serves the late move gate, it spares 8.1% of the bench's evaluations (296,542 of 3,665,549), 9.8% of the games suite's (505,451 of 5,155,451), and 15.9% over eight games of twelve moves played with one table kept (1,860,914 of 11,697,747), where the entries outlive the search that wrote them. The difference from the roadmap's count has not been traced further than that. The roadmap entry goes; the finding about the late move gate it also held stays as an entry of its own. The evaluation reads the placement and the side to move, which the key covers, and not the fifty move counter, the castling rights, the en passant square or the line, so an entry found for the position holds the position's evaluation. The table's replacement and its bounds read none of the new bytes, so every store, probe and cutoff is as before. A false accept (another position under the same 32 bit signature in the same bucket) hands the node that position's evaluation for its pruning decisions, as it already hands over that position's move and, where the bound allows, its score. The entry holds the raw evaluation, so a correction that reads anything else belongs after the read. The read takes the key, and a debug build asserts that the last probe was of it, so a search added between the node's probe and its read fails there rather than handing the node another position's evaluation. A new test searches four positions and holds every evaluation stored within two plies of the root to the evaluation computed afresh. Callgrind, against this commit's parent: the games suite 11,504,829,960 instructions to 11,398,942,426 (0.901% of the tree search before the series), the bench 8,927,303,119 to 8,830,624,901 (1.063%). Every per position line of both suites is unchanged. Bench: 5965973 Speed: +0.7% (bench nps, 95% interval -1.2% to +2.5%, 60 interleaved rounds over shuffled layouts vs a620e4b) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
1dd84a9 to
dbe4207
Compare
Four exact changes to the evaluation and the table, one commit each. The search visits the same tree: every per position line of the bench and the games suite is unchanged (5,965,973 and 9,409,127 nodes).
Rebased onto
adfd966. Commit 4's stores go through master's newNode, and a test master added that callssearch_table_movepassesNO_EVAL.Instructions. Callgrind, the tree search, against master
adfd966: the games suite 11,748,841,555 to 11,398,942,426 (2.978% fewer), the bench 9,096,745,929 to 8,830,624,901 (2.925% fewer). By commit, in points of master's games suite: 0.893, 0.521, 0.663 and 0.901. On the old baseb6e82bbthe series saved 3.077%.Clock.
scripts/speed.shon the rebased head againstadfd966: +0.8% (95% interval −0.4% to +1.9%, 120 interleaved rounds over shuffled layouts), not resolved. The same comparison before the rebase read +2.2% (+0.8% to +3.8%) againstb6e82bb. Each commit'sSpeed:trailer is 60 rounds against its own parent, on the rebased commits: +0.0%, +1.6%, +1.0% and +0.7%. A step of half a point to a point is below what 60 rounds resolve. An earlier clock of the same code over shuffled layouts read the whole series at +1.81% (+0.92% to +2.73%) on the games suite over 450 rounds, and +2.52% (+0.40% to +4.63%) on the bench over 80.Games. Four Strength runs of 500 games at 10+0.1 on
8moves_v3againstb6e82bb, before the rebase, the decision fixed before the first game: the mean rate ratio of the four and its 95% interval. The runs (#258 to #261) read 1.0024, 1.0112, 1.0117 and 1.0115, so r = 1.0092 [1.0020, 1.0164], wholly above 1.000. Every game ended normally. The games returned about a third of the instruction saving, less than the search node change did, most likely because most of what this removes is arithmetic that overlaps with the walk.Non-regression. An sprt
[-10, 0]at 10+0.1 on8moves_v3, the rebased head againstadfd966, in batches of 500 up to 2,000 games: Strength #263 read +0 ±10 over 2,000 games, inconclusive at the cap (LLR 1.88 against ±2.94). The batches read −24, +8, +10 and +6, and the test moved toward the upper bound in each batch after the first. Every game ended normally.Checked, on the rebased head. The release, debug,
machine-testand baseline target test runs pass, as do the tactical and strategic suites. Master and the head search the 1,800 positions oftactics.epdandstrategy.epdto depth 7 with identical nodes, scores and moves (101,372,856 nodes each), and 150 play-outs of up to 60 plies with the table kept agree at every ply (8,460 plies, 171,842,739 nodes each).Constraints this leaves:
in_rangechecks at compile time. On the shipped factors the i32 limit on the sums of squares still binds first (a uniform scale of 1.275 against 1.314 for the lanes).Machine::ofkeeps two index loops under anallow: the iterator form clippy asks for moved the search's code by 0.2% to 0.8% of its instructions.🤖 Generated with Claude Code
https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
Generated by Claude Code