The generator wrote every move twice: into one of two 512 move arrays in
its frame, and then into the caller's list when it finished. Each
array's count lived beside it in the same struct, so llvm kept the
counts in memory, and every push loaded, incremented and stored its
count and checked it against 512.
The captures now go straight into the caller's list through a cursor,
with the count and the capacity in registers. A push that finds the list
full grows it on a cold path, so a list of any width is still built. The
quiet moves go into one buffer that no position can fill (a side has at
most 63 pieces besides its king and no piece more than 27 moves), so
their push is not checked, and one copy puts them behind the captures.
With the pawns walked as sets, the captures door gives its promoting
pushes after every capture, so it writes its list once, in place, with
no buffer and no copy.
The list is unchanged move for move; only where a move waits before the
list is built has changed. The quiet buffer is 10,368 bytes, so the
frame of the full and evasions doors grows from about 6.2 KB to 10.6 KB
and takes a second probed page. A 512 move buffer for a side of sixteen
pieces or fewer, with the full one on a cold path, keeps the frame under
a page; it was exact but not shown faster on the clock (+0.37% over the
games suite at 16 MB, -0.61% at 256 MB), so the frame stays.
The move list's writes into its buffer before `set_len` were one unsafe
block in `board.rs`. The generator now has six: taking the cursor from
the list, the cursor's write, the quiet buffer's write, the copy behind
the captures with its `set_len`, the other `set_len`, and the cold
growth of the list. docs/ROADMAP.md counts them, sixteen in the crate.
The cursor and the list are held as raw pointers: a list short of a
spill keeps its moves inside itself, so a `&mut` to it reborrowed after
the cursor was taken would invalidate the cursor. Miri passes a test of
all three doors and both growth paths under Stacked Borrows and Tree
Borrows, and reports undefined behaviour on the first form, which held
the list as a `&mut`. A debug build checks every quiet push against the
buffer's room, which is none in the captures door. A new test builds a
captures list of more than 64 moves, so the cursor grows the list while
it writes.
Callgrind, against this commit's parent: the games suite 11,227,678,882
instructions to 11,077,306,805 (-1.339%; -2.730% for this commit and the
last together), the bench 8,629,364,315 to 8,526,614,230 (-1.191%;
-3.430% together). Writing the captures into the list with the pawns
walked one at a time saved 0.19%, so the two changes pay mostly
together. Every per position line of both suites is unchanged.
Bench: 5965973
Speed: +2.1% (bench nps, 95% interval -0.9% to +4.8%, 60 interleaved rounds over shuffled layouts vs ccc1d04)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae
Two exact changes to move generation, one commit each. Every list holds the same moves in the same order, so the search visits the same tree: every per position line of the bench and the games suite is unchanged (5,965,973 and 9,409,127 nodes).
Instructions. Callgrind, the tree search against master
ec2abac: the games suite 11,388,195,036 to 11,077,306,805 (2.730% fewer), the bench 8,829,440,763 to 8,526,614,230 (3.430% fewer). By commit: 1.409% and 1.339% on the games suite. The second alone, on the old pawn walk, saved 0.19% (measured onadfd966): the two pay mostly together.Clock.
scripts/speed.shon the head against masterec2abac: +4.1% (95% interval +2.2% to +6.0%, 120 interleaved rounds over shuffled layouts). An interleaved clock of the same code onadfd966over 100 layouts read the games suite at +3.16% (+2.46% to +3.85%) over 450 rounds at a 16 MB table, and +2.50% (+1.49% to +3.55%) at 256 MB over 120. The clock reads more than the instructions because the pawn loop's branches go with it: 9.4% fewer mispredicted branches over the games suite under callgrind's simulation.Games. Four Strength runs of 500 games at 10+0.1 on
8moves_v3against master, run before the rebase (the headb7d96b3against587f1f6, the same two changes), the decision fixed before the first game: the mean rate ratio of the four and its 95% interval. The runs (#264 to #267) read 1.0201, 1.0182, 1.0194 and 1.0179, so r = 1.0189 [1.0172, 1.0206], wholly above 1.000. Every game ended normally. CI's speed job read that head at +3.3% (+2.9% to +3.6%) on the bench.Checked (on the rebased head unless said otherwise).
adfd966: a build that runs master's generator beside the new one at every call asserted the same list (every field, in order), the same capture count and the samehas_legal_moveanswer over the bench, the games suite and a 448 search UCI workload ending in a 1 MB table.adfd966: a million random positions went through every door, among them lists over 64 moves, en passant into and out of check, promotions with captures, and pawns on the back ranks. All 7,063,989 distinct positions in traced searches matched an independent reference generator.machine-testand baseline target tests pass, as do the tactical and strategic suites. Master and the head search 1,800 suite positions to depth 7 and 150 play-outs of up to 60 plies identically.Constraints this leaves:
🤖 Generated with Claude Code
https://claude.ai/code/session_01DtepPHhwporfDvjPSnX4ae