Skip to content

refactor(search): Hold a node's facts and its answer in one value and ask for a child search with the loop's decision - #369

Merged
aywrite merged 2 commits into
masterfrom
search/one-node
Oct 4, 2026
Merged

aywrite merged 2 commits into
masterfrom
search/one-node

Conversation

@aywrite

@aywrite aywrite commented Oct 4, 2026

Copy link
Copy Markdown
Owner

Two commits, meant to land as two (rebase merge), because each carries its
own measurement against its own parent.

55d19a3 refactor(search): Hold a node's facts and its answer in one value
the reductions read.
A full width node's state was held three times:
NodeFacts for the recorders, NodeAnswer for the bounds and the fail soft
best, and late_move::Node, which repeated six of the seven facts beside its
memos and the two rule halves. The loop kept them in step by handing the late
move node alpha after every rise, checked only by debug assertions. Now one
Node holds the facts and the answer, alpha lives there alone, and the late
move struct keeps only its memos and rule halves as Rules, reading the node
per call. raised and the assertions go. New tests read the node with no
search behind it, and two pin what the census and the ledger record after a
rise.

fb42c41 refactor(search): Ask for a child search with the loop's
decision.
search_child and windowed took first: bool, reduction: u8, staged: Option<&Staged> and an assertion that a first move carried no
reduction. Decision gains First beside Skip and Search { reduction, staged }, and both functions take it. The first move no longer asks the late
move rules, which could not act on it.

What a reviewer should know

  • The tree does not move. arche bench counts 5,965,973 on the base and on
    both commits, and every instrument's rows at depth 5 (cutoffs, reductions,
    effort with and without null_move off, residuals) print identically, at
    the default sampling and at every 1. A second review's own builds also
    held the reference configuration, each switch off, each shallow rule alone
    and bench 7 identical.
  • Callgrind at bench 5, each commit against its parent: the node +0.47%,
    the child value -0.58%, together -0.11%. Of the node's rise, 175,428 is the
    mate test on alpha, asked per admitted move where it was per rise; the rules'
    own lines fell 155,063 and the rest is register allocation in lines the
    change does not touch. A one-bit memo of a mated alpha on the node, to make
    the test per rise again, was measured and read 67,365 worse, so it is not
    taken. A speed round on the pair read -0.8% with a 95% interval from -2.3%
    to +0.9%, not resolved, as a change this size should read.
  • The review found no behaviour defect. It found two mutants that change
    instrument rows yet passed every test: with the opening and the risen alpha
    now in one struct, the census or the ledger reading the wrong one is a one
    word slip, and the recorder tests only built nodes that had never risen.
    Both have tests now, as does the admission's mate guard.
  • Quiescence and the root keep their own locals, and the evaluation memo is
    filled as before. Both are deliberately out of scope.

Tests go from 770 release and 772 debug in the workspace to 776 and 778.

🤖 Generated with Claude Code

… reductions read

A full width node's state was held three times. `NodeFacts` had the seven
facts the census and the effort instrument record. `NodeAnswer` had the
bounds as they stood, the opening alpha, the root bounds, the fail soft
best, the taint and the searched count. `late_move::Node` repeated six of
the seven facts beside its three memos and the two rule halves. The loop
kept the copies in step by handing the late move node alpha after every
rise, and only debug assertions checked that it had.

Now one `Node` in engine.rs holds the facts and the answer, and alpha lives
there alone. `open`, `absorb` and `raised_alpha` are its methods. The late
move struct keeps only the memos and the rule halves, and is renamed
`Rules`. Its calls take the node, so each rule half reads alpha and the
searched count off it as the move is reached: the mate test on alpha is
asked per move rather than per rise, and the futility memo records the
alpha it was found short at and asks again at any other. `raised` goes,
with the assertions that kept the copies in step and the `search`
parameter only one of them read. `record_node`, `census_event`,
`ledger_skip` and `search_table_move` read the node. The census still
records the window the node opened with, and the ledger the bounds as the
move is reached.

New tests read the node with no search behind it: a rise moves alpha and
the root bounds together and nothing lowers either, the answer is a
ceiling until a rise, and the searched count takes a legal table move and
not a pinned piece's.

The tree is unchanged. The bench counts 5,965,973, and every instrument's
rows at depth 5 print as on the parent: cutoffs, reductions, effort with
no switch off and with null_move off, and residuals, both at their default
sampling and at every 1.

Callgrind over the depth 5 bench counts 201,360,782 instructions on the
parent and 202,312,034 here, +951,252 (+0.47%). All but 12,103 of it is in
the body of alpha_beta, where every node path is inlined. There the rules'
own lines are 155,063 fewer and the mate test on alpha is 175,428 more
(is_mate and the abs it calls). The rest is spread over engine.rs,
board.rs and value.rs lines that did not change, which reads as register
allocation: a variant that changed only how the cold recorders take the
node moved the same body by a further 364,571.

Bench: 5965973
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A child search was asked for with three parameters, `first`, `reduction`
and `staged`, through `search_child` and `windowed`, and an assertion held
that a first move carried no reduction. The move loop already had the
later move's case as a value, `Decision::Search { reduction, staged }`.

`Decision` gains `First` beside `Skip` and `Search`, and both functions
take the decision by reference in place of the three. A first move cannot
carry a reduction, so the assertion goes. `late_move_decision` answers
`First` for the node's first searched move without asking the late move
rules, which could not act on it (neither shallow rule reaches a node that
has searched nothing, and the reduction starts at the fifth). The table's
move is asked for as `First`, and the root builds the same two cases from
its own locals. A `Skip` reaching `search_child` is a programming error: a
debug assertion says so there, and `windowed` treats it as unreachable.
The `too_many_arguments` allowances on the two functions go.

The tree is unchanged. The bench counts 5,965,973, and every instrument's
rows at depth 5 print as on the parent: cutoffs, reductions, effort with
no switch off and with null_move off, and residuals, both at their default
sampling and at every 1.

Callgrind over the depth 5 bench counts 202,312,034 instructions on the
parent and 201,147,133 here, -1,164,901 (-0.58%). The body of alpha_beta
is 1,658,535 fewer across its two copies, and windowed, which now matches
on the decision, is 481,276 more. Of the fall in alpha_beta, 502,559 is on
the lines of std's option.rs and 385,287 on the late move rules' lines,
which the first move no longer asks.

Bench: 5965973
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Speed against 55ac7f86

Measured on AMD EPYC 7763 64-Core Processor.
Both sides built with rustc 1.98.1 (48a229cea 2026-09-01).

Both sides built and run on this runner in this job, the way
scripts/speed.sh measures a perf commit. Each round runs
both sides on a layout of its own, the same compiled code with
its code and data shuffled and moved, so the 95% interval carries
where the code landed as well as the run. The verdict holds the
interval against a 1% threshold. The default layout, the one a
release ships, is measured after as a diagnostic. The node
counts are the search's: they move when the search does, and a
speed change leaves them alone.

round     base nps  candidate nps  change
    1      4582318        4543090   -0.9%
    2      4375908        4627507   +5.7%
    3      4428138        4627274   +4.5%
    4      4424853        4629755   +4.6%
    5      4571942        4634913   +1.4%
    6      4686409        4493346   -4.1%
    7      4563367        4613833   +1.1%
    8      4629712        4688984   +1.3%
    9      4573273        4590128   +0.4%
   10      4614372        4633625   +0.4%
   11      4558318        4622907   +1.4%
   12      4641018        4534017   -2.3%
   13      4604457        4598304   -0.1%
   14      4557037        4615525   +1.3%
   15      4610068        4649409   +0.9%
   16      4564048        4601900   +0.8%
   17      4558743        4599197   +0.9%
   18      4564048        4559628   -0.1%
   19      4680472        4539716   -3.0%
   20      4654189        4606345   -1.0%
   21      4632185        4610435   -0.5%
   22      4630424        4607544   -0.5%
   23      4667571        4632498   -0.8%
   24      4611982        4687676   +1.6%
   25      4693263        4640981   -1.1%
   26      4625796        4652491   +0.6%
   27      4637936        4670121   +0.7%
   28      4626725        4640252   +0.3%
   29      4648767        4419556   -4.9%
   30      4607110        4447843   -3.5%
   31      4702782        4663514   -0.8%
   32      4647815        4634672   -0.3%
   33      4585037        4588387   +0.1%
   34      4691757        4607814   -1.8%
   35      4636985        4643000   +0.1%
   36      4648945        4620358   -0.6%
   37      4678317        4580844   -2.1%
   38      4674834        4722063   +1.0%
   39      4623548        4619882   -0.1%
   40      4644312        4618966   -0.5%

             nodes    time  median nps  faster half
base       5965973  1.29 s     4626260      4657721
candidate  5965973  1.29 s     4619424      4647094
change                           -0.1%        -0.2%

paired change +0.0%, 95% interval -0.5% to +0.5%
diagnostic, on the default layout alone +0.5%, 95% interval -0.3% to +1.8%, 13 rounds

no change beyond ±1.0%: the whole interval is inside it

Speed: +0.0% (bench nps, 95% interval -0.5% to +0.5%, 40 interleaved rounds over shuffled layouts vs 55ac7f86)

The instructions each side's bench executed, counted under
cachegrind. The count repeats to within a few hundred
instructions, so a small change here is a real one, but it
prices instructions only: cache misses, mispredicted branches
and where the code lands are the speed's to show.

instructions against 55ac7f86, one cachegrind run a side, bench

            instructions    nodes  per node
base       9,380,787,624  5965973    1572.4
candidate  9,367,191,560  5965973    1570.1
change            -0.14%             -0.14%

within ±0.7%, as far as edits not made for speed have moved the count

@aywrite
aywrite merged commit e47ac13 into master Oct 4, 2026
21 checks passed
@aywrite
aywrite deleted the search/one-node branch October 4, 2026 05:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant