Skip to content

refactor(search): Answer quiescence and the root through one fail soft value - #374

Merged
aywrite merged 1 commit into
masterfrom
search/one-answer
Oct 5, 2026
Merged

aywrite merged 1 commit into
masterfrom
search/one-answer

Conversation

@aywrite

@aywrite aywrite commented Oct 5, 2026

Copy link
Copy Markdown
Owner

A search's fail soft answer (the window as it stands, the best score and its
move, the taint, the count searched, whether alpha rose) was written three
times: Node held it, quiescence kept six locals and repeated absorb by
hand, and search_root kept seven and did the same. It is now one value,
FailSoft, which Node holds and quiescence and the root absorb through.
Quiescence's stand pat is a named step on it, and the root's store is a named
store_root_answer with its documented exemption from the taint policy
stated where it happens.

What a reviewer should know

  • The tree does not move. arche bench counts 5,965,973 on the base and the
    branch, and its per position rows are identical under each of the four taint
    policies. Every instrument's rows at depth 5 are identical at the default
    sampling and at every 1: cutoffs, reductions, effort with and without
    null_move off, residuals, and the new forced instrument with every kind
    and three selections of kinds. Every info line of go depth 12 from the
    start position through the protocol matches. A second review's own builds
    also held bench 7, the reference configuration at depths 7 and 8, the
    games suite under three policies, seven switches off, and node limit
    sessions including an aborted iteration reported as a floor.
  • Callgrind at bench 5 against the parent: -61,670 instructions (-0.03%).
    alpha_beta and the root fell; quiescence rose 1.2% of its own cost, which
    is integer comparisons credited to the operator's line rather than the
    caller's. Two shapes were measured on the way: absorb asking beta first
    cost 0.1% more, almost all in quiescence, so it asks beta only of a score
    above alpha, as quiescence and the root did before; and quiescence's mate
    test written flag first made the compiler spill the check flag to the
    stack, seen in the disassembly, so the count is asked first. A speed round
    on an idle machine read -0.7% with a 95% interval from -2.9% to +0.6%, not
    resolved, as a change this size should read. A first round under load was
    set aside.
  • The review proved the old and new quiescence and root equivalent case by
    case (the stand pat's baseline, the order of the beta test, the best move
    and the store, the abort rule, the bound, the game over read, the taint)
    and found no behaviour defect. It found the root's abort rule had no test
    off the full window, since the bench and a fixed depth search open the root
    at the full window. Two tests now drive the root at an aspiration window no
    move beats: the search answers a ceiling, the same search stopped one node
    short answers nothing, and a root whose every move scores exactly alpha
    answers a ceiling.
  • The root's store skipping the taint policy is intended and was already
    documented on TaintPolicy::Skip; a test now holds it.

Tests go from 784 release and 786 debug in the workspace to 789 and 791.

🤖 Generated with Claude Code

…t value

A search's fail soft answer (the window as it stands, the best score and
the move that scored it, the taint, the count of moves searched, and
whether alpha rose) was written three times. `Node` held it with `absorb`
and `raised_alpha`. Quiescence kept six locals for it and repeated
`absorb` by hand in its capture loop. `search_root` kept seven and did the
same in its move loop, with `alpha != opening_alpha` as whether an aborted
iteration may answer.

The answer half of `Node` is now its own value, `FailSoft`, which `Node`
holds as `answer`. It has `open`, `absorb` (which still moves alpha and the
root bounds together at a rise) and `raised_alpha`, and a `stand_pat` step
for quiescence that sets the best score and alpha from the static score
with no move and no count. Quiescence opens one before its stand pat test,
absorbs each capture, reads a side in check with no legal move off a
searched count of zero, and stores as before (its standing evaluation with
the entry) with `raised_alpha` from the answer. The root opens one at the
aspiration window with both root bounds, absorbs each move, stops at beta,
answers with the best move and score, and reads whether an aborted
iteration may answer off `raised_alpha`. The late
move rules, the census, the effort instrument and the ledger read the
bounds and the count through `node.answer`, a field path that costs
nothing over the old one. `Node::open` takes the opened answer, and its
`too_many_arguments` allowance goes. The name keeps clear of the protocol's
`Answer` and the effort instrument's `Answered`, which are finished
answers rather than one being built.

`absorb` asks beta only of a score above alpha, as quiescence and the root
did, where `Node` asked beta first. With alpha under beta the two give the
same answer, and a move that fails low (most of them) asks one question.
A debug assertion holds the window open. Asking beta first measured
206,734 more instructions over the depth 5 bench (0.1%), most of it in
quiescence.

The root's store keeps its exemption from the taint policy, now as
`store_root_answer` with the reason beside it: the reported line is read
back from the root's slot, so the root stores whatever its taint, and
under the skip policy the rare tainted cutoff that slot offers is refused
as under refuse. No test held the exemption, so one now searches a root a
half move short of the fifty move draw under the skip policy and reads
that its one tainted store landed.

The answer's tests move to the new value, and two more read it with no
board: the stand pat sets the best score and alpha and counts no move, and
a fresh answer counts nothing searched until a move is absorbed. The bench
and a fixed depth search open the root at the full window, so no test
reached the root's ceiling or its abort below alpha. Two more call the
root at a window no move beats. A depth five search of the opening at 500
to 560 answers a ceiling, and the same search stopped one node short
answers nothing rather than the closest move. A root whose every move
scores exactly alpha answers a ceiling too, since meeting alpha does not
raise it.

The tree is unchanged. The bench counts 5,965,973, and its per position
rows print as on the parent under each of the four taint policies, with
the time masked. Every instrument's rows at depth 5 print as on the
parent: cutoffs, reductions, effort with no switch off and with null_move
off, residuals, and forced with every kind and with null_move,
reverse_futility, and skip and trusted_scout from depth 2 alone, at their
default sampling and at every 1. The info lines of `go depth 12` from the
start position through the protocol match the parent's, with the time
masked.

Callgrind over the depth 5 bench counts 194,139,961 instructions on the
parent and 194,078,291 here, -61,670 (-0.03%). By symbol, summing each
function's two copies, alpha_beta is 180,733 fewer and search_root 21,250,
windowed is unchanged, and quiescence is 140,252 more (1.2% of its own).
Quiescence's own lines in engine.rs lost 228,276 and the integer
comparisons inlined into it gained 267,714, much of that the same
comparisons credited to the operator's line rather than the caller's. Its
mate test asks the searched count before the check flag. Written the
other way round, the compiler kept the flag on the stack rather than in a
register, every test of it read memory, and quiescence ran 45,834 more.

Bench: 5965973
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Speed against 31eb33bc

Measured on AMD EPYC 7763 64-Core Processor.
Both sides built with rustc 1.98.1 (48a229cea 2026-09-01).

Both sides built and run on this runner in this job, the way
scripts/speed.sh measures a perf commit. Each round runs
both sides on a layout of its own, the same compiled code with
its code and data shuffled and moved, so the 95% interval carries
where the code landed as well as the run. The verdict holds the
interval against a 1% threshold. The default layout, the one a
release ships, is measured after as a diagnostic. The node
counts are the search's: they move when the search does, and a
speed change leaves them alone.

round     base nps  candidate nps  change
    1      4637475        4532869   -2.3%
    2      4667421        4448908   -4.7%
    3      4500291        4493045   -0.2%
    4      4413028        4536392   +2.8%
    5      4100502        4244042   +3.5%
    6      4224532        4320323   +2.3%
    7      4091217        4230025   +3.4%
    8      4217436        4285832   +1.6%
    9      4185132        4109832   -1.8%
   10      4405357        4207169   -4.5%
   11      4240178        4231843   -0.2%
   12      4215136        4188564   -0.6%
   13      4177366        4291946   +2.7%
   14      4380029        4330499   -1.1%
   15      4315042        4340616   +0.6%
   16      4333638        4262209   -1.6%
   17      4411289        4341222   -1.6%
   18      4461401        4476243   +0.3%
   19      4440709        4347575   -2.1%
   20      4135179        4483590   +8.4%
   21      4433646        4563852   +2.9%
   22      4617769        4555930   -1.3%
   23      4691971        4675256   -0.4%
   24      4682444        4678673   -0.1%
   25      4548910        4519796   -0.6%
   26      4728313        4750840   +0.5%
   27      4855532        4848772   -0.1%
   28      4880516        4833028   -1.0%
   29      4785488        4760458   -0.5%
   30      4729827        4777793   +1.0%
   31      4819924        4725036   -2.0%
   32      4842270        4786202   -1.2%
   33      4870921        4777552   -1.9%
   34      4855472        4897102   +0.9%
   35      4852471        4773985   -1.6%
   36      4829120        4866193   +0.8%
   37      4720512        4787865   +1.4%
   38      4651943        4795551   +3.1%
   39      4785039        4809360   +0.5%
   40      4836394        4818313   -0.4%

             nodes    time  median nps  faster half
base       5965973  1.30 s     4583340      4767041
candidate  5965973  1.32 s     4534630      4750908
change                           -1.1%        -0.3%

paired change -0.1%, 95% interval -0.7% to +0.7%
diagnostic, on the default layout alone +0.7%, 95% interval -0.1% to +1.4%, 13 rounds

no change beyond ±1.0%: the whole interval is inside it

Speed: -0.1% (bench nps, 95% interval -0.7% to +0.7%, 40 interleaved rounds over shuffled layouts vs 31eb33bc)

The instructions each side's bench executed, counted under
cachegrind. The count repeats to within a few hundred
instructions, so a small change here is a real one, but it
prices instructions only: cache misses, mispredicted branches
and where the code lands are the speed's to show.

instructions against 31eb33bc, one cachegrind run a side, bench

            instructions    nodes  per node
base       9,106,070,966  5965973    1526.3
candidate  9,098,728,027  5965973    1525.1
change            -0.08%             -0.08%

within ±0.7%, as far as edits not made for speed have moved the count

@aywrite
aywrite merged commit ff07eb7 into master Oct 5, 2026
21 checks passed
@aywrite
aywrite deleted the search/one-answer branch October 5, 2026 01:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant