Skip to content

Add a forced decision instrument that inverts one shortcut at a time - #372

Merged
aywrite merged 1 commit into
masterfrom
claude/fragile-root-trigger-c6yfo8
Oct 4, 2026
Merged

aywrite merged 1 commit into
masterfrom
claude/fragile-root-trigger-c6yfo8

Conversation

@aywrite

@aywrite aywrite commented Oct 4, 2026

Copy link
Copy Markdown
Owner

arche forced searches each root of a suite under the default, samples the shortcut decisions it takes, and searches the root again for each sampled decision with that one decision inverted. It prints what the root answered both times. The ledger and the residuals say whether a decision was wrong at its own node; this says whether it cost the root anything.

Four kinds can be inverted: a reverse futility cut, a null move cut, a skipped late quiet, and a trusted fail low of a reduced scout. A decision is addressed by its kind, a position key and the node's depth, and it is inverted wherever the search meets that address. kinds and from <depth> narrow the sampling. docs/INSTRUMENTS.md has the rows and what each inversion does.

The default search is unchanged. Bench is 5965973. Callgrind over bench 5 reads +0.22% instructions against master with the same nodes; the hooks are bare is_some checks in front of cold calls.

Tests: an armed recording search answers as an unarmed one does (move, score, nodes); an address never met changes nothing; every kept decision is met when forced; one row per address; the filters; the model's score only where the gate reads one; a session test against the binary. fmt, both clippy runs and both test profiles pass.

Elo: not measured (an instrument; no search change).

🤖 Generated with Claude Code

https://claude.ai/code/session_01NgTGRjqAYAXR7VCxCRqhjh


Generated by Claude Code

…ut at a time

The ledger and the residuals label a decision at its own node. Whether
being wrong there cost the root anything is a different question, and most
such errors cost nothing: a re-search recovers them, or the root's move
survives them. `arche forced` searches each root of a suite under the
default, samples the shortcut decisions it takes, and searches the root
again for each sampled decision with that one decision inverted, printing
what the root answered both times.

Four kinds can be inverted: the reverse futility margin answering a node,
the null move's cut, a late quiet skipped (by the model at depth four and
up, or a shallow rule below), and a reduced scout's fail low trusted. A
decision is addressed by its kind, a position key and the deciding node's
depth; a move decision is keyed by the position the move leaves, as the
ledger keys its rows. The inversion applies wherever the search meets the
address. `kinds` and `from <depth>` narrow the sampling, since shallow
decisions are most of them.

The hooks sit where each decision is taken, behind a bare is_some with the
body cold and out of line, as the sampler's are. While armed, the move loop
asks the shallow rules move by move as it does under the ledger, and
reduced scouts are staged so their features reach the row. The model's
score is exposed for the row where the gate reads one.

The default search is unchanged: the bench is 5965973, and callgrind over
`bench 5` reads 195,641,600 instructions against master's 195,202,622
(+0.22%), with the same nodes. The tests hold a recording arm's move,
score and nodes to an unarmed search's, an address never met to no visit
and no change, every kept decision to at least one visit, one row an
address, the kinds and depth filters, and the model's score to the depths
the gate reads it at. A session test runs the binary.

Bench: 5965973
Elo: not measured
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NgTGRjqAYAXR7VCxCRqhjh
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Speed against af1c8d4a

Measured on AMD EPYC 7763 64-Core Processor.
Both sides built with rustc 1.98.1 (48a229cea 2026-09-01).

Both sides built and run on this runner in this job, the way
scripts/speed.sh measures a perf commit. Each round runs
both sides on a layout of its own, the same compiled code with
its code and data shuffled and moved, so the 95% interval carries
where the code landed as well as the run. The verdict holds the
interval against a 1% threshold. The default layout, the one a
release ships, is measured after as a diagnostic. The node
counts are the search's: they move when the search does, and a
speed change leaves them alone.

round     base nps  candidate nps  change
    1      4264732        4222259   -1.0%
    2      4273350        4340170   +1.6%
    3      4284173        4243064   -1.0%
    4      4264287        4307636   +1.0%
    5      4193457        4291100   +2.3%
    6      4192820        4265345   +1.7%
    7      4257577        4256903   -0.0%
    8      4241797        4222800   -0.4%
    9      4312568        4314926   +0.1%
   10      4274452        4254155   -0.5%
   11      4156021        4317565   +3.9%
   12      4250967        4291270   +0.9%
   13      4254280        4338397   +2.0%
   14      4286685        4267053   -0.5%
   15      4344421        4266001   -1.8%
   16      4306010        4242249   -1.5%
   17      4280632        4257966   -0.5%
   18      4246304        4411077   +3.9%
   19      4310957        4240789   -1.6%
   20      4240117        4282211   +1.0%
   21      4289653        4211282   -1.8%
   22      4281083        4286608   +0.1%
   23      4287964        4264751   -0.5%
   24      4320348        4309254   -0.3%
   25      4223622        4258312   +0.8%
   26      4310073        4294946   -0.4%
   27      4342919        4309260   -0.8%
   28      4187405        4307841   +2.9%
   29      4198432        4266672   +1.6%
   30      4241540        4286771   +1.1%
   31      4280727        4246065   -0.8%
   32      4171037        4208529   +0.9%
   33      4232041        4328617   +2.3%
   34      4266754        4312883   +1.1%
   35      4297864        4400879   +2.4%
   36      4321202        4300181   -0.5%
   37      4206784        4317699   +2.6%
   38      4339839        4351561   +0.3%
   39      4353692        4300810   -1.2%
   40      4363818        4399487   +0.8%

             nodes    time  median nps  faster half
base       5965973  1.40 s     4273901      4309454
candidate  5965973  1.39 s     4288936      4327278
change                           +0.4%        +0.4%

paired change +0.4%, 95% interval -0.1% to +1.0%
diagnostic, on the default layout alone +1.2%, 95% interval +0.7% to +1.9%, 13 rounds

no change beyond ±1.0%: the whole interval is inside it

Speed: +0.4% (bench nps, 95% interval -0.1% to +1.0%, 40 interleaved rounds over shuffled layouts vs af1c8d4a)

The instructions each side's bench executed, counted under
cachegrind. The count repeats to within a few hundred
instructions, so a small change here is a real one, but it
prices instructions only: cache misses, mispredicted branches
and where the code lands are the speed's to show.

instructions against af1c8d4a, one cachegrind run a side, bench

            instructions    nodes  per node
base       9,084,213,258  5965973    1522.7
candidate  9,106,073,042  5965973    1526.3
change            +0.24%             +0.24%

within ±0.7%, as far as edits not made for speed have moved the count

@aywrite
aywrite merged commit 31eb33b into master Oct 4, 2026
21 checks passed
@aywrite
aywrite deleted the claude/fragile-root-trigger-c6yfo8 branch October 4, 2026 22:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants