Skip to content

Validate each input signature once instead of every port on every call - #83

Merged
martyn merged 3 commits into
developfrom
perf/check-once
Sep 26, 2026
Merged

martyn merged 3 commits into
developfrom
perf/check-once

Conversation

@martyn

@martyn martyn commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

GraphModule._execute() used to check every port of every node, and replay the registered-state walk, on every forward call. That was a fixed ~6 µs per node on top of the modules themselves. This PR checks everything once per input signature, keeps the hot path close to a plain loop of module calls, and makes failures readable.

What is checked when

When What runs
First call for a signature: each external input's (shape, dtype, device), plus training and autocast state The full checked program: every port's shape, dtype and device before and after its node, then the registered-state check (same _state_matches replay as c78a43e). The signature is remembered (up to 64 per model, then the cache is cleared).
Later calls with a remembered signature Build the signature tuple (~0.6 µs), a dict lookup, one integer compare of the watch generation, then the modules from a slot-indexed program. Single-input nodes call module(values[i]) directly.
Any registration hook fires during such a call The state check runs at the end of that call, as before.

What invalidates it

  • _apply (.to(), .cuda(), .double(), .half(), ...) clears every validated signature and forces a full state re-check before the next call.
  • Registering state on any module of the graph. register_parameter, register_buffer, add_module/register_module, and attribute assignment all go through torch's global registration hooks (register_module_{parameter,buffer,module}_registration_hook). A hook filtered to modules of built graphs advances a watch generation, so unrelated modules constructed elsewhere cost nothing. On the next call the graph re-checks state against the build, clears its signatures (a parameter replaced under the same name can change any shape), and revalidates every port.
  • deepcopy gives the clone a fresh cache and watches the clone's modules. load_state_dict needs nothing: it copies in place, and assign=True goes through setattr, so the hooks fire.

What the state check protects against (after reading c78a43e / 5b880dd)

c78a43e restored detection of: filling a declared-None parameter/buffer/submodule slot (the lazy-init bug), self.weight = None, and replacing or grafting a submodule. In this PR:

  • During forward on the first call for a signature, all of these are caught exactly as before, with the same _state_matches walk. That's where lazy init happens.
  • During a later forward, or between calls, anything that goes through a registration hook is caught: filling slots, grafting or replacing submodules, and None over a buffer or submodule. test_every_kind_of_state_mutation_during_forward_is_rejected passes unchanged.
  • The gap (please review): PyTorch runs no hook for del module.x, for module.param = None (register_parameter(name, None) skips hooks), or for writes straight to _parameters/_buffers/_modules. If one of these happens after a signature has been validated, it is reported at the next full check (new signature, .to()/cast, any hooked registration) or as soon as a node then raises. The failure path always re-runs the state check, so del model["head"].weight gives a readable "parameter 'nodes.n_head.weight' was removed" error rather than a bare AttributeError. The only silent case is param = None on a module that tolerates it, e.g. linear.bias = None, which runs bias-free until the next revalidation. test_a_parameter_set_to_none_is_reported_at_the_next_revalidation pins this down. Closing it would need a per-call sweep of store.get(key) is not None over every registered tensor, about 50 ns each. I left that out to keep the hot path a straight loop, but it's a small follow-up if you want it.

Errors

Port mismatch after a validated call:

E_RUNTIME: input 'x': expected shape [B=3, 4], got [3, 5]
  expected [B=3, 4]:float32 on cpu
  got      [3, 5]:float32 on cpu

An exception raised inside a node propagates as itself, with the same type, message and traceback, so except RuntimeError and OOM handling keep working. It gains one add_note(), which the traceback prints under the message. The note is built only in the except path, so there's no per-call cost:

RuntimeError: mat1 and mat2 shapes cannot be multiplied (2x4 and 3x3)
HNDL: raised inside node 'odd' (flaky, line 2, column 1)
  input 'x': got [2, 4]:float32 on cpu; contract [B=2, 4]:float32 on cpu

A deleted parameter (no hook fires) explains itself in the note:

AttributeError: 'Linear' object has no attribute 'weight'
HNDL: raised inside node 'head' (linear, line 3, column 1)
  registered state no longer matches the build: parameter 'nodes.n_head.weight' was removed
  input 'x': got [3, 8]:float32 on cpu; contract [B=3, 8]:float32 on cpu

An upstream node that changes its output after validation is named at the port it breaks, in the note on the downstream node's error (linear 'head' input 'x' (from node:d/out): expected shape [B=2, 4], got [2, 2] plus an explanation). An exception passing out through nested graphs (a built graph used as a node of another) keeps only the innermost node's note. Errors that come from dynamo or inductor get a line saying so. HNDLError/E_RUNTIME is kept for what HNDL's own checks find: first-call validation and registered-state mismatches.
Registered state changed:

E_RUNTIME (node n0; line 1, column 1): registered state of lazy 'n0' changed during forward: submodule 'nodes.n_n0.inner.proj' was filled in with Linear (it was None at build)
  HNDL fixes every parameter, buffer and submodule when it builds a plan, so state added or removed later is not trained, seeded or saved as the plan describes; create it in __init__, or resolve and build a new plan to change the architecture

Other details:

  • A second input that disagrees on batch gets the hint B=3 is this call's batch size, read from input 'x'.

Behavior changes to review

  1. Exception types are unchanged for anything raised inside a node; HNDL only adds a note. (The first revision of this PR wrapped them in HNDLError. That broke except RuntimeError callers and was replaced.)
  2. Dtype spelling in messages: expected dtype float64, got float32 (HNDL notation) instead of torch.float64. The dtype assertions in tests/test_torch.py were updated for this.
  3. Per-call guarantees: after the first call for a signature, a node whose output shape depends on data is not re-checked on every call. It's caught when a later node raises, and reported at the port. External outputs are not re-checked per call either.
  4. Autocast: it's part of the signature, so the first autocast call is validated (and, as before, fails the dtype check). A signature validated without autocast is never reused under it.

Tests

  • test_deepcopy_keeps_lookup_moves_and_the_state_consistency_check used to tamper with _state_program after a successful call and expect the next call to notice. With per-call checking gone that no longer holds, so I rewrote it to do a real mutation on the clone (clone["head"].ghost = nn.Parameter(...)). It now asserts that the clone reports it, that the original keeps running, and that the original still catches its own mutations.
  • Dtype message assertions: torch.float64 became float64 (see above).
  • New tests: once-per-signature caching and its invalidation (batch, training mode, .to(), and that unrelated module construction does not invalidate); a broken input after the first call (exact message); a second input's batch hint; a torch error inside a node on the validating call and on the fast path (type(...) is RuntimeError, exact note); OOM keeping its type; an HNDLError inside a node keeping its message and gaining a note; one note only through nested graphs; upstream output drift named in the note; five mutations after the first call (grafted submodule, added parameter and buffer set to None → HNDLError; deleted parameter → AttributeError with a state note; wrong-shaped parameter → RuntimeError), each reported on every call; the param = None gap; state registered during a later forward; and torch.compile(backend="eager", fullgraph=True) tracing the checked program with no graph breaks.
  • Full suite after rebasing onto Keep 0.6.0 saved plans loading when operators gain arguments #81 and Relax strictness that costs more than it protects #82: PYTHONPATH=src python -m pytest -q -m "not network and not benchmark" gives 3586 passed, 80 skipped.

Benchmarks

CPU, torch.set_num_threads(1), torch 2.14. The machine was shared (load average 16–25), so each row is the best of 9 interleaved hndl/hand-written measurements, with ranges over 3 alternating old/new process runs. "Before" is origin/develop at 7f7cd3d.

Case Before: hndl / hand Before: ratio After: hndl / hand After: ratio Overhead before → after
MLP (5 nodes), batch 1 53 / 23 µs 2.28–2.35x 28 / 22 µs 1.21–1.30x 30 → 5–7 µs
MLP, batch 32 80–92 / 46–49 µs 1.73–1.89x 51–54 / 44–47 µs 1.15–1.21x 33–43 → 7–9 µs
MLP, batch 256 207–210 / 173–175 µs 1.19–1.21x 183–185 / 172–177 µs 1.04–1.06x 33–37 → 7–11 µs
transformer_block, batch 8 269–279 / 205–210 µs 1.29–1.33x 243–247 / 207–210 µs 1.16–1.18x 61–69 → 34–38 µs
20× relu, batch 1 102–109 / 45–48 µs 2.27–2.29x 51–54 / 44–46 µs 1.17–1.21x 57–61 → 7–9 µs
conv_stack, batch 16 ~5 ms 0.96–1.10x ~5 ms 0.97–1.6x (noise) lost in ms-level noise

The fixed cost is now ~2 µs per call (the signature, a dict lookup and the output dict) plus ~0.35 µs per node. The ~35 µs left on transformer_block sits inside that single operator's own implementation compared with nn.MultiheadAttention, not in per-node validation.

pytest -m benchmark tests/benchmark -q -s passes before and after (28 passed, 1 skipped). Parity ratios from one run each: mlp 1.22x → 1.09–1.14x, conv_stack 0.74x → 1.06–1.07x (noise), transformer_block 1.13x → 1.08–1.09x. All compile-compat networks compile as before. moe_transformer is skipped in both runs by an inductor CppCompileError, which is environmental.

TOLERANCE: I did not change it. The MLP now runs ~1.2x at batch 32, so the parity suite could time the MLP at batch 32, the case this issue was about, and TOLERANCE could come down from 2.0 to about 1.5. I haven't done that because of the shared-machine noise on CI hosts. Your call.

Other notes

  • The global registration hooks are installed once at import hndl.torch, for the whole process. Each is a WeakValueDictionary lookup that runs only on registrations, never in forward. Returning None means they never alter a registration.
  • Under torch.compiler.is_compiling(), _execute traces the full checked program and state check exactly as before (dynamo turns them into guards), and skips the signature cache.
  • The docs changes are the CHANGELOG bullet under ## Unreleased, README, IMPLEMENTATION.md, SPEC.md §3/§8/§11/§13, and the parity benchmark docstring.

🤖 Generated with Claude Code

martyn and others added 3 commits September 26, 2026 10:13
GraphModule._execute() checked the shape, dtype and device of every port of
every node, and replayed the registered-state walk, on every forward call:
a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x
hand-written PyTorch.

The first call for a given input signature (every external input's shape,
dtype and device, plus training mode and autocast state) now runs the full
checked program and the state check, as before. The signature is then
remembered, and later calls with it compare the signature and run the
modules back to back from a slot-indexed program.

What was validated is invalidated on the events that can break it:
- .to()/.cuda()/.double()/... (_apply) drops every signature and forces a
  full state re-check;
- registering a parameter, buffer or submodule on any module in the graph
  (register_* or attribute assignment) advances a watch generation through
  torch's global registration hooks, filtered to modules of built graphs, so
  the next call re-checks state and revalidates every port; a registration
  during a fast call is checked at the end of that call, as before.

Errors are readable: port mismatches name the node, operation and source
line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu")
beside the tensor that arrived; torch errors raised inside a node are
wrapped as E_RUNTIME with the node's inputs and any state change that
explains them; OOM errors pass through unwrapped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are
checked, what invalidates a validated signature, the new error wording, and
the registered-state edits PyTorch runs no hook for. The parity benchmark's
docstring now describes the cheap path it actually times.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Wrapping a torch error raised inside a node in HNDLError (a ValueError)
broke callers that catch RuntimeError around a forward call. The original
exception now propagates unchanged --- type, message and traceback --- and
gains one add_note() naming the node, its operation and source line, its
inputs against their contract, any registered-state change that explains
the failure, the upstream port that drifted after validation, and a hint
when the error came from torch.compile. An exception passing out through
nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME)
is kept for what HNDL's own checks find: first-call validation and state
mismatches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@martyn
martyn merged commit 17532f1 into develop Sep 26, 2026
5 checks passed
@martyn
martyn deleted the perf/check-once branch September 26, 2026 16:16
This was referenced Sep 26, 2026
martyn added a commit that referenced this pull request Sep 26, 2026
* Add equalized learning rate to linear and transformer projections

* Add HNDL primitives for style transformer coordinate rendering

* Bump README status to 0.6.0 and drop host-application references

The README status line was missed in the 0.6.0 release; a test now ties it
to hndl.__version__. HNDL is a base library, so release notes, SPEC, and
test docstrings no longer name a particular downstream consumer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep 0.6.0 saved plans loading when operators gain arguments (#81)

Adding equalized= to linear, attention and feed_forward without a version
bump made every 0.6.0 plan holding those operators fail to load with
E_INTEGRITY and changed their semantic digests.

Arg gains since="<release>" for arguments added to a released operator.
A resolved node that holds such an argument's default omits it from its
args and argument origins; construct() fills it back in. Plans that do not
use the new argument keep their 0.6.0 bytes and digests, and 0.6.0 plans
load. equalized is declared since="0.7.0".

Adds plan fixtures generated by the 0.6.0 release and a regression test
that loads, builds and re-resolves them, and documents the rule in
docs/ADDING_OPERATORS.md.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Relax strictness that costs more than it protects (#82)

* Parse configs in process instead of in a subprocess worker

The AST allowlist is unchanged and source is still never compiled or
executed. hndl/_parser.py replaces the python -I -S worker, its JSON wire
protocol, response validation and resource limits. Before ast.parse, a
token-stream screen bounds bracket nesting (50) and Python operators or
keywords (32; valid configs use none): on CPython 3.11-3.13 a few thousand
chained operators otherwise crash ast.parse with SIGSEGV on a small thread
stack. Parser MemoryError/RecursionError become E_RESOURCE.

Statement handling is one dispatch table on each side: Validator.STATEMENTS
in _parser.py and _Interpreter.STATEMENTS in config.py.

The concat timing test now counts solver shape updates instead of racing a
5 s subprocess timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Explain plan mismatches and fill values a saved plan determines

Revalidation lists every differing port shape and argument per node instead
of a one-line E_INTEGRITY. A plan missing only an operator default or an
argument bound to a verified port dimension loads completed, with a warning
naming the filled values and the new semantic digest. Missing values only a
policy or relation search would choose are still refused. Digest mismatches
say the file is corrupted or was hand-edited.

Raise the default max_state_bytes from 1 GiB to 64 GiB (a 405M-parameter
network failed to build by default), drop from_json's 16 MiB cap and its
duplicate max_nodes check, and have capture read max_nodes from the shared
limits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document in-process parsing, plan completion and limit defaults

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make the attention deepcopy test perturb bias entries independently

Adding a constant to the whole relative-position bias table shifts every
logit equally, which softmax cancels, so the assertion passed or failed on
rounding noise (flaky on CI 3.14). Independent normal perturbations change
the output for every input tried (0 of 500 indistinguishable).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Validate each input signature once instead of every port on every call (#83)

* Validate each input signature once instead of every port on every call

GraphModule._execute() checked the shape, dtype and device of every port of
every node, and replayed the registered-state walk, on every forward call:
a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x
hand-written PyTorch.

The first call for a given input signature (every external input's shape,
dtype and device, plus training mode and autocast state) now runs the full
checked program and the state check, as before. The signature is then
remembered, and later calls with it compare the signature and run the
modules back to back from a slot-indexed program.

What was validated is invalidated on the events that can break it:
- .to()/.cuda()/.double()/... (_apply) drops every signature and forces a
  full state re-check;
- registering a parameter, buffer or submodule on any module in the graph
  (register_* or attribute assignment) advances a watch generation through
  torch's global registration hooks, filtered to modules of built graphs, so
  the next call re-checks state and revalidates every port; a registration
  during a fast call is checked at the end of that call, as before.

Errors are readable: port mismatches name the node, operation and source
line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu")
beside the tensor that arrived; torch errors raised inside a node are
wrapped as E_RUNTIME with the node's inputs and any state change that
explains them; OOM errors pass through unwrapped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document validation once per input signature

CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are
checked, what invalidates a validated signature, the new error wording, and
the registered-state edits PyTorch runs no hook for. The parity benchmark's
docstring now describes the cheap path it actually times.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep exceptions raised inside a node as their own type, with a note

Wrapping a torch error raised inside a node in HNDLError (a ValueError)
broke callers that catch RuntimeError around a forward call. The original
exception now propagates unchanged --- type, message and traceback --- and
gains one add_note() naming the node, its operation and source line, its
inputs against their contract, any registered-state change that explains
the failure, the upstream port that drifted after validation, and a hint
when the error came from torch.compile. An exception passing out through
nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME)
is kept for what HNDL's own checks find: first-call validation and state
mismatches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add bounded for loops to the config grammar (#84)

`for _ in range(N):` with a positive int literal N repeats its body, which
may hold any top-level statement including nested loops. The interpreter
unrolls it before resolution, so a loop yields the same node IDs, plan,
semantic digest and state_dict keys as the statements written out by hand.
Explicit names gain the per-level iteration suffix (block3, res1_3), node
source metadata records the iterations, and configuration, resolution and
runtime errors report them. The multiplied-out node count is checked against
max_nodes before the first node is emitted, loops nest at most 8 levels, and
every other loop, conditional and comprehension form is rejected by name.
The ViT and GPT example networks now use loops for their blocks.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Release 0.7.0

Bump version, README status, and rename the Unreleased changelog heading for
bounded config loops, once-per-signature contract checks, in-process config
parsing, forward-compatible saved plans, and the equalized and coordinate
rendering primitives.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
martyn added a commit that referenced this pull request Sep 26, 2026
* Add equalized learning rate to linear and transformer projections

* Add HNDL primitives for style transformer coordinate rendering

* Bump README status to 0.6.0 and drop host-application references

The README status line was missed in the 0.6.0 release; a test now ties it
to hndl.__version__. HNDL is a base library, so release notes, SPEC, and
test docstrings no longer name a particular downstream consumer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep 0.6.0 saved plans loading when operators gain arguments (#81)

Adding equalized= to linear, attention and feed_forward without a version
bump made every 0.6.0 plan holding those operators fail to load with
E_INTEGRITY and changed their semantic digests.

Arg gains since="<release>" for arguments added to a released operator.
A resolved node that holds such an argument's default omits it from its
args and argument origins; construct() fills it back in. Plans that do not
use the new argument keep their 0.6.0 bytes and digests, and 0.6.0 plans
load. equalized is declared since="0.7.0".

Adds plan fixtures generated by the 0.6.0 release and a regression test
that loads, builds and re-resolves them, and documents the rule in
docs/ADDING_OPERATORS.md.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Relax strictness that costs more than it protects (#82)

* Parse configs in process instead of in a subprocess worker

The AST allowlist is unchanged and source is still never compiled or
executed. hndl/_parser.py replaces the python -I -S worker, its JSON wire
protocol, response validation and resource limits. Before ast.parse, a
token-stream screen bounds bracket nesting (50) and Python operators or
keywords (32; valid configs use none): on CPython 3.11-3.13 a few thousand
chained operators otherwise crash ast.parse with SIGSEGV on a small thread
stack. Parser MemoryError/RecursionError become E_RESOURCE.

Statement handling is one dispatch table on each side: Validator.STATEMENTS
in _parser.py and _Interpreter.STATEMENTS in config.py.

The concat timing test now counts solver shape updates instead of racing a
5 s subprocess timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Explain plan mismatches and fill values a saved plan determines

Revalidation lists every differing port shape and argument per node instead
of a one-line E_INTEGRITY. A plan missing only an operator default or an
argument bound to a verified port dimension loads completed, with a warning
naming the filled values and the new semantic digest. Missing values only a
policy or relation search would choose are still refused. Digest mismatches
say the file is corrupted or was hand-edited.

Raise the default max_state_bytes from 1 GiB to 64 GiB (a 405M-parameter
network failed to build by default), drop from_json's 16 MiB cap and its
duplicate max_nodes check, and have capture read max_nodes from the shared
limits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document in-process parsing, plan completion and limit defaults

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Make the attention deepcopy test perturb bias entries independently

Adding a constant to the whole relative-position bias table shifts every
logit equally, which softmax cancels, so the assertion passed or failed on
rounding noise (flaky on CI 3.14). Independent normal perturbations change
the output for every input tried (0 of 500 indistinguishable).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Validate each input signature once instead of every port on every call (#83)

* Validate each input signature once instead of every port on every call

GraphModule._execute() checked the shape, dtype and device of every port of
every node, and replayed the registered-state walk, on every forward call:
a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x
hand-written PyTorch.

The first call for a given input signature (every external input's shape,
dtype and device, plus training mode and autocast state) now runs the full
checked program and the state check, as before. The signature is then
remembered, and later calls with it compare the signature and run the
modules back to back from a slot-indexed program.

What was validated is invalidated on the events that can break it:
- .to()/.cuda()/.double()/... (_apply) drops every signature and forces a
  full state re-check;
- registering a parameter, buffer or submodule on any module in the graph
  (register_* or attribute assignment) advances a watch generation through
  torch's global registration hooks, filtered to modules of built graphs, so
  the next call re-checks state and revalidates every port; a registration
  during a fast call is checked at the end of that call, as before.

Errors are readable: port mismatches name the node, operation and source
line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu")
beside the tensor that arrived; torch errors raised inside a node are
wrapped as E_RUNTIME with the node's inputs and any state change that
explains them; OOM errors pass through unwrapped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Document validation once per input signature

CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are
checked, what invalidates a validated signature, the new error wording, and
the registered-state edits PyTorch runs no hook for. The parity benchmark's
docstring now describes the cheap path it actually times.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep exceptions raised inside a node as their own type, with a note

Wrapping a torch error raised inside a node in HNDLError (a ValueError)
broke callers that catch RuntimeError around a forward call. The original
exception now propagates unchanged --- type, message and traceback --- and
gains one add_note() naming the node, its operation and source line, its
inputs against their contract, any registered-state change that explains
the failure, the upstream port that drifted after validation, and a hint
when the error came from torch.compile. An exception passing out through
nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME)
is kept for what HNDL's own checks find: first-call validation and state
mismatches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Add bounded for loops to the config grammar (#84)

`for _ in range(N):` with a positive int literal N repeats its body, which
may hold any top-level statement including nested loops. The interpreter
unrolls it before resolution, so a loop yields the same node IDs, plan,
semantic digest and state_dict keys as the statements written out by hand.
Explicit names gain the per-level iteration suffix (block3, res1_3), node
source metadata records the iterations, and configuration, resolution and
runtime errors report them. The multiplied-out node count is checked against
max_nodes before the first node is emitted, loops nest at most 8 levels, and
every other loop, conditional and comprehension form is rejected by name.
The ViT and GPT example networks now use loops for their blocks.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Close extension API gaps for operators defined outside HNDL (#86)

* Close extension API gaps for operators defined outside HNDL

- The global `ops` namespace resolves aliases against the registry of the
  active capture, so custom operators work as `ops.my_op()` under
  resolve_callable/network_from_callable(registry=...). A registry.ops
  factory whose exact declaration the capture's registry lacks now fails
  with E_CAPTURE and a message naming the registry to pass.
- `hndl.relations` publishes the convolution arithmetic, spatial,
  elementwise_join and broadcast relations, plus conv_input_range,
  conv_transpose_input, conv_axis and conv_transpose_axis. Built-ins import
  from it; operators._relations re-exports it.
- `hndl.testing` publishes the operator harness (check_operator and the
  per-check functions, example_params/operator_params for pytest). The
  built-in test_all_operators runs through it. pytest is imported lazily.
- Tests in tests/test_extension_api.py exercise all three as an external
  package would. SPEC, README, ADDING_OPERATORS and CHANGELOG updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Test the harness's failure paths and address review nits

- Exercise every check_build_and_run, check_declaration, check_round_trip
  and check_reference failure with a deliberately broken operator (or, for
  checks that guard HNDL itself, a patched resolver/repr/replay), asserting
  each message, so a check that stops checking fails the suite.
- check_reference no longer crashes when a reference drops the gradient the
  module carries; it reports it.
- check_declaration requires the class's own docstring instead of accepting
  one inherited from nn.Module.
- check_operator accepts devices="cpu" / dtypes="float32".
- A declaration missing from the registry fails with AssertionError, like
  one registered with a different declaration.
- The E_CAPTURE hint only suggests ops.<alias> when that alias binds the
  same identity in the capture's registry.
- conv_axis / conv_transpose_axis fail with E_CONSTRAINT for a port whose
  rank lacks the axis instead of leaking IndexError.
- The README's my_silu declares an Example so check_operator runs on it,
  with a test that executes the README block through the harness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Detect cleared parameters on the fast path and tighten the parity tolerance (#87)

* Detect registered-state edits PyTorch runs no hook for on the fast path

`del module.weight`, `linear.bias = None` and direct writes to a module's
`_parameters`, `_buffers` or `_modules` fire none of torch's global
registration hooks, so after the first validated call the unchecked path
kept running (a layer silently without its bias) until a new signature, a
move or cast, or a failing layer forced a full check.

Each graph now records every module's three stores together with a copy of
each whenever its state is found to match the build, and every call compares
the two tuples in one C-level comparison: values are compared by identity
first, so an unchanged store costs about 10-17 ns and never touches its
tensors. A mismatch takes the validating path, which reports removed or
added state with the usual E_RUNTIME naming the node and the change, and
re-checks every port when a tensor was swapped in under the same name.
.to()/casts drop the recorded copies so replaced tensors are released
immediately. torch.compile tracing is unchanged: the compiling path returns
before the comparison.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Tighten the parity benchmark tolerance from 2.0x to 1.5x

Sixteen runs of the opt-in benchmark suite with the registered-state
comparison in place measured mlp 1.00-1.13x, conv_stack 0.97-1.17x and
transformer_block 1.10-1.24x of hand-written PyTorch. 1.5x leaves about 20%
over the worst of those. The docstring records the runs, the machine, and the
one noisy conv_stack reading (1.91x, on the code before this branch) that a
loaded machine can produce.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Compare registered state without holding or comparing tensors

The per-call store check compared copies of every _parameters/_buffers
dict with ==, so a tensor swapped in under its name (functional_call)
ran an elementwise Tensor.__eq__ and a full re-check on every call,
kept the swapped tensors and their autograd graphs alive, and made
copy.deepcopy fail on non-leaf, grad-tracking or vmap-batched tensors.

Submodule dicts and empty tensor dicts are still compared by identity;
non-empty tensor dicts are compared by length and which values are None,
so removals, additions and None assignments are caught and swaps are
not. __deepcopy__ skips the record, _validate records it once, and the
docs list what is not detected.

The parity benchmark now takes the fastest of ten interleaved rounds per
side and keeps 2.0x for the arithmetic-bound convolution case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Release 0.8.0

Bump version, README status, and rename the Unreleased changelog heading for
the extension API (registry-aware ops, public hndl.relations, hndl.testing)
and fast-path detection of hookless state edits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant