Validate each input signature once instead of every port on every call - #83
Merged
Merged
Conversation
GraphModule._execute() checked the shape, dtype and device of every port of
every node, and replayed the registered-state walk, on every forward call:
a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x
hand-written PyTorch.
The first call for a given input signature (every external input's shape,
dtype and device, plus training mode and autocast state) now runs the full
checked program and the state check, as before. The signature is then
remembered, and later calls with it compare the signature and run the
modules back to back from a slot-indexed program.
What was validated is invalidated on the events that can break it:
- .to()/.cuda()/.double()/... (_apply) drops every signature and forces a
full state re-check;
- registering a parameter, buffer or submodule on any module in the graph
(register_* or attribute assignment) advances a watch generation through
torch's global registration hooks, filtered to modules of built graphs, so
the next call re-checks state and revalidates every port; a registration
during a fast call is checked at the end of that call, as before.
Errors are readable: port mismatches name the node, operation and source
line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu")
beside the tensor that arrived; torch errors raised inside a node are
wrapped as E_RUNTIME with the node's inputs and any state change that
explains them; OOM errors pass through unwrapped.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are checked, what invalidates a validated signature, the new error wording, and the registered-state edits PyTorch runs no hook for. The parity benchmark's docstring now describes the cheap path it actually times. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Wrapping a torch error raised inside a node in HNDLError (a ValueError) broke callers that catch RuntimeError around a forward call. The original exception now propagates unchanged --- type, message and traceback --- and gains one add_note() naming the node, its operation and source line, its inputs against their contract, any registered-state change that explains the failure, the upstream port that drifted after validation, and a hint when the error came from torch.compile. An exception passing out through nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME) is kept for what HNDL's own checks find: first-call validation and state mismatches. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
martyn
force-pushed
the
perf/check-once
branch
from
September 26, 2026 16:14
2d65b4d to
5de1dea
Compare
This was referenced Sep 26, 2026
Merged
martyn
added a commit
that referenced
this pull request
Sep 26, 2026
* Add equalized learning rate to linear and transformer projections * Add HNDL primitives for style transformer coordinate rendering * Bump README status to 0.6.0 and drop host-application references The README status line was missed in the 0.6.0 release; a test now ties it to hndl.__version__. HNDL is a base library, so release notes, SPEC, and test docstrings no longer name a particular downstream consumer. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep 0.6.0 saved plans loading when operators gain arguments (#81) Adding equalized= to linear, attention and feed_forward without a version bump made every 0.6.0 plan holding those operators fail to load with E_INTEGRITY and changed their semantic digests. Arg gains since="<release>" for arguments added to a released operator. A resolved node that holds such an argument's default omits it from its args and argument origins; construct() fills it back in. Plans that do not use the new argument keep their 0.6.0 bytes and digests, and 0.6.0 plans load. equalized is declared since="0.7.0". Adds plan fixtures generated by the 0.6.0 release and a regression test that loads, builds and re-resolves them, and documents the rule in docs/ADDING_OPERATORS.md. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Relax strictness that costs more than it protects (#82) * Parse configs in process instead of in a subprocess worker The AST allowlist is unchanged and source is still never compiled or executed. hndl/_parser.py replaces the python -I -S worker, its JSON wire protocol, response validation and resource limits. Before ast.parse, a token-stream screen bounds bracket nesting (50) and Python operators or keywords (32; valid configs use none): on CPython 3.11-3.13 a few thousand chained operators otherwise crash ast.parse with SIGSEGV on a small thread stack. Parser MemoryError/RecursionError become E_RESOURCE. Statement handling is one dispatch table on each side: Validator.STATEMENTS in _parser.py and _Interpreter.STATEMENTS in config.py. The concat timing test now counts solver shape updates instead of racing a 5 s subprocess timeout. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Explain plan mismatches and fill values a saved plan determines Revalidation lists every differing port shape and argument per node instead of a one-line E_INTEGRITY. A plan missing only an operator default or an argument bound to a verified port dimension loads completed, with a warning naming the filled values and the new semantic digest. Missing values only a policy or relation search would choose are still refused. Digest mismatches say the file is corrupted or was hand-edited. Raise the default max_state_bytes from 1 GiB to 64 GiB (a 405M-parameter network failed to build by default), drop from_json's 16 MiB cap and its duplicate max_nodes check, and have capture read max_nodes from the shared limits. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Document in-process parsing, plan completion and limit defaults Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make the attention deepcopy test perturb bias entries independently Adding a constant to the whole relative-position bias table shifts every logit equally, which softmax cancels, so the assertion passed or failed on rounding noise (flaky on CI 3.14). Independent normal perturbations change the output for every input tried (0 of 500 indistinguishable). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Validate each input signature once instead of every port on every call (#83) * Validate each input signature once instead of every port on every call GraphModule._execute() checked the shape, dtype and device of every port of every node, and replayed the registered-state walk, on every forward call: a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x hand-written PyTorch. The first call for a given input signature (every external input's shape, dtype and device, plus training mode and autocast state) now runs the full checked program and the state check, as before. The signature is then remembered, and later calls with it compare the signature and run the modules back to back from a slot-indexed program. What was validated is invalidated on the events that can break it: - .to()/.cuda()/.double()/... (_apply) drops every signature and forces a full state re-check; - registering a parameter, buffer or submodule on any module in the graph (register_* or attribute assignment) advances a watch generation through torch's global registration hooks, filtered to modules of built graphs, so the next call re-checks state and revalidates every port; a registration during a fast call is checked at the end of that call, as before. Errors are readable: port mismatches name the node, operation and source line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu") beside the tensor that arrived; torch errors raised inside a node are wrapped as E_RUNTIME with the node's inputs and any state change that explains them; OOM errors pass through unwrapped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Document validation once per input signature CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are checked, what invalidates a validated signature, the new error wording, and the registered-state edits PyTorch runs no hook for. The parity benchmark's docstring now describes the cheap path it actually times. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep exceptions raised inside a node as their own type, with a note Wrapping a torch error raised inside a node in HNDLError (a ValueError) broke callers that catch RuntimeError around a forward call. The original exception now propagates unchanged --- type, message and traceback --- and gains one add_note() naming the node, its operation and source line, its inputs against their contract, any registered-state change that explains the failure, the upstream port that drifted after validation, and a hint when the error came from torch.compile. An exception passing out through nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME) is kept for what HNDL's own checks find: first-call validation and state mismatches. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add bounded for loops to the config grammar (#84) `for _ in range(N):` with a positive int literal N repeats its body, which may hold any top-level statement including nested loops. The interpreter unrolls it before resolution, so a loop yields the same node IDs, plan, semantic digest and state_dict keys as the statements written out by hand. Explicit names gain the per-level iteration suffix (block3, res1_3), node source metadata records the iterations, and configuration, resolution and runtime errors report them. The multiplied-out node count is checked against max_nodes before the first node is emitted, loops nest at most 8 levels, and every other loop, conditional and comprehension form is rejected by name. The ViT and GPT example networks now use loops for their blocks. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Release 0.7.0 Bump version, README status, and rename the Unreleased changelog heading for bounded config loops, once-per-signature contract checks, in-process config parsing, forward-compatible saved plans, and the equalized and coordinate rendering primitives. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
martyn
added a commit
that referenced
this pull request
Sep 26, 2026
* Add equalized learning rate to linear and transformer projections * Add HNDL primitives for style transformer coordinate rendering * Bump README status to 0.6.0 and drop host-application references The README status line was missed in the 0.6.0 release; a test now ties it to hndl.__version__. HNDL is a base library, so release notes, SPEC, and test docstrings no longer name a particular downstream consumer. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep 0.6.0 saved plans loading when operators gain arguments (#81) Adding equalized= to linear, attention and feed_forward without a version bump made every 0.6.0 plan holding those operators fail to load with E_INTEGRITY and changed their semantic digests. Arg gains since="<release>" for arguments added to a released operator. A resolved node that holds such an argument's default omits it from its args and argument origins; construct() fills it back in. Plans that do not use the new argument keep their 0.6.0 bytes and digests, and 0.6.0 plans load. equalized is declared since="0.7.0". Adds plan fixtures generated by the 0.6.0 release and a regression test that loads, builds and re-resolves them, and documents the rule in docs/ADDING_OPERATORS.md. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Relax strictness that costs more than it protects (#82) * Parse configs in process instead of in a subprocess worker The AST allowlist is unchanged and source is still never compiled or executed. hndl/_parser.py replaces the python -I -S worker, its JSON wire protocol, response validation and resource limits. Before ast.parse, a token-stream screen bounds bracket nesting (50) and Python operators or keywords (32; valid configs use none): on CPython 3.11-3.13 a few thousand chained operators otherwise crash ast.parse with SIGSEGV on a small thread stack. Parser MemoryError/RecursionError become E_RESOURCE. Statement handling is one dispatch table on each side: Validator.STATEMENTS in _parser.py and _Interpreter.STATEMENTS in config.py. The concat timing test now counts solver shape updates instead of racing a 5 s subprocess timeout. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Explain plan mismatches and fill values a saved plan determines Revalidation lists every differing port shape and argument per node instead of a one-line E_INTEGRITY. A plan missing only an operator default or an argument bound to a verified port dimension loads completed, with a warning naming the filled values and the new semantic digest. Missing values only a policy or relation search would choose are still refused. Digest mismatches say the file is corrupted or was hand-edited. Raise the default max_state_bytes from 1 GiB to 64 GiB (a 405M-parameter network failed to build by default), drop from_json's 16 MiB cap and its duplicate max_nodes check, and have capture read max_nodes from the shared limits. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Document in-process parsing, plan completion and limit defaults Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Make the attention deepcopy test perturb bias entries independently Adding a constant to the whole relative-position bias table shifts every logit equally, which softmax cancels, so the assertion passed or failed on rounding noise (flaky on CI 3.14). Independent normal perturbations change the output for every input tried (0 of 500 indistinguishable). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Validate each input signature once instead of every port on every call (#83) * Validate each input signature once instead of every port on every call GraphModule._execute() checked the shape, dtype and device of every port of every node, and replayed the registered-state walk, on every forward call: a fixed ~10 us per node that made a five-node MLP at batch 32 run ~1.7x hand-written PyTorch. The first call for a given input signature (every external input's shape, dtype and device, plus training mode and autocast state) now runs the full checked program and the state check, as before. The signature is then remembered, and later calls with it compare the signature and run the modules back to back from a slot-indexed program. What was validated is invalidated on the events that can break it: - .to()/.cuda()/.double()/... (_apply) drops every signature and forces a full state re-check; - registering a parameter, buffer or submodule on any module in the graph (register_* or attribute assignment) advances a watch generation through torch's global registration hooks, filtered to modules of built graphs, so the next call re-checks state and revalidates every port; a registration during a fast call is checked at the end of that call, as before. Errors are readable: port mismatches name the node, operation and source line and show the contract in HNDL notation ("[B=32, 64]:float32 on cpu") beside the tensor that arrived; torch errors raised inside a node are wrapped as E_RUNTIME with the node's inputs and any state change that explains them; OOM errors pass through unwrapped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Document validation once per input signature CHANGELOG, README, IMPLEMENTATION and SPEC describe when contracts are checked, what invalidates a validated signature, the new error wording, and the registered-state edits PyTorch runs no hook for. The parity benchmark's docstring now describes the cheap path it actually times. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep exceptions raised inside a node as their own type, with a note Wrapping a torch error raised inside a node in HNDLError (a ValueError) broke callers that catch RuntimeError around a forward call. The original exception now propagates unchanged --- type, message and traceback --- and gains one add_note() naming the node, its operation and source line, its inputs against their contract, any registered-state change that explains the failure, the upstream port that drifted after validation, and a hint when the error came from torch.compile. An exception passing out through nested graphs keeps only the innermost node's note. HNDLError (E_RUNTIME) is kept for what HNDL's own checks find: first-call validation and state mismatches. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Add bounded for loops to the config grammar (#84) `for _ in range(N):` with a positive int literal N repeats its body, which may hold any top-level statement including nested loops. The interpreter unrolls it before resolution, so a loop yields the same node IDs, plan, semantic digest and state_dict keys as the statements written out by hand. Explicit names gain the per-level iteration suffix (block3, res1_3), node source metadata records the iterations, and configuration, resolution and runtime errors report them. The multiplied-out node count is checked against max_nodes before the first node is emitted, loops nest at most 8 levels, and every other loop, conditional and comprehension form is rejected by name. The ViT and GPT example networks now use loops for their blocks. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Close extension API gaps for operators defined outside HNDL (#86) * Close extension API gaps for operators defined outside HNDL - The global `ops` namespace resolves aliases against the registry of the active capture, so custom operators work as `ops.my_op()` under resolve_callable/network_from_callable(registry=...). A registry.ops factory whose exact declaration the capture's registry lacks now fails with E_CAPTURE and a message naming the registry to pass. - `hndl.relations` publishes the convolution arithmetic, spatial, elementwise_join and broadcast relations, plus conv_input_range, conv_transpose_input, conv_axis and conv_transpose_axis. Built-ins import from it; operators._relations re-exports it. - `hndl.testing` publishes the operator harness (check_operator and the per-check functions, example_params/operator_params for pytest). The built-in test_all_operators runs through it. pytest is imported lazily. - Tests in tests/test_extension_api.py exercise all three as an external package would. SPEC, README, ADDING_OPERATORS and CHANGELOG updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Test the harness's failure paths and address review nits - Exercise every check_build_and_run, check_declaration, check_round_trip and check_reference failure with a deliberately broken operator (or, for checks that guard HNDL itself, a patched resolver/repr/replay), asserting each message, so a check that stops checking fails the suite. - check_reference no longer crashes when a reference drops the gradient the module carries; it reports it. - check_declaration requires the class's own docstring instead of accepting one inherited from nn.Module. - check_operator accepts devices="cpu" / dtypes="float32". - A declaration missing from the registry fails with AssertionError, like one registered with a different declaration. - The E_CAPTURE hint only suggests ops.<alias> when that alias binds the same identity in the capture's registry. - conv_axis / conv_transpose_axis fail with E_CONSTRAINT for a port whose rank lacks the axis instead of leaking IndexError. - The README's my_silu declares an Example so check_operator runs on it, with a test that executes the README block through the harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Detect cleared parameters on the fast path and tighten the parity tolerance (#87) * Detect registered-state edits PyTorch runs no hook for on the fast path `del module.weight`, `linear.bias = None` and direct writes to a module's `_parameters`, `_buffers` or `_modules` fire none of torch's global registration hooks, so after the first validated call the unchecked path kept running (a layer silently without its bias) until a new signature, a move or cast, or a failing layer forced a full check. Each graph now records every module's three stores together with a copy of each whenever its state is found to match the build, and every call compares the two tuples in one C-level comparison: values are compared by identity first, so an unchanged store costs about 10-17 ns and never touches its tensors. A mismatch takes the validating path, which reports removed or added state with the usual E_RUNTIME naming the node and the change, and re-checks every port when a tensor was swapped in under the same name. .to()/casts drop the recorded copies so replaced tensors are released immediately. torch.compile tracing is unchanged: the compiling path returns before the comparison. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Tighten the parity benchmark tolerance from 2.0x to 1.5x Sixteen runs of the opt-in benchmark suite with the registered-state comparison in place measured mlp 1.00-1.13x, conv_stack 0.97-1.17x and transformer_block 1.10-1.24x of hand-written PyTorch. 1.5x leaves about 20% over the worst of those. The docstring records the runs, the machine, and the one noisy conv_stack reading (1.91x, on the code before this branch) that a loaded machine can produce. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Compare registered state without holding or comparing tensors The per-call store check compared copies of every _parameters/_buffers dict with ==, so a tensor swapped in under its name (functional_call) ran an elementwise Tensor.__eq__ and a full re-check on every call, kept the swapped tensors and their autograd graphs alive, and made copy.deepcopy fail on non-leaf, grad-tracking or vmap-batched tensors. Submodule dicts and empty tensor dicts are still compared by identity; non-empty tensor dicts are compared by length and which values are None, so removals, additions and None assignments are caught and swaps are not. __deepcopy__ skips the record, _validate records it once, and the docs list what is not detected. The parity benchmark now takes the fastest of ten interleaved rounds per side and keeps 2.0x for the arithmetic-bound convolution case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Release 0.8.0 Bump version, README status, and rename the Unreleased changelog heading for the extension API (registry-aware ops, public hndl.relations, hndl.testing) and fast-path detection of hookless state edits. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
GraphModule._execute()used to check every port of every node, and replay the registered-state walk, on every forward call. That was a fixed ~6 µs per node on top of the modules themselves. This PR checks everything once per input signature, keeps the hot path close to a plain loop of module calls, and makes failures readable.What is checked when
(shape, dtype, device), plustrainingand autocast state_state_matchesreplay as c78a43e). The signature is remembered (up to 64 per model, then the cache is cleared).module(values[i])directly.What invalidates it
_apply(.to(),.cuda(),.double(),.half(), ...) clears every validated signature and forces a full state re-check before the next call.register_parameter,register_buffer,add_module/register_module, and attribute assignment all go through torch's global registration hooks (register_module_{parameter,buffer,module}_registration_hook). A hook filtered to modules of built graphs advances a watch generation, so unrelated modules constructed elsewhere cost nothing. On the next call the graph re-checks state against the build, clears its signatures (a parameter replaced under the same name can change any shape), and revalidates every port.deepcopygives the clone a fresh cache and watches the clone's modules.load_state_dictneeds nothing: it copies in place, andassign=Truegoes throughsetattr, so the hooks fire.What the state check protects against (after reading c78a43e / 5b880dd)
c78a43e restored detection of: filling a declared-
Noneparameter/buffer/submodule slot (the lazy-init bug),self.weight = None, and replacing or grafting a submodule. In this PR:_state_matcheswalk. That's where lazy init happens.Noneover a buffer or submodule.test_every_kind_of_state_mutation_during_forward_is_rejectedpasses unchanged.del module.x, formodule.param = None(register_parameter(name, None)skips hooks), or for writes straight to_parameters/_buffers/_modules. If one of these happens after a signature has been validated, it is reported at the next full check (new signature,.to()/cast, any hooked registration) or as soon as a node then raises. The failure path always re-runs the state check, sodel model["head"].weightgives a readable "parameter 'nodes.n_head.weight' was removed" error rather than a bareAttributeError. The only silent case isparam = Noneon a module that tolerates it, e.g.linear.bias = None, which runs bias-free until the next revalidation.test_a_parameter_set_to_none_is_reported_at_the_next_revalidationpins this down. Closing it would need a per-call sweep ofstore.get(key) is not Noneover every registered tensor, about 50 ns each. I left that out to keep the hot path a straight loop, but it's a small follow-up if you want it.Errors
Port mismatch after a validated call:
An exception raised inside a node propagates as itself, with the same type, message and traceback, so
except RuntimeErrorand OOM handling keep working. It gains oneadd_note(), which the traceback prints under the message. The note is built only in theexceptpath, so there's no per-call cost:A deleted parameter (no hook fires) explains itself in the note:
An upstream node that changes its output after validation is named at the port it breaks, in the note on the downstream node's error (
linear 'head' input 'x' (from node:d/out): expected shape [B=2, 4], got [2, 2]plus an explanation). An exception passing out through nested graphs (a built graph used as a node of another) keeps only the innermost node's note. Errors that come from dynamo or inductor get a line saying so.HNDLError/E_RUNTIMEis kept for what HNDL's own checks find: first-call validation and registered-state mismatches.Registered state changed:
Other details:
B=3 is this call's batch size, read from input 'x'.Behavior changes to review
HNDLError. That brokeexcept RuntimeErrorcallers and was replaced.)expected dtype float64, got float32(HNDL notation) instead oftorch.float64. The dtype assertions intests/test_torch.pywere updated for this.Tests
test_deepcopy_keeps_lookup_moves_and_the_state_consistency_checkused to tamper with_state_programafter a successful call and expect the next call to notice. With per-call checking gone that no longer holds, so I rewrote it to do a real mutation on the clone (clone["head"].ghost = nn.Parameter(...)). It now asserts that the clone reports it, that the original keeps running, and that the original still catches its own mutations.torch.float64becamefloat64(see above)..to(), and that unrelated module construction does not invalidate); a broken input after the first call (exact message); a second input's batch hint; a torch error inside a node on the validating call and on the fast path (type(...) is RuntimeError, exact note); OOM keeping its type; anHNDLErrorinside a node keeping its message and gaining a note; one note only through nested graphs; upstream output drift named in the note; five mutations after the first call (grafted submodule, added parameter and buffer set toNone→HNDLError; deleted parameter →AttributeErrorwith a state note; wrong-shaped parameter →RuntimeError), each reported on every call; theparam = Nonegap; state registered during a later forward; andtorch.compile(backend="eager", fullgraph=True)tracing the checked program with no graph breaks.PYTHONPATH=src python -m pytest -q -m "not network and not benchmark"gives 3586 passed, 80 skipped.Benchmarks
CPU,
torch.set_num_threads(1), torch 2.14. The machine was shared (load average 16–25), so each row is the best of 9 interleaved hndl/hand-written measurements, with ranges over 3 alternating old/new process runs. "Before" isorigin/developat 7f7cd3d.The fixed cost is now ~2 µs per call (the signature, a dict lookup and the output dict) plus ~0.35 µs per node. The ~35 µs left on
transformer_blocksits inside that single operator's own implementation compared withnn.MultiheadAttention, not in per-node validation.pytest -m benchmark tests/benchmark -q -spasses before and after (28 passed, 1 skipped). Parity ratios from one run each: mlp 1.22x → 1.09–1.14x, conv_stack 0.74x → 1.06–1.07x (noise), transformer_block 1.13x → 1.08–1.09x. All compile-compat networks compile as before.moe_transformeris skipped in both runs by an inductorCppCompileError, which is environmental.TOLERANCE: I did not change it. The MLP now runs ~1.2x at batch 32, so the parity suite could time the MLP at batch 32, the case this issue was about, and TOLERANCE could come down from 2.0 to about 1.5. I haven't done that because of the shared-machine noise on CI hosts. Your call.
Other notes
import hndl.torch, for the whole process. Each is aWeakValueDictionarylookup that runs only on registrations, never in forward. ReturningNonemeans they never alter a registration.torch.compiler.is_compiling(),_executetraces the full checked program and state check exactly as before (dynamo turns them into guards), and skips the signature cache.## Unreleased, README, IMPLEMENTATION.md, SPEC.md §3/§8/§11/§13, and the parity benchmark docstring.🤖 Generated with Claude Code