Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,30 @@
`ResolvedPlan.from_json` drops its 16 MiB input cap and duplicate `max_nodes`
check; revalidation still applies `max_nodes`.

- **Contracts are checked once per input signature, not on every call.** The
first forward whose inputs have a given shape, dtype and device (in a given
training mode and autocast state) still checks every node's ports and that
no module created or removed registered state; later calls with the same
signature compare only the inputs and run the layers back to back. Moving
or casting the model, or registering a parameter, buffer or submodule on
any of its modules, makes the next call check everything again. The fixed
per-call cost of a five-node MLP drops from ~30-40 us to ~6-9 us: at batch
32 it runs ~1.2x hand-written PyTorch instead of ~1.8x.
Errors say more: `E_RUNTIME` names the node, its operation and source line,
and prints the contract in HNDL notation beside the tensor that arrived
(`expected [B=32, 64]:float32 on cpu` / `got [32, 63]:float32 on cpu`);
dtypes are spelled `float32`, not `torch.float32`. An exception raised
inside a layer keeps its type, message and traceback (a `RuntimeError` or
out-of-memory error is still caught as one) and gains a note, via
`add_note`, naming the node, its operation and source line, its inputs
against their contract, and any registered-state change that explains it;
an exception passing out through nested graphs carries only the innermost
node's note. Registered
state edits PyTorch runs no hook for --- `del` of a registered name,
`module.param = None`, direct writes to `_parameters`/`_buffers`/`_modules`
--- are reported at the next full check or when a layer then fails, not on
the very next call.

- Add `broadcast_mul`, `coordinate_grid`, `fourier_features`, and `grid_sample`
for style-conditioned coordinate renderers composed in HNDL. Fourier tables
and coordinate grids are persistent buffers; sampling follows PyTorch semantics.
Expand Down
19 changes: 19 additions & 0 deletions IMPLEMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,25 @@ graph input is integer. Edge dtypes are checked at resolution (`E_DTYPE`), so
an integer tensor cannot reach a floating-point port. Reduced precision is
qualified on CUDA.

A built model checks its contracts once per input signature, not on every
call. The first call whose inputs have a given shape, dtype and device (in a
given training mode and autocast state) checks every node's ports and that no
module created or removed registered state; later calls with the same
signature compare only the inputs and run the layers back to back. Moving or
casting the model, or registering a parameter, buffer or submodule on any of
its modules, makes the next call check everything again. Failures are
`E_RUNTIME` errors that name the node, its operation and its source line and
print the contract beside the tensor that arrived, for example
`linear 'head' input 'x' (from node:hidden/out): expected shape [B=32, 64], got [32, 63]`;
an exception raised inside a layer propagates unchanged (a `RuntimeError`
stays a `RuntimeError`) with a note, printed under its message in the
traceback, that names the node, its operation and source line, the tensors it
received, and any registered-state change that explains it. PyTorch runs no registration hook for `del` of a
registered name, for assigning `None` over a registered parameter, or for
writing `_parameters`/`_buffers`/`_modules` directly, so those edits are
reported at the next full check or when a layer then fails, not on the very
next call.

Built-in unary operations take an optional leading tensor or `x=`. Custom
unary operations use their declared input-port keyword. Every operator's
arguments, defaults, bounds, shape relation, and examples are listed in
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,7 @@ last_layer = model[-1]
features = model[:2] # nn.Sequential sharing these layers
```

Unnamed operations receive IDs such as `n0`; `model["n0"]` and `model[0]` return the same module. Optional names appear in the branching example below. Slices reuse their parameters, so training a slice also updates the original model. The complete model retains its resolved shape contract; a slice is a regular PyTorch sequence. Standard `state_dict()`, `train()`, and `eval()` remain available. Moving the model with `.cpu()` or `.cuda()` and casting it with `.double()`, `.half()`, `.bfloat16()`, `.float()`, or `.to(dtype=...)` retarget the runtime checks too, so a cast model takes tensors of its new compute dtype; ports declared as integers, such as token ids, keep the dtype the plan declared. And `copy.deepcopy(model)` returns an independent model --- its own parameters and buffers, the same trainability flags and training mode, and no draw on the random state --- which is what a moving-average copy of a model needs. The resolved plan is immutable, so the copy shares it. To store a model, save `model.plan.to_json()` next to `torch.save(model.state_dict())` and rebuild it; pickling the module itself is not supported.
Unnamed operations receive IDs such as `n0`; `model["n0"]` and `model[0]` return the same module. Optional names appear in the branching example below. Slices reuse their parameters, so training a slice also updates the original model. The complete model retains its resolved shape contract; a slice is a regular PyTorch sequence. Standard `state_dict()`, `train()`, and `eval()` remain available. Moving the model with `.cpu()` or `.cuda()` and casting it with `.double()`, `.half()`, `.bfloat16()`, `.float()`, or `.to(dtype=...)` retarget the runtime checks too, so a cast model takes tensors of its new compute dtype; ports declared as integers, such as token ids, keep the dtype the plan declared. Those checks run once: the first call with a given input shape, dtype and device checks every layer's inputs and outputs against the plan, and later calls with the same inputs run the layers directly, at close to hand-written speed. A mismatch is an `E_RUNTIME` error naming the layer, its operation and its source line, with the expected and actual shapes; an error raised inside a layer keeps its own type and gains a note with the same details. And `copy.deepcopy(model)` returns an independent model --- its own parameters and buffers, the same trainability flags and training mode, and no draw on the random state --- which is what a moving-average copy of a model needs. The resolved plan is immutable, so the copy shares it. To store a model, save `model.plan.to_json()` next to `torch.save(model.state_dict())` and rebuild it; pickling the module itself is not supported.

Initial weights use PyTorch’s normal random state. For repeatable initialization in the same environment, call [`torch.manual_seed(7)`](https://docs.pytorch.org/docs/stable/notes/randomness.html#pytorch-random-number-generator) before constructing the network.

Expand Down
8 changes: 4 additions & 4 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ A tensor contract is an ordered tuple of dimensions including batch, interpreted

Batch passes through every operator unchanged except the two that move tensors across axis 0. Joining along that axis stacks examples, so the result holds several batches at once: the entry is then written `"k*B"`, meaning `k` times the plan's batch. `"B"` is one batch and has no other spelling — `"1*B"`, `"0*B"`, `"B*2"` and leading zeros are rejected with `E_SCHEMA` — so a plan that never touches the batch axis carries exactly the entries, and therefore the digests, it carried before multiples existed. Multiples run from 2 to 1024; exceeding that fails with `E_RESOURCE`. A `k*B` port requires `k` times the runtime batch of the call, checked like any other extent. External contracts — `input_shape` and `output_shape`, named or not — stay one plan batch: a graph must chunk a joined tensor back before publishing it, and a declared `"2*B"` contract fails with `E_SCHEMA`. Internal port contracts in a saved plan may carry multiples, and a malformed entry at axis 0 fails `E_SCHEMA` on load.

A plan carries a compute `dtype` of `float32` (the default), `float16`, or `bfloat16`, and an `input_dtype` that defaults to the compute dtype. The graph input may instead be an integer contract — `int64`, `int32`, or `bool` — for token ids and masks; a floating-point `input_dtype` must equal the compute dtype. Operator ports may declare their own dtype, including `any` for a port that accepts whatever its producer carries. Every edge's dtype is checked once during resolution and a mismatch fails with `E_DTYPE`, so an integer tensor cannot reach a floating-point port. The backend constructs parameters in the compute dtype and checks each port's dtype, shape, and device at runtime.
A plan carries a compute `dtype` of `float32` (the default), `float16`, or `bfloat16`, and an `input_dtype` that defaults to the compute dtype. The graph input may instead be an integer contract — `int64`, `int32`, or `bool` — for token ids and masks; a floating-point `input_dtype` must equal the compute dtype. Operator ports may declare their own dtype, including `any` for a port that accepts whatever its producer carries. Every edge's dtype is checked once during resolution and a mismatch fails with `E_DTYPE`, so an integer tensor cannot reach a floating-point port. The backend constructs parameters in the compute dtype and checks each port's dtype, shape, and device at runtime, once per input signature (§11).

Both frontend APIs accept `input_shape` and `output_shape`, both including batch, plus optional `dtype` and `input_dtype` keywords. Structured graph contracts use the same tuples.

Expand Down Expand Up @@ -631,7 +631,7 @@ linear()

Declarations are trusted code, for custom operators as much as for built-ins: relation functions and module constructors run in the host's process. A registered relation is allowed to be arbitrary Python, and custom operators may therefore express any rule a built-in can. What remains untrusted is data: configuration source and saved plans can name only an alias or identity the host already registered, can never import a class, a module path, or a callback, and are bounded by the same parser and resolver limits. A plan stores `example.silu@1` plus concrete arguments and port shapes, not executable Python, and building or restoring it without that registration fails with `E_STATE_VERSION`.

The constructor receives every declared argument by keyword, plus any shape symbol it names as a keyword-only parameter (`*, D`) and, when requested by name, `input_shapes` / `output_shapes` mappings of resolved port shapes. It runs under `torch.device(device)`, so tensors may be created normally; parameters are then cast to the plan's compute dtype. `forward` receives tensors positionally in declared input-port order and returns one tensor, or a tuple/list in declared order or a dictionary with exactly the declared names for multiple outputs. HNDL validates each result's shape, dtype, and device. Within a built model each node owns an independent module instance and registered state; reusing a module or registered storage across nodes is rejected. Sharing a producer tensor across branches, or obtaining a slice of an already built chain, remains supported. A genuine unary chain supports indexing and shared sequential slices even when its ports use names other than `x` and `out`.
The constructor receives every declared argument by keyword, plus any shape symbol it names as a keyword-only parameter (`*, D`) and, when requested by name, `input_shapes` / `output_shapes` mappings of resolved port shapes. It runs under `torch.device(device)`, so tensors may be created normally; parameters are then cast to the plan's compute dtype. `forward` receives tensors positionally in declared input-port order and returns one tensor, or a tuple/list in declared order or a dictionary with exactly the declared names for multiple outputs. HNDL validates each result's shape, dtype, and device on the first call for each input signature (§11). Within a built model each node owns an independent module instance and registered state; reusing a module or registered storage across nodes is rejected. Sharing a producer tensor across branches, or obtaining a slice of an already built chain, remains supported. A genuine unary chain supports indexing and shared sequential slices even when its ports use names other than `x` and `out`.

Saved plans contain concrete shapes and arguments plus exact operator identities and versions. Declarations remain in the explicitly supplied registry; artifacts cannot import them. Restoring a custom plan revalidates its concrete equations against that registry, and any mismatch fails with `E_INTEGRITY` rather than being re-inferred (§12 lists the two kinds of missing value restoration fills in). An operator must raise its semantic version when it changes a shape relation, an argument schema, its numerical meaning, or its parameter layout.

Expand Down Expand Up @@ -727,7 +727,7 @@ The **artifact digest** covers the complete saved plan except its own digest fie
- Construction revalidates the plan, then runs an allocation-free `meta` pass over every node to bound registered storage before allocating anything (§7). The build receipt records `torch_version`, `device`, `dtype`, `initialization_seed`, `seed_mode`, and `state_bytes`.
- Parameters and buffers are constructed in the plan's compute dtype. A node is constructed under the target device so constructors may allocate normally.
- A `GraphModule` validates declared external input contracts in `forward(**inputs)`, which accepts exactly the declared input names, and returns a dictionary keyed by public output names in declaration order, including for a single output. A missing, repeated, or undeclared runtime input fails with `E_BINDING` before any node runs. Both network facades instead accept the declared inputs positionally in declaration order, by keyword, or both, and return the selected output tensor for a single public output or the same dictionary for several, preserving the contract checks even for a branched graph.
- Every port is checked for shape, dtype, and device before and after applying its node, and a module that creates or removes registered state during forward fails with `E_RUNTIME`.
- Contracts are checked once per input signature: the first call whose external inputs have a given shape, dtype, and device — in a given training mode and autocast state — checks every port for shape, dtype, and device before and after applying its node, and then checks that no module created or removed registered state; a module that did fails with `E_RUNTIME`. Later calls with a validated signature compare only that signature and then run the nodes. Moving or casting the module, and registering a parameter, buffer, or submodule on any module of the graph, invalidate what was validated: the next call re-checks registered state against the build and every port again. An exception raised inside a node's module propagates unchanged, keeping its type, message, and traceback, with one note (`add_note`) naming the node, its operation, and its source position, the tensors it received, and any change to registered state that explains it; an exception passing out through nested graphs keeps only the innermost node's note. `E_RUNTIME` reports what the backend's own checks find.
- Modules register once under `nodes.n_<node_id>`; the prefix avoids collisions with module attribute names. Stateless nodes keep execution/diagnostic identities without state entries. Two nodes sharing a module instance, tensor, or storage fail with `E_REGISTRY`.
- Nodes run in stable topological order, breaking ties by declaration order. The order is saved in the plan.
- Parameters/buffers are fully materialized before optimizer or distributed setup. No first-forward parameter creation is allowed.
Expand Down Expand Up @@ -837,7 +837,7 @@ Failures must expose a stable code, source/node/field location, affected constra
| `E_INTEGRITY` | A saved plan's digests or concrete equations do not match its contents |
| `E_RESOURCE` | Source size, nesting depth, graph, state-size, or resolution budget is exceeded |
| `E_STATE_VERSION` | Saved state requires an unavailable compatible operator implementation or plan version |
| `E_RUNTIME` | A runtime tensor violates its declared shape, dtype, or device, or the device is unavailable |
| `E_RUNTIME` | A runtime tensor violates its declared shape, dtype, or device; registered state no longer matches the build; or the device is unavailable |

Configuration failures report original source line/column locations through the dedent map. Native capture locations are best effort and may use call-site information, but must always identify the affected node/field without requiring function-source inspection. `print(model)` and `repr(model)` must include every layer, its ID/operator, and complete input/output shapes without executing the network. `print(plan)` exposes the equivalent named-port table without a backend. `plan.describe()` must additionally expose provenance sufficiently to explain why a field changed between separately resolved specifications.

Expand Down
Loading
Loading