Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 54 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,60 @@ Versioning: [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

### Added — `check_boundary_types`, the 13th rule

A `GraphSpec`'s declared `input_type` / `output_type` are now compared against the variables the
edges at START and END actually carry. They were never checked, and they go straight to
`GraphBuilder`:

```python
class Lying(Greeting): # every edge in Greeting carries str
input_type, output_type = int, int

Lying().coherence_check() # [] — before
Lying().render(...).run_sync(inputs=" a b ") # 'Hello, a b!' — a str
```

**A dead field would have been the small version.** `_port_type` reads those two fields as the
ORACLE for `check_subgraphs`, so a check that does run — whose whole job is *"the child fits the
node"* — was comparing a parent node's contract against the child's unverified claim about
itself. Measured: a child declaring `output_type=int` while every edge in it carries `str` passed
against a parent node declaring an int output, and the graph returned `'HELLO'`. Issue #21.

Assignability, not identity, and the direction differs per side: the input may be **widened** on
the way in (`input_type=list[int]` into an edge carrying `numbers: list` — `examples/parallel.py`
does this), the output **narrowed** on the way out (`report: str` reaching `output_type=object` —
`examples/ladder/stage10_no_basenode.py` does this).

⚠️ **The default `type(None)` is a claim, not an absence.** A design that never declares a
boundary and then wires a `str` across it is now reported, because that default reaches the
engine as the graph's real signature.

### Fixed — `_produces` called two spellings of one type undecidable

`list[int] is list[int]` is `False`, so identity alone reported *not decidable* for literally the
same type; and a parameterised alias was never compared against a bare declared type, although
`list[int]` plainly IS a `list`. Both now decide. This can only turn an undecidable into a
verdict — it cannot manufacture a finding where there was none — and it is why
`check_boundary_types` is clean on all 28 boundary crossings in `examples/` rather than printing
a permanent `NOT CHECKED` line on `parallel.py`.

`check_variable_types` reads the same helper and gains the same decidability.

⚠️ Only a **runtime class** origin is compared. `get_origin` is `typing.Literal` for
`Literal['ok']` and `typing.Annotated` for `Annotated[int, 'tag']`, and `issubclass` on either
raises — which the first cut of this did, through `coherence_check()`, a method documented
*"Never raises."* Those wrappers stay undecidable rather than being unwrapped; nothing has needed
unwrapping yet.

### Fixed — `_type_name` rendered `list[int]` and `list` identically

Its docstring said generic aliases have no `__name__`. Since 3.10 they do, and it is the bare
origin — so the bug was a WRONG name rather than a missing one, and a finding comparing those two
types read as a complaint that `list` is not `list`. Found while writing the message for the
check above, which compares exactly that pair.


## [0.3.0] — 2026-10-01

### ⛔ Breaking — `check()` is renamed to `coherence_check()`
Expand Down
34 changes: 27 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ shape and its data contracts as **data**, before any step exists — which is wh
of defect findable:

```python
spec.coherence_check() # 12 well-formedness rules, 7 of them with nothing implemented
spec.coherence_check() # 13 well-formedness rules, 8 of them with nothing implemented
spec.diagram() # a picture of the same declaration
spec.render(strategy) # refuses outright if anything blocks
```
Expand All @@ -39,7 +39,7 @@ Four problems, and the same declaration answers all four:

| | |
|---|---|
| **An agent's output works and is incoherent.** Each piece is locally fine; the whole does not add up. | `coherence_check()` — 12 [well-formedness rules](docs/glossary.md#well-formedness-rule), 7 needing nothing implemented |
| **An agent's output works and is incoherent.** Each piece is locally fine; the whole does not add up. | `coherence_check()` — 13 [well-formedness rules](docs/glossary.md#well-formedness-rule), 8 needing nothing implemented |
| **A reasoning strategy cannot be asserted correct — only compared.** There is no right answer to diff against, so "better" is an empirical question. | [`eval_battle()`](docs/glossary.md#battle) — same cases, same evaluators, plus a replicate arm as the [noise floor](docs/glossary.md#noise-floor) |
| **Complexity grows unless pieces are reused.** Two arms that differ in one stage should say so, not be two files. | the [data language](docs/glossary.md#deep-embedding): declare a role once, bind it many ways; `SubgraphBinding` reuses a whole child design as one node |
| **You cannot see what you built.** | `diagram()` and `diff_diagram()`, from the declaration alone |
Expand Down Expand Up @@ -116,8 +116,27 @@ One workflow — normalize a name, then compose a greeting from it:
Desired behaviour: preserve the name's words, trim surrounding whitespace, collapse repeated
internal whitespace, return `Hello, {name}!`.

Two strategies disagree about how much of that `normalize` does. `compose` is the same function
in both, so the comparison diagram highlights the one node that varies:
That table is the whole declaration, and it draws itself. **Nothing is implemented at this
point** — no `normalize` body, no `compose` body, no strategy, nothing an agent has written:

```mermaid
flowchart TD
START([START])
normalize["normalize"]
compose["compose"]
END([END])
START -- raw_name --> normalize
normalize -- clean_name --> compose
compose -- greeting --> END
```

Bare boxes, because nothing is bound to them yet. This picture and `coherence_check()` are what
you review *before* asking an agent for a line of code — which is the one thing a drawing taken
from a built graph cannot do, since building it requires the code to already exist.

Two strategies disagree about how much of that `normalize` does. Same graph, two implementations
bound: `compose` is the same function in both, so the comparison greys it and highlights the one
node that varies:

```mermaid
flowchart TD
Expand Down Expand Up @@ -194,7 +213,7 @@ Everything it produces goes to the terminal; no files are written. Excerpt:
normalize_spaces 1.00
```

Two mermaid blocks go past on the way: the specification, and the comparison above. A browser
Both mermaid blocks above go past on the way — the specification, then the comparison. A browser
viewer is available as a separate process — `uv run python3 -m workflow_workbench.cli serve`, see
[`serve.py`](workflow_workbench/serve.py) — and nothing in the quickstart needs it.

Expand All @@ -219,9 +238,9 @@ missed.
<summary><strong>Every rule — generated from each check's own docstring</strong></summary>

<!-- rules:start -->
**12 rules.** `coherence_check()` returns one finding per violation and an empty list for a clean design; `render()` refuses on any finding that blocks.
**13 rules.** `coherence_check()` returns one finding per violation and an empty list for a clean design; `render()` refuses on any finding that blocks.

**7 need no implementations at all** — runnable the moment `nodes` and `edges` are written.
**8 need no implementations at all** — runnable the moment `nodes` and `edges` are written.

| check | rule |
|---|---|
Expand All @@ -231,6 +250,7 @@ missed.
| `check_step_arity` | A step body receives exactly ONE value, so a node cannot consume two inputs at once. |
| `check_decisions` | `when` appears exactly on the edges leaving a decision, and nowhere else. |
| `check_transform_edges` | A transform edge is fixed (`apply=`) or a variation point (bound) — exactly one. |
| `check_boundary_types` | The graph's declared `input_type` / `output_type` match what crosses START and END. |
| `check_fan_out_rejoins` | Everything a fan-out produces must reach a join before it reaches END. |

**5 more once a strategy exists**, checking the implementations against the roles they fill.
Expand Down
Loading
Loading