Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
51 commits
Select commit Hold shift + click to select a range
78460c1
PACT: object model & metered universality — recursion fuel, agent sha…
jithinAB Aug 13, 2026
06ac849
PACT: values, argument-taking patterns, and the net that says nothing…
jithinAB Aug 22, 2026
2b3b735
PACT: a carried file says what is in it
jithinAB Aug 22, 2026
bf9a29b
PACT: work handed to an agent by name — the dynamic-bottom rule
jithinAB Aug 22, 2026
498875c
PACT: what the audit found — figures and patterns, hardened
jithinAB Aug 22, 2026
295db17
PACT: what an author allows is what a cycle can do
jithinAB Aug 22, 2026
f1fbdc8
PACT: a program is declared, and checking one runs nothing
jithinAB Aug 22, 2026
c18fa28
PACT: a program a run cannot start is said out loud
jithinAB Aug 22, 2026
26993a1
PACT: a program an agent may use without wrapping it in a tool
jithinAB Aug 22, 2026
3aace86
PACT: what the second audit found — two of them mine
jithinAB Aug 22, 2026
ffd1540
PACT: a figure written where one cannot stand is said out loud
jithinAB Aug 22, 2026
fcc412c
docs: the plan says what was built, and where it was wrong
jithinAB Aug 22, 2026
3985b5d
PACT: a desk that remembers what it was told
jithinAB Aug 22, 2026
6a62b4e
PACT: a grader the author carries
jithinAB Aug 22, 2026
30e4efc
PACT: an answer a person types is checked before the run goes on with it
jithinAB Aug 22, 2026
6c8aa91
PACT: a result shortened before the model reads it
jithinAB Aug 22, 2026
2269090
PACT: a rule that rewrites what it is given
jithinAB Aug 22, 2026
c4f0f31
PACT: a stage that decides where to go
jithinAB Aug 22, 2026
ae124ba
PACT: a stage that runs what the model wrote
jithinAB Aug 22, 2026
357773a
PACT: a tool an agent wrote for itself
jithinAB Aug 22, 2026
b2e1184
PACT: an author's reserved space is not a template, and a list is a f…
jithinAB Aug 22, 2026
7a88f47
PACT: where a figure landed, recorded because it cannot be recomputed
jithinAB Aug 22, 2026
771de2b
PACT: a rewriting rule no longer deletes the hiding rules beside it
jithinAB Aug 22, 2026
f225aa4
PACT: every door to a carried program is held to the rules that door …
jithinAB Aug 22, 2026
3cdd582
PACT: the eighth part of the boundary reaches somebody, or says it do…
jithinAB Aug 22, 2026
27e2632
PACT: reading somebody else's tree finishes, and costs what the tree …
jithinAB Aug 22, 2026
bfa2663
PACT: memory really is a variable a run can read and write
jithinAB Aug 22, 2026
d1b7911
PACT: a document built from a pattern says which pattern, and three p…
jithinAB Aug 22, 2026
b463ba4
PACT: a written rule needs a person, whichever direction it is edited
jithinAB Aug 22, 2026
1948a9f
PACT: a snippet and what the room printed meet the rules like everyth…
jithinAB Aug 22, 2026
2610b95
PACT: the laundering channel was already closed, and now it is pinned
jithinAB Aug 22, 2026
d2f6f3f
PACT: the tier discipline the no-code badge rests on is checked
jithinAB Aug 22, 2026
f30442a
PACT: the room runs what the model wrote, the transcript shows what w…
jithinAB Aug 22, 2026
3950614
PACT: one declared room is not a claim about every room
jithinAB Aug 22, 2026
231f394
PACT: a code fence is a sample, not a document
jithinAB Aug 22, 2026
869fb60
PACT: a program a suite grades with is a program something names
jithinAB Aug 23, 2026
f3a1396
PACT: a rewriting rule rewrites on a real run, and says so when it ca…
jithinAB Aug 23, 2026
0d9c7f0
PACT: a checked answer is checked on a real run, not refused on princ…
jithinAB Aug 23, 2026
81d9fac
PACT: a run knows about the programs it reaches without a tool
jithinAB Aug 23, 2026
f79f2ad
PACT: parking is not a reason for the rules to stop
jithinAB Aug 23, 2026
993b966
PACT: a name two collections could answer to is read the way the sche…
jithinAB Aug 23, 2026
638a127
PACT: four documents that said something the code stopped doing
jithinAB Aug 23, 2026
22a688d
PACT: the workspace floor is what the words meet last
jithinAB Aug 23, 2026
ebfb9bd
PACT: a rewriter cannot slip past the stop rule beside it
jithinAB Aug 23, 2026
bbac17f
PACT: the reading budget is a total, not a per-file allowance
jithinAB Aug 23, 2026
d6b2b7d
PACT: granting a carried body the outside world is not silence
jithinAB Aug 23, 2026
1ba4f31
PACT: the second port is told what a tool is
jithinAB Aug 23, 2026
9c4d9a6
PACT: an author can see where the rules in a skill begin
jithinAB Aug 23, 2026
c02065c
PACT: a reworded heading is reported as a reworded heading
jithinAB Aug 23, 2026
5bd53e0
PACT: a carried program is not a written procedure, and a registry is…
jithinAB Aug 23, 2026
0387260
PACT: the front page says what the format grew
jithinAB Aug 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 1 addition & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

194 changes: 161 additions & 33 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,23 +70,33 @@ build step that *executes author code*. PACT's tree is executable as-is.
|---|---|
| **Design** | Thesis, 28 binding decisions, FRD (120 requirements), implementation plan — complete |
| **Research** | 14 source-grounded studies, ~15,750 lines, over 140 repos (~15 GB) + 57 papers |
| **Code** | Loader, diagnostics, schema engine, CLI, harness, resolver, evals, SLO — **2028 tests (744 Rust + 1284 adapter), clippy clean, TypeScript type-checked** |
| **Code** | Loader, diagnostics, schema engine, CLI, harness, resolver, evals, SLO — **3188 tests (1095 Rust + 2093 adapter), clippy clean, TypeScript type-checked** |
| **Adapters** | **All 7 named targets**, proven against one shared conformance suite |

### What works today

```bash
./scripts/test-all.sh # 2028 tests, Rust + 7 adapters, fully offline
./scripts/test-all.sh # 3188 tests, Rust + 7 adapters, fully offline
```

**Framework portability is proven, not asserted.** One folder — loaded by the
Rust CLI — executes over **all seven targets** and produces *byte-identical
traces, tool sequences and model-call counts*. That claim is bounded and the
bound is written down: it covers `instructions:`, `tools:`, `skills:`,
`knowledge:`, `team:`, `answers-with:`, the `loop:` and the `limits:` ceilings,
and **not** the nine
governance keys the TypeScript port reports on `unenforced` — see
and **not** the ten governance keys the TypeScript port reports on
`unenforced` — nine written at the top of the agent's own file, plus the `asks:`
line on a loop stage that stops to ask a person, which stops both ports in the
same stage and is put to somebody only by the Python one — nor the documents
under `knowledge/`, which nothing on that port
can look anything up in and which it names on `unretrieved` for every set an
answer did not come from — see
[§7.28 *What "byte-identical across all seven targets" covers, and what it does not*](docs/20-ARCHITECTURE-DRAFT.md).
That ten counts the *excluded keys*, not the lines a run prints: the same channel
also carries `team:` — which **is** inside the claim, since the names are offered
to the model, and says only that asking one comes back as an error here — and one
line per `limits:` key this port does not read, each under its own `limits.`
prefix.
Each takes the framework's *lowest* seam, so none of them gets to own the loop:

| Target | Seam taken |
Expand Down Expand Up @@ -115,15 +125,15 @@ Each publishes a capability lattice, and the lattice records real differences
rather than flattering uniformity:

```
model_cal | tool_call | text_with | parallel_ | streaming | durable_r
reference native | native | native | native | unsupport | unsupport
pydantic-ai native | native | native | native | emulated | unsupport
langgraph native | emulated | native | emulated | emulated | native
langchain native | native | native | native | emulated | unsupport
autogen native | native | emulated | native | emulated | unsupport
openai-agents native | native | native | native | emulated | unsupport
anthropic native | native | native | native | emulated | unsupport
vercel-ai native | native | native | native | emulated | unsupport
model_cal | tool_call | text_with | parallel_ | streaming | durable_r | connected
reference native | native | native | native | unsupport | unsupport | unsupport
pydantic-ai native | native | native | native | emulated | unsupport | native
langgraph native | emulated | native | emulated | emulated | native | unsupport
langchain native | native | native | native | emulated | unsupport | unsupport
autogen native | native | emulated | native | emulated | unsupport | unsupport
openai-agents native | native | native | native | emulated | unsupport | unsupport
anthropic native | native | native | native | emulated | unsupport | unsupport
vercel-ai native | native | native | native | emulated | unsupport | unsupport
```

Seven rows for seven targets, plus `reference` — the framework-free control arm,
Expand All @@ -135,6 +145,13 @@ LangGraph is the only target with native durable resume. AutoGen cannot carry an
assistant sentence *and* tool calls in one result — its text rides in `thought`,
so the value survives and the difference is declared instead of hidden.

`connected` is `connected_tools`, truncated like the other headings: whether a
tool's `connect:` line reaches the system it names. Pydantic AI is the only
target that can — `mcp_bridge` turns a `connect:` into that runtime's own client
— and every other target binds a model and nothing else, so a `connect:` tool
reaches the model as a name and the call comes back `error: no tool named …`.
Declared here rather than discovered there.

**The HITL kill test passes on all six Python targets.** The Vercel target has
no durable resume and declares `durable_resume: unsupported`, so it is not in
that suite — counting the framework-free control arm as the seventh made "all
Expand All @@ -144,15 +161,16 @@ each tool executes exactly once, the decision is honoured, and the resulting
history is identical across frameworks. The architecture named this the one
fixture that decides everything: *"if it passes on both adapters, D12 is proven."*

**Model portability is measured — but has no shipped command yet.** The
**Model portability is measured, and `--choose-model` is the door onto it.** The
resolver filters the catalogue on `needs:`, runs the author's eval cases against
each candidate strategy, refuses when none passes, and names the cheapest that
would. All of that is real and tested. Its only caller today is
`tests/test_model_portability.py`: `scoring.py` imports `load_catalogue` and
`needs_of` from that module and not `resolve()` itself. So the report below is
produced by the test suite, and a user can currently ask *"does model X pass?"*
but not *"which model should I use?"* — tracked as A2 in
[the gap register](docs/70-PRODUCTION-GAP-REGISTER.md).
would. It has a shipped caller: `scoring.py:1005` calls `resolve()`, behind the
`--choose-model` flag its own usage text advertises, so a user can now ask both
*"does model X pass?"* and *"which model should I use?"*. That closes A2 in
[the gap register](docs/70-PRODUCTION-GAP-REGISTER.md), and the report below is
the shape that command prints. The transcript itself is still produced by the
test suite rather than pasted from a run, because it needs a machine serving
those models; what is shown is the renderer's own output.

```
PORTABILITY: PASS for claude-haiku-4-5 (agent Refund Desk, strategy decomposed)
Expand Down Expand Up @@ -232,6 +250,98 @@ Every diagnostic carries **where, what, why, and how** — a fix is required by
the constructor's signature, so an unactionable error cannot be built. The
intended reader cannot write code, so a message they cannot act on is a defect.

Not everything a check has to say is a complaint. A **note** says something true
about a document that is not wrong with it, and `--deny-warnings` stays green
for one — a fact printed as a problem teaches an author to stop reading the
output, which is the one thing this format cannot afford:

```
note: 'refund-policy' holds 5 written rules under `## Rules` — lines a person
has to approve before they change, whoever or whatever proposes it.
fix: Nothing to do. Move a rule out from under `## Rules`, or reword that
heading, and it stops being one — which is how you move this boundary.
```

---

## The object model, and exact logic

Two things an author reaches for once a workspace stops being one desk: *"these
five desks are the same shape"* and *"this bit must be exactly right, not
approximately right."* Both are authored in YAML, and both resolve to nothing —
which is the point.

**One shape, several desks.** `expects:` declares the holes; `with:` fills them;
`based-on:` inherits by shallow merge; `base: yes` marks a shape nothing runs.
Together they are inheritance, encapsulation and polymorphism, written the way
the rest of the format is written.

```yaml
# agents/desk-pattern/agent.yaml — a pattern, not a desk
expects:
domain: { shape: text, help: what this desk answers questions about }
daily-cap: { shape: money, help: the most one request may cost }
description: A desk that answers questions about <domain>.
limits:
cost-per-request-under: <daily-cap>

# agents/refunds/agent.yaml — one of the desks
based-on: desk-pattern
with: { domain: refunds, daily-cap: 0.05 USD }
```

**The pattern is resolved and removed.** A tree written this way and one written
out longhand are not similar — they are the same document, down to the digest.
Both shipped fixtures prove it, and you can run this yourself:

```bash
pact discover tests/trees/two-desks-one-pattern # sha256:554f1ac71be9a23f…
pact discover tests/trees/two-desks-longhand # sha256:554f1ac71be9a23f…
```

`values:` does the same for a single figure — a spend cap written once and used
in three files digests identically to the same cap typed three times
(`one-figure-in-three-places` against `one-figure-longhand`). Nothing below the
loader ever learns these features exist, which is what keeps every downstream
claim — portability, digests, conformance — true of trees that use them.

Because both are erased, `pact waits` records **where each one landed**, since
by then the reference is gone and nothing else could say.

**Exact logic, when getting it right matters more than reading it.** A carried
program is a file in the folder with a declared engine, a determinism promise
and a fuel ceiling. It runs in a locked room the host supplies, and a person
consents before it runs at all:

```bash
pact waits tests/trees/a-desk-with-a-program
# waits: [('may-we-run', 'needs-permission')]
```

Seven lines reach one: a tool's `program:`, an agent's `uses:`, an action's
`projects-with:`, a question's `checked-by:`, a stage's `decided-by:`, an eval's
`uri: program:<name>`, and a rewriting interceptor sentence. Each is held to the
rules its own line claims — a router and a rewriter must be `pure`, because
where a run goes and what it says have to be the same twice; a checker and a
grader are deliberately not, because looking something up is what they are for.

A stage may also write its own code and have the room run it (`does: run-code`,
the sixth of FR-6.1.5's loop patterns). The snippet lands in the transcript
verbatim, holds no structural authority, and is refused at check time in a
workspace that declares no room. What the model wrote is what the room runs;
what the transcript shows is what the hiding rules left, and the run says so
when those differ.

**All of it is `tier: expert`.** No core capability requires a program, deleting
`programs/` leaves a working agent, and a workspace that carries one simply does
not earn the `no-code` badge — held by a test over the specification's own tiers
rather than by anybody remembering.

**And memory is a variable the format can name.** `bind: remembers.<name>` reads
what the conversation established into an argument the model never sees;
`remember-as:` writes a tool's answer back. `never-from: tool output` is what
stops one filling the other.

---

## Documents
Expand All @@ -242,6 +352,8 @@ intended reader cannot write code, so a message they cannot act on is a defect.
| [`docs/01-DECISIONS.md`](docs/01-DECISIONS.md) | 28 binding decisions — **read this before proposing anything** |
| [`docs/30-FRD.md`](docs/30-FRD.md) | 120 functional requirements, each traced to its justification |
| [`docs/40-IMPLEMENTATION-PLAN.md`](docs/40-IMPLEMENTATION-PLAN.md) | Milestones M0–M8, gates, risks, and what would falsify the approach |
| [`docs/50-NOT-COPIED.md`](docs/50-NOT-COPIED.md) | The refusal ledger — what was deliberately not copied from prior art, and what would bring each back |
| [`docs/70-PRODUCTION-GAP-REGISTER.md`](docs/70-PRODUCTION-GAP-REGISTER.md) | Every gap between what is claimed and what is built, with its measurement |
| [`research/notes/README.md`](research/notes/README.md) | Index of the research, with the 12 findings that changed the design |
| [`.claude/workflows/pact-architecture.js`](.claude/workflows/pact-architecture.js) | The research → critique → reflect workflow that produced it |

Expand All @@ -250,18 +362,25 @@ intended reader cannot write code, so a message they cannot act on is a defect.
## Layout

```
crates/ Rust core
adapters/python/ the harness + Pydantic AI and LangGraph transports
pact-diag/ diagnostics — a fix is mandatory by construction
pact-doc/ span-preserving YAML / JSON / Markdown
pact-loader/ the Expansion Rule (tree → document)
pact-schema/ validation; the schema is data, not code
pact-cli/ `pact check`, `pact show`, `pact waits`, `pact discover`, `pact card`
— never executes author code
spec/schema.yaml the specification, written in PACT
examples/ the worked no-code multi-agent example
research/ 14 studies + the 140-repo corpus (gitignored)
docs/ thesis, decisions, FRD, plan
crates/ Rust core
pact-diag/ diagnostics — a fix is mandatory by construction
pact-doc/ span-preserving YAML / JSON / Markdown
pact-loader/ the Expansion Rule (tree → document), and every check a
document needs that a schema cannot make
pact-schema/ validation; the schema is data, not code
pact-cli/ `pact check`, `pact show`, `pact waits`, `pact discover`,
`pact card` — never executes author code
adapters/python/ the reference harness, six framework transports, the
resolver, evals, learning and the port boundary
adapters/typescript/ the second, independent port — Node, and smaller on
purpose; it says which parts it is smaller by
spec/schema.yaml the specification, written in PACT
spec/loops/ the six loop shapes, authored the way anybody's are
examples/ the worked no-code multi-agent example, and eight
orchestration patterns
tests/trees/ small fixtures, each the smallest tree that shows one thing
research/ 14 studies + the 140-repo corpus (gitignored)
docs/ thesis, decisions, FRD, plan, refusal ledger, gap register
```

---
Expand Down Expand Up @@ -289,7 +408,16 @@ harness-vs-native.

**Learning emits source.** Every self-improvement is a signed, reviewable,
revertible diff to a spec file. An agent that grows is still an agent you can
read, fork, and port.
read, fork, and port. A written rule under a `## Rules` heading in a skill needs
a person however the edit is made — added, removed, reworded, or moved by
renaming the heading over it — and `pact check` shows the author where that
boundary falls rather than leaving them to trip over it.

**Convenience erases itself.** Patterns, inherited shapes and shared figures all
resolve and disappear before anything downstream reads the tree, so a workspace
that uses them digests identically to one written out longhand. Every claim this
project makes about portability, digests and conformance therefore covers trees
that use them, without one line of special-casing anywhere below the loader.

---

Expand Down
46 changes: 43 additions & 3 deletions adapters/out-of-tree/echo_adapter/transport.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,11 +48,51 @@ def lattice(self) -> dict[str, str]:
"parallel_tool_calls": "unsupported",
"streaming": "unsupported",
"durable_resume": "unsupported",
# An out-of-tree adapter is not obliged to track the core's key set
# — E-1 says adding one costs zero core changes, and no test here
# compares this dict against a core transport's. It is written all
# the same, because P-2's *"an adapter that omits a feature cannot
# be compared"* is the reason the key set matters, and an exemplar
# that quietly omitted the newest one would teach the omission.
#
# `emulated` rather than `unsupported`, and an eighth adapter's
# author does not have to do anything to earn it: a `connect:` tool
# reaches its server through PACT's own client and the `tool_impls`
# seam `harness.run` already has, above whatever this binds. Only a
# transport that takes the LOOP away — `a2a_transport.py` — is
# `unsupported` there.
"connected_tools": "emulated",
}

def usage(self) -> tuple[int, float]:
"""Tokens and money for the last call. Free, and counted honestly."""
return (len(self.prefix) // 4, 0.0)
#: Whether ANYTHING can put a price on what a call here carried. Nothing
#: can: this is a scripted echo bound to no model, so no row in
#: `models/catalog.yaml` describes it and none ever will.
#:
#: Separate from `usage()` on purpose, and the reason is the same one the
#: `lattice()` comment gives about the newest key — an exemplar that quietly
#: omitted this would teach the omission. `harness.run` reads two facts in
#: two steps: does `usage()` exist (tokens), and can anything price it
#: (money). It defaults the second to `False`, so an undeclared transport is
#: told nothing on its behalf and the author's money ceilings arrive on
#: `RunResult.unmetered` — which is the honest report here. This file said
#: nothing for a round and returned `0.0` below, and a run over it reported
#: `cost-per-request-under: 0.05 USD` as ENFORCED against a meter that read
#: 0.00 for the life of the workspace. That is B6, in the one file a third
#: party is told to copy.
prices_money = False

def usage(self) -> tuple[int, float | None]:
"""Tokens for the last call, and no price — because there is none.

`None`, never `0.0`. `transports/_metering.py` states the rule in the
imperative: *"An unpriced row yields `None`, never zero … Metering a
ceiling at 0.0 USD is a spend cap that can never be reached, under an
author who believes they capped their spend."* `harness._meter_usage`
adds the token count and leaves the money meter alone when the second
half is `None`, so a cap over this transport moves nothing and says so,
rather than sitting at 0.00 and looking held.
"""
return (len(self.prefix) // 4, None)

async def model_call(
self,
Expand Down
5 changes: 5 additions & 0 deletions adapters/python/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,11 @@ pact-pipeline = "pact_adapters.pipeline:main"
pact-explode = "pact_adapters.exploding:main"
pact-import = "pact_adapters.importing:main"
pact-export = "pact_adapters.exporting:main"
# The other direction for the one target that has a declarative format of its
# own. `pact-export` writes a registry record, which is an index entry; this
# writes an agent spec, which is a runnable agent minus what the format has no
# field for — and the report is where the difference is said.
pact-pydantic-ai = "pact_adapters.pydantic_ai_interop:main"
pact-improve = "pact_adapters.optimising:main"

[build-system]
Expand Down
Loading
Loading