Skip to content

Repository files navigation

Agentic Development Framework

A control plane for repositories worked on by AI agents. People and agents build contracts together - what this repository holds true, who decided it, and what evidence closes it - and every change is worked against them and adds to them.

日本語の解説は docs/concepts.ja.md にあります。

The problem

An agent asked to implement something will implement something. What it will not do reliably is notice that the specification it needs was never written down, and stop. It fills the gap - plausibly, invisibly, and in code that nobody reviews as a decision, because it does not look like one.

The gap is not the agent's to fill. Whether deleting a task should cascade to its attachments is a product decision. An agent can find that the question exists, lay out the options, and say which it would pick. It cannot be the one who decides, and no amount of prompting makes it the right party.

And when someone does decide, the answer usually lives in a conversation that ends. The next change asks again, or does not ask and answers differently.

Contracts are the point

A contract is what this repository currently holds true: the rule, who had the authority to set it, and what evidence closes it. For the deletion question above, that is three parts: attachments go with the task, the recorded decision that settled it, and the test that proves it. Contracts are the durable artifact here. Everything else - the actions, the roles, the checks - exists to build them, use them, and keep them honest.

People and agents write them together. The agent investigates the code, finds that a rule is missing, lays out the options and their impact, and says which it would pick. A person decides. That decision is recorded, and the rule it settled becomes a contract clause citing it as authority. Neither party does this alone: the agent cannot decide, and the person should not have to reconstruct the question from scratch.

Every change is worked against them. Before implementation, the contracts that govern the intended effects are resolved and challenged. During it, they are what the implementation must satisfy. After it, the change completes only when every direct or inherited clause affected by the change has traceable evidence and every review clause has been available to the independent challenge.

Contract authors choose that cost explicitly with evidence_mode. direct requires an Evidence claim that names the clause and states what the artifact demonstrates. inherited may reuse an Evidence artifact that covers the clause. review requires no Evidence Record and remains in Challenger context. A Contract without evidence_mode keeps the original all-direct behavior, so a framework update never weakens an existing Contract silently.

A Contract's change_id identifies its originating Change, not its applicability. Snapshots retain Contracts across Changes; action Context selects the applicable clauses by applies_to or an explicit clause reference in verification_scope. Revalidation therefore includes the selected clause's text, authority reference, evidence mode, and source digests even when another Change created the Contract. When upgrading from Change-filtered Snapshots, an existing Impact Assessment may need to be refreshed because its Contract index now includes those Contracts.

They grow. A question answered once does not come back - the next change that touches deletion finds the clause and resolves against it instead of asking again. An incident becomes a clause with a test behind it. The repository ends each change knowing more about itself than it did before, and that accumulation is the actual product.

They are layered, because a rule that governs one feature and a rule that governs the system are not the same kind of thing:

project        system-wide invariants, and where a human must decide
  domain       meaning, ownership, relationships, lifecycle
  capability   behaviour shared across features
  architecture standard implementations and dependency rules
  data         states that must hold across every operation sequence
  operation    one read or write: preconditions, effects, failure, retry
feature        what this change adds, and which of the above govern it

A feature contract is not a copy of the ones above it. It states the delta and names which ones apply. When a decision turns out to hold beyond the feature - say the deletion rule is really "no child record outlives its parent", a fact about the data rather than about tasks - it belongs in the layer that already governs that, and moving it there is how the repository stops relearning it.

The alternative is what usually happens: the knowledge lives in whoever was in the room, the agent re-derives it every session, and the two disagree.

What makes that trustworthy

Contracts are only worth building if an agent cannot talk its way past the process that enforces them. Work is issued one action at a time - an agent asks what to do, does that one thing, submits the result, and asks again - and each of those steps is decided by the control plane rather than by what the agent remembers.

The order of work is computed, not prompted. A kernel derives the next action from the records in the repository, and the action's identity is a digest of what it was derived from. An agent cannot skip a step by forgetting one, or invent one the state does not call for. It also means work survives a crash: a different process reaches the same action from the same records.

Authority is a checked property, not a judgement call. Reporting a requirement as met requires a clause in an accepted contract, an explicit requirement in the request, a recorded human decision, or an accepted decision record. An agent's own reasoning is evidence and never authority - not by convention, but because the submission is rejected without one of those four.

What the code does is read, not described. Detectors parse the actual source in sixteen languages. A call they cannot account for stops the change instead of being passed over, and nothing is classified by name: save and execute mean different things in different frameworks, so they need a binding a person reviewed.

Intent is assessed before source detection decides the scope. Every newly created Change begins with an Impact Assessment. Its result is explicitly one of impacts-identified, no-impact, or inconclusive; an empty result is never silently treated as no impact. This also gives an empty repository a valid bootstrap path: the first Change declares its intended effects, then creates only the Contracts needed to govern them before implementation starts.

Candidate reviews are resolved from their evidence, independently of Result ID or file order. A current not-applicable review takes precedence over stale reviews and proceeds to independent challenge. A confirmed review remains binding across evidence changes and takes precedence if reviews disagree. CLI, MCP, explanations, and challenge context use the same selection.

Contracts going stale is a feature. Each result is bound to digests of what it was based on, so changing a contract, the code, or the authority behind it marks the work that depended on it stale and asks for it again. A contract that drifted from the code does not silently keep passing. Committing only ADF Result or Evidence Records does not invalidate the product inputs those records verify. After adf_add_evidence, a refreshed adf_next Context may include that Evidence and other Evidence for the same Requirement. They are not retroactive prerequisites of the verification record. Evidence dependencies captured when the record was created still have to match. Submit truthful unsatisfied or inconclusive outcomes with the corresponding Evidence in basis_refs and output_refs; no commit or manual rewriting of Evidence is needed.

Nothing reviews its own work. Implementation and challenge are separate roles, and a post-build challenge runs in a context that did not build the change. The rules deciding all of this come from a signed framework release the project pins - a project cannot quietly widen what counts as verified.

And the boundary, which matters as much as the rest: the control plane checks structure, references, state, digests, and coverage. It does not judge meaning. A contract that says the wrong thing is accepted. Knowing exactly what it does not verify is what makes the rest worth trusting.

Installing

No release has been published yet, so today the way to get adf is to build it. The download below works from the first release onward.

Build it. You need Rust 1.89 or newer, and nothing else:

git clone https://github.com/piaro/agentic-development-framework
cd agentic-development-framework
cargo build --release

A project pins a signed framework release - the rules and schemas it is evaluated against - and a downloaded binary arrives with one beside it. A binary you built does not, so build one to develop against. The key here is throwaway; a real one belongs in the publishing job:

SEED=$(openssl rand -hex 32)
PUBLIC_KEY=$(ADF_RELEASE_SIGNING_KEY_HEX=$SEED ./target/release/adf release public-key)
ADF_RELEASE_SIGNING_KEY_HEX=$SEED \
ADF_RELEASE_SIGNING_PUBLIC_KEY_HEX=$PUBLIC_KEY \
  sh scripts/release-ci.sh
# the release is in dist/framework

Then point initialization at it with --candidate-dir dist/framework.

Or download it, once a release exists. The bootstrap script fetches the binary for your platform along with the framework release it pins, checks that GitHub attested both to a build of this repository, and installs them together. It needs the GitHub CLI for that check:

sh bootstrap/install.sh --tag framework-<release-id>
# add the printed bin directory to PATH

It refuses to install anything whose attestation does not verify, so a tampered or unattested download stops there rather than landing on your machine. Offline installs, key rotation, and rolling back are in docs/implementation.md.

Running users need no Python and no Rust toolchain - the binary carries everything, including the agent skills it places into a project.

Getting started

You need adf and a git repository.

adf project init --project /path/to/project

That places the configuration, the pinned framework release, the three agent skills, a guide at docs/adf/README.md, and a block in AGENTS.md. Nothing existing is overwritten, and nothing is committed for you.

If the repository already has code, let the detectors list what they can see and review it before it counts:

adf project observe --project . --output .adf/repository-observation.draft.yaml
# fill in the logical IDs, owners, and the accepted decisions that authorize them
adf project validate-bindings --draft .adf/repository-observation.draft.yaml
adf project promote-bindings --draft .adf/repository-observation.draft.yaml

Candidates are never applied on their own. A name that looks like a database write is a candidate, not a fact.

Then start a change:

adf change init change.first-feature \
  --title "First feature" \
  --intent "Why this change exists"

adf next change.first-feature

To verify behavior that already exists, select only the Contract clauses whose Evidence you will register:

adf change init change.verify-existing-orders \
  --title "Verify existing order behavior" \
  --intent "Register reproducible evidence for existing behavior" \
  --verify-clause contract.order-lifecycle#orders-source-of-truth

The selected clauses enter the normal Evidence workflow even when they have not been verified before. Other unverified clauses do not block this Change. Existing behavior is not reported as a new implementation impact merely to select it for verification. Clauses whose evidence_mode is review remain human-review work and cannot be selected with this option.

From here, next says what to do. Agents normally reach it over MCP:

adf mcp --project /path/to/project

The loop

        adf change init
                │
                ▼
    ┌──▶ adf next ──── one action, with only its required Context
    │           │
    │           ├─ Analyst    assess intended impact, review detected signals,
    │           │             write contracts, ask a person when no authority
    │           │             settles it, record what they answered
    │           ├─ Builder    implement, then record required clause evidence
    │           └─ Challenger try to falsify it, before and after the build
    │           │
    └───────────┴─ adf_submit ──── validated and stored
                │
                ▼
          ready to merge

Three skills, one per role, are placed into the project by project init. The order of work is not in them - it comes from next. A challenge after the build runs in a context independent of the one that built it.

State What is assigned
needs-impact-assessment assess intended effects before implementation
needs-post-build-impact-assessment reassess because code or governance changed
needs-analysis review detected candidates, answer the requirements
needs-human-decision put the question to a person
needs-decision-recording record their answer as a decision and a contract
needs-pre-build-challenge falsify the request, the authority, the contracts
ready-to-build implement
needs-evidence record evidence for affected direct and inherited clauses
needs-post-build-challenge falsify the implementation
ready-to-merge nothing

adf explain <change-id> says why the change is where it is.

Context reuse and model guidance

next compiles a Context for one action instead of handing every action the entire repository history. Impact assessment receives compact repository, Contract, and Decision indexes plus at most three prior assessments. After an assessment is accepted, implementation receives that Result, matching governance, and matching artifacts. This makes the assessment a reusable input instead of asking later actions to rediscover the same scope.

While impact assessment is still pending, next and explain select that action before deriving repository-wide Contract health. Unrelated Result and Evidence history is not loaded for this first step.

When Contract health is required, ADF maintains persistent Evidence and Result indexes under .adf/cache/runtime/. Unchanged tracked records are identified by their Git blob IDs; changed and untracked records use content hashes. ADF parses only records whose source identity changed, and it does not scan every Result again for each Contract clause. Corrupt cache entries are rebuilt from source. ADF writes runtime caches only when Git confirms that the cache path is ignored.

Repository observation uses a separate cache tied to the current revision, analysis configuration, signal catalog, and source identities. Source changes invalidate the observation without making cache files authoritative.

New Results store shared input and freshness references once at the Result level. An outcome carries its own references only when they differ from those shared values. Existing Results remain readable and are not rewritten, so their identities and downstream freshness checks remain stable.

adf_submit returns after the Result is stored. Its response includes result_id, already_completed, next_required, and per-stage timings_ms. Call adf_next separately when next_required is true. This keeps a slow next evaluation from obscuring whether submission itself succeeded. adf_next also reports per-stage timings.

Each action also carries advisory execution guidance. Impact assessment normally recommends an economy model, while challenge recommends a high-accuracy model. The listed escalation conditions tell an orchestrator when to choose a more capable model. ADF does not invoke or select the model itself.

Lightweight execution log

An orchestrator may attach measurements it already has to adf_submit: duration, model, input and output tokens, tool calls, retries, and timestamps. ADF records the serialized Context size while it is already validating the submission. It does not start a timer, call a model, or run another tracing pass to collect metrics, and it never estimates missing values.

An external runner can also bracket an attempt with adf_begin_execution and adf_complete_execution. These append-only events are separate from Results, so a failed, interrupted, or still-incomplete attempt remains visible. A completion can be attached after adf_submit, when a non-interactive agent has reported its final token counts. External completion records may additionally carry cache-creation tokens, cached-input tokens, reasoning-output tokens, and provider-reported USD cost. ADF never estimates a missing cost. Runner events do not affect Kernel state, Result identity, freshness, or Evidence validation. If a runner completion and the legacy adf_submit.execution describe the same Result, the execution log uses the runner completion and does not count the Result metrics twice.

Read the per-action entries and totals with adf_execution_log over MCP or:

adf execution-log <change-id> --format json

Experimental agent runners

adf-codex-runner and adf-claude-runner are optional adapters, not part of the ADF control plane. They run only when a person or a primary agent invokes one of them. Both accept Challenger Actions only and start one independent non-interactive session per invocation.

The primary session first obtains the expected Action ID and Context digest, then invokes:

adf-codex-runner run \
  --project /path/to/project \
  --change change.example \
  --expected-action action.example \
  --expected-context sha256:...

Use the same identifiers with Claude Code:

adf-claude-runner run \
  --project /path/to/project \
  --change change.example \
  --expected-action action.example \
  --expected-context sha256:...

Each runner re-evaluates adf next before launch. The child agent must also call adf_next through its own MCP session and stop if the identifiers differ. It receives the complete Generated Context from ADF; the runner does not summarize the primary chat or turn that summary into authority. Durable requirements must already be in the Change, accepted Contracts, or accepted Decisions.

The Codex adapter uses codex exec --json --ephemeral, an explicit workspace-write sandbox, and a JSON Schema for the final response. It records the turn.completed input, cached-input, output, and reasoning-output token counts after the Result has been submitted. Raw JSONL and the primary chat are not stored. See the official Codex non-interactive mode documentation for the underlying CLI event contract. Build the experimental binary from source.

The Claude Code adapter uses claude -p --output-format json --json-schema with --no-session-persistence. It deliberately does not use --bare, because the independent execution needs the project's MCP server, Skills, and instructions. It records input, cache-creation, cache-read, and output tokens, the actual model names, and total_cost_usd reported by Claude Code. The full Claude response and the primary chat are not stored. See the official Claude Code programmatic execution documentation for the underlying CLI contract.

Build either adapter from source with:

cargo build --locked --bin adf-codex-runner
cargo build --locked --bin adf-claude-runner

The signed binary release currently continues to publish only adf; runner distribution is a later compatibility milestone.

Reducing stored Record size

ADF can store identical input and freshness reference maps once within each Result. Each file remains self-contained JSON. Reading it restores the exact logical Record, including explicit versus omitted fields, before validation and digest calculation. Record IDs, evidence, explanations and the pinned Framework identity are preserved. Small Records remain plain JSON when sharing would cost more space.

Existing projects keep their current write format until explicitly migrated. Before migration, stop agents, MCP sessions and CI writers for that working directory and update every reader/writer to a version supporting adf-record-refmaps-v1. Older CLIs reject the new project configuration on most paths, but some old execution commands bypass that check; configuration alone cannot stop a running old writer.

adf project storage inspect --format json
adf project storage migrate --to adaptive-refmaps-v1 --dry-run
adf project storage migrate --to adaptive-refmaps-v1
adf project storage verify
adf project storage export --record result.<id> --format json

These commands accept --project <root> and the existing offline --release <root> option. Inspect, dry-run, verify and export leave Records, configuration and derived indexes unchanged; they acquire a small maintenance lock under .adf/cache/locks/. Reports describe Result/Evidence JSON bytes only, excluding execution events, non-JSON evidence, caches, Git history and other working directories. A skipped-file list identifies non-JSON files and interrupted temporary files.

Migration validates Records against the pinned signed release and checks lossless roundtrips before writing. It enables project config version 2, atomically replaces individual files, and keeps a small Git-ignored recovery journal under .adf/local/storage-migrations/. A local .gitignore is created there if needed; existing ignore rules are never overwritten. Configuration YAML may be reformatted. Normal operations cooperate with the migration lock. Do not edit files, switch Git revisions, or remove lock files while migration is running.

Repeat the same migration command after interruption. If a Record changed since an interrupted migration, the command stops instead of overwriting it. Inspect that conflict before continuing. After verifying the changed Records, move the recovery journal aside and run dry-run again to review a new plan; do not rewrite its hashes to conceal the conflict. To restore compatibility with older CLIs, expand all Records before downgrading the configuration:

adf project storage migrate --to plain-json-v1 --dry-run
adf project storage migrate --to plain-json-v1

Rollback restores equivalent JSON, including untracked Records, without keeping second copies of their contents. Original whitespace is not retained. Free space is checked before either direction. Completed replacement files are synced before publication; an interruption during a temporary write can leave a temporary file, which is reported and must be reviewed before removal. A restored Record may be up to 256 MiB; a sharing envelope may be up to 64 MiB. Excessive expansion is rejected before copying shared maps. JSON nesting is subject to the parser's depth limit.

No commits, Git history rewrites, cache deletion or automatic legacy Kit migration are performed. Existing indexes rebuild from restored Records when their physical source changes; their own reference duplication is a separate capacity cost. A Git revision or source change still triggers the normal freshness rules.

Commands

Command What it does
project init set a repository up
project observe list what the detectors can see, for review
project validate-bindings report missing bindings and what cannot be checked
project promote-bindings make a reviewed draft the observation of record
change init start a change, optionally selecting existing Contract clauses to verify
next issue the next action
explain say why the change is where it is
execution-log aggregate Context size and any execution metrics already reported
execution begin/complete append an external runner attempt without launching an agent
contract-health check the contracts across the repository
mcp serve the same operations to agents over MCP
release build, fetch, install, switch, and roll back framework releases
binary install, update, and roll back the CLI itself

migration, benchmark, detector-audit, and catalog also exist and are experimental - see COMPATIBILITY.md.

What it supports, and what it does not

Detectors read the source and report what they cannot account for rather than passing over it. See docs/limits.md for the supported languages, the calls that are not resolved, and what a blocked change means.

The short version: sixteen languages have detectors, C++ deliberately does not, aliases and dynamic dispatch are not resolved, and a gap stops the change instead of being ignored.

Documentation

Document What is in it
docs/limits.md supported languages, known gaps, what a stop means
COMPATIBILITY.md what is promised and what may change
docs/publishing.md releasing, for whoever holds the signing key
CONTRIBUTING.md building, testing, and what is out of scope
SECURITY.md reporting a vulnerability, and what counts as one
docs/concepts.ja.md what it solves and how to use it, in Japanese
docs/implementation.md the implementation of record, in Japanese

License

MIT or Apache-2.0, at your option. See LICENSE-MIT and LICENSE-APACHE. Contributions are accepted under both.

Published binaries link their dependencies statically and ship the terms those dependencies require, as THIRD-PARTY-NOTICES.md.

About

Control plane that connects specifications, decisions, implementation, and evidence in a repository worked on by AI agents

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages