Skip to content

Repository files navigation

driftlock

driftlock — Checkpoint, detect drift, roll back, learn

CI Latest release Python 3.13 No runtime dependencies 170 archived trials MIT license

A long-horizon coding agent that checkpoints its work, rolls back when it drifts,
and distils what it learns into skills that must earn their place.

Quick start · Results · Architecture · Usage · Contributing


Checkpoint Detect Roll back Learn
Snapshot filesystem and agent state together Zero-token heuristics escalate to a trajectory-aware judge Return to the last healthy state with bounded retries Distil localized failures into skills, then admit only validated gains

Important

The published 170-trial result measures Terminus-2 wrapped in driftlock's checkpoint, rollback, distillation, and validation machinery. driftlock's native tool-calling agent ships, but does not yet have a published measurement. The distinction is retained in every run record and explained in RESULTS.md.

The long-horizon problem

Agents fail differently on long tasks than on short ones. Frontier models solve near-100% of tasks a human expert finishes in under four minutes and under 10% of tasks that take a human more than four hours; a recurring rule of thumb is that doubling a task's length roughly quadruples its failure rate.

Two failure modes dominate that regime. Context rot — relevant information gets harder to retrieve as history grows. Compounding error and goal drift — small early mistakes snowball until the agent is working on the wrong thing. driftlock attacks the second, and turns what it learns there into skills it carries into the next task.

Quick start

Python 3.13 and uv. The library is stdlib-only at runtime; pytest and ruff are dev extras, and sentence-transformers is optional — needed only by the pinned-embedder integration test.

git clone https://github.com/hyj28/driftlock
cd driftlock
uv sync --extra dev
uv run pytest

Then watch the thing the first line of this page claims — no API key, no network, 0.1 seconds:

uv run python examples/rollback_demo.py

A scripted stand-in drifts into a loop, the coarse detector names the signal, and the run returns to the last checkpoint still eligible under the detector's lookback — skipping one that captured the drifted workspace. The retried step prints what it found there, so the output is evidence rather than narration.

From there: the usage guide for the agent and every optional component, RESULTS.md for the measured evidence, or the runner API below.


Undo is the missing lever

The standard vocabulary for context engineering has four levers, and every one of them assumes the agent keeps moving forward:

Lever Mechanism in driftlock
write Rollback-grounded skill distillation into a persistent library
select Agent-initiated retrieval over skills, workspace and memory in one corpus
compress Context compaction at checkpoint boundaries
isolate Fresh bounded sub-agents with their own conversation and no recursion
undo Checkpoint + progress-aware rollback

Snapshot the filesystem and agent state together, then let a two-tier judge periodically ask whether the current state is still a sound basis for continuing. The coarse tier is zero-token heuristics — no file changes for N steps, action loops, error spikes, reward stalls. The fine tier is a cheap model reading goal, plan, recent trajectory and diff. If the answer is no, roll back to the last healthy checkpoint and retry from there.

Skills have to earn their place

Recent work found that self-evolving agents improve through validation-filtered search, not accumulation: only 55 of 388 candidate skills produced a real gain, and every selected improvement was grounded in a failed trajectory.

driftlock narrows the grounding further. Because every checkpoint can be scored by the task's own verifier for free, a trajectory becomes a timeline, and a flat segment is a stretch of steps that provably bought nothing. That localized evidence — a bounded region with a diff attached — becomes the supervision signal for distillation, in place of a whole failed trajectory.

A candidate enters the library only after a paired validation run: ten replicates of its own source task, with and without the skill, differenced per replicate, against a control shared by every candidate from that task. When retrieval selects nothing, the treatment prompt is byte-identical to its control — which yields a measured noise floor at no extra cost.

What was measured, and what was not

A 170-trial validation run cost $15.06 and is reported in full, including the parts that did not work, in RESULTS.md. One candidate of fourteen was admitted against a chance expectation of 0.150; the two distillation arms were indistinguishable at this sample size; and six candidates were never retrieved at all, which is the finding that drove the agentic-RAG work.

That run measured Terminus-2 wrapped in driftlock's checkpoint, rollback, distillation and validation machinery. All 287 archived job configs used driftlock.harbor_agent:LHTBDriftlockAgent. driftlock's own tool-calling agent, and the components built on it, have no published measurement yet. Both paths ship, and which one a run used is recoverable from its record.


What is built

The loop itself. Checkpointing, progress-aware rollback and the two-tier judge; free checkpoint scoring through the task's own verifier; failure localization to a checkpoint segment; skill distillation, retrieval and once-per-task injection; paired validation and admission with a measured noise floor; resumable runs, bounded retries and degraded-observation reporting; and a tool-calling agent with five tools — run_shell, read_file, write_file, search_files, complete. Context compaction at checkpoint boundaries is always on, bounded by driftlock_max_history_characters rather than switched.

Optional components, each behind its own flag:

driftlock_agentic_retrieval Agentic RAG — retrieval as a tool, over code and skills
driftlock_planning Planning and task decomposition
driftlock_memory Memory that persists across runs
driftlock_delegation Bounded sequential sub-agent delegation
driftlock_parallel_reads Bounded parallel workspace reads
driftlock_self_verification Output self-verification
driftlock_edit_file Bounded exact-string file edit
driftlock_prompt_cache Prompt-cache management

The MCP client — stdio and Streamable HTTP, with host-supplied authorization — is available on the agent but deliberately has no harness flag: LHTB tasks provide no server, so an enabled arm would be indistinguishable from an empty one. The reason is asserted in the code, not just written here.

Everything past the tool-calling loop is opt-in. An agent built without them offers exactly the five historical tools and sends a byte-identical request, which is what keeps the archived experiment replayable. Each carries a driftlock_* flag into the experiment harness and appears in the run record's active-component set, so a trial can always be attributed to the configuration that produced it.

Runner API

from driftlock.runner import DriftlockRunner, RunnerConfig
from driftlock.checkpoints import DirectoryCheckpointStore
from driftlock.heuristics import HeuristicJudge, HeuristicConfig

result = await DriftlockRunner(
    DirectoryCheckpointStore(workspace, store_dir),
    HeuristicJudge(HeuristicConfig()),
    config=RunnerConfig(max_steps=50, max_rollbacks=3, checkpoint_interval=5),
).run(
    goal="repair the parser",
    plan="inspect, patch, verify",
    step=agent,
    initial_state=agent.initial_state(),
)

Full usage — the native agent, every optional component, remote and Harbor environments, and the one-command LHTB harness — is in docs/usage.md.

Documentation

RESULTS.md The 170-trial run, its numbers, and its limits
docs/architecture.md How the pieces fit, and the invariant each one holds
docs/usage.md Running the library, the agent, and the experiment harness
docs/design-journal.md The dated working plan, kept as a record of how the design moved
CONTRIBUTING.md The engineering discipline this repository holds itself to
CHANGELOG.md What landed, in order
CODE_OF_CONDUCT.md How review here treats work, and people

Repo layout

src/driftlock/    # the library: runner, agent, checkpoints, judges, skills, components
tests/            # no network: real subprocesses and real loopback servers
examples/         # runnable with no provider key
integrations/lhtb/  # the Terminus-2 adapter the published results came from
scripts/          # round orchestration for a remote experiment server
docs/             # architecture, usage, design journal
RESULTS.md        # the measured run

What this repository optimizes for

Every component states what it does not guarantee as plainly as what it does, and that statement lives in the code rather than only in the docs. Self-verification says it is not adversarially sound, and why. The MCP client says a credential reaches it only from an injected supplier. The local environment says it is not a process sandbox. The file edit says hard links break.

The same standard applies to measurement. A status that cannot distinguish we could not observe this from we observed nothing is treated as a defect, because a blind channel reading as a negative result has destroyed real measurements here — five separate times, catalogued in RESULTS.md §6.

License

MIT — see LICENSE.

About

Long-horizon coding agent with checkpointing, drift detection, rollback, and validation-gated skill learning. Python 3.13, stdlib-only runtime.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages