Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,8 +81,8 @@ use the documentation by job:
- [tutorials](docs/tutorials/README.md) — review, approval, environment
repair, bounded debugging, and migration;
- [examples](docs/examples.md) — working programs in all four languages;
- [evaluations](evals/README.md) — pinned conversion cases, reproducible
measurements, and the raw-evidence publication boundary;
- [evaluations](evals/README.md) — first-party workflow conformance and runtime
invariant results, including the exact claim boundary;
- [convert an existing skill](docs/convert-existing-skill.md) — move
control flow into code without claiming that fixture execution proves
every reading of the original prose;
Expand Down
41 changes: 11 additions & 30 deletions UPSTREAM.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,11 +2,11 @@
"files": {
".gitignore": "803c5f79d6da7f2c5a1dc0ce27c53b8e5c059309b165782831c4c472af058a4c",
"LICENSE": "fff261ce507eabd57666c283a621f33e183a3aedebda04c4ecbc6309a62f5edf",
"README.md": "d3a13ebcd016ab1ba6a4fac64ca4a555ff738e9ce4c5c9744ee4e046e9dc08ef",
"README.md": "acd9f37f0f7013a4bf99df382fd3b5c14d070914ec958e3198d9c70aa8317ce6",
"cmd/yskill/main.go": "a88bdd133118aa7b86cdf021e67052bd8e64e63b3792c091f7612699af89648b",
"docs/README.md": "cd8a4105f4f04143172661b39789e3c533d031ad46437aa2c078625a29a0a9db",
"docs/convert-existing-skill.md": "0dd538e908a1f2d3a1958e4b2cc1efede2a8e7a73341e32c1a75fec69a9acb8a",
"docs/examples.md": "8c81f990b8d3a06b42cedcf34813a11dbe0606201782f72903a8c14d91d1ae23",
"docs/examples.md": "3c19bacae7ec7b31bf93417e61d01cd1872f26228fedf0ed6decfce7c41ed274",
"docs/locus-conformance.md": "a71fd5a72678dcad344afc71cd7a788090ef008a8e5e82d925da5a3db41af6a0",
"docs/locus-converter.md": "b8d1bdb63cb0283ace524f572fb13f0d1cb535b130c5ea3e112d2efae55a8681",
"docs/locus-yield.md": "c29d0c5801f32591828fc348385fb25c1850b39afc22770098eae0614fc3a417",
Expand Down Expand Up @@ -40,33 +40,13 @@
"docs/tutorials/data-migration.md": "6699fc47cda84c4f46a3703ca4a69bcced7cc4a9df339a308e807ea3a99578bd",
"docs/tutorials/environment-repair.md": "289c9b261e3768c2c58d60d9ef7a09e85f89038438a08febb4b9c12ebada88d1",
"evals/.gitignore": "0d5020173666118bafe31c857b96aa325809f41d159ca51324ceaf239e043347",
"evals/README.md": "c6fb08381877d3cddce22afe59746902ba73deebd78e52a414f85bcbbf441b22",
"evals/cases/README.md": "03bd5a5bca02c9052a8413b5d30f463d37e30d130a99093d9642a1e0bbede29d",
"evals/cases/actions-auditor/README.md": "e9efdfa7657840ed1829b2831abc1ef988c7f3faac5de98e9444d5751df9b913",
"evals/cases/actions-auditor/SKILL.md": "bdc81816f25144bcc44badb14ff3b46a234365466a8eeb6b3a6f6128ed57fd9b",
"evals/cases/actions-auditor/workflow.ts": "622f6dd037548a5ae0eb097a77392dbcc92eaa294691a4c4e30c4abfaa9544ec",
"evals/cases/doc-coauthoring/README.md": "d72fc22a0d7134c0764efb190fad7f7a6a7d93b0f55ce09447f469cb1711ecda",
"evals/cases/doc-coauthoring/SKILL.md": "8ab3e342ef202f22fe2763d1d5e0d5d6918ccdf5a30426abfc02f71b8a0b00d8",
"evals/cases/doc-coauthoring/workflow.ts": "df78a605f5c0d7e52ba63ef2b91757dda1f4a5229fbe66fcb7c9adcb9aac9efb",
"evals/cases/gstack-review/README.md": "fbfcba7aca73c0fc13813e3e85edf86b4988ac3e0e538209d8d3a89df7e35cd2",
"evals/cases/gstack-review/SKILL.md": "9b4914d4e569cf70fdc794d88c1c1714d4e88b4646ab9762d29da2cc2b892fba",
"evals/cases/gstack-review/workflow.ts": "0743492bc14aeacd941cfbe90ef909ffe47f474941fd1cc28c33a38147ae6910",
"evals/cases/index.json": "b363426b292939f977360daaaad3ea4875a447f59657fed74275f6cc4084de43",
"evals/cases/mcp-builder/README.md": "db708ccf258d6888173e573c02455a40c689ab51e09f7c57311fe1cefea32f14",
"evals/cases/mcp-builder/SKILL.md": "0a5c45ae7ad2e902f2a5c2c782e78db995b8e42869f36fad42839993e346dd5e",
"evals/cases/mcp-builder/workflow.ts": "fdf1ed76dc0cbe0eafd95ce1c8da7f8e65419e1f655c2df9ce88eaf285eee8a5",
"evals/cases/systematic-debugging/README.md": "5654154dd172c6405fcd8833d73c2ec2deb273416208b29890eb35b84fa0b438",
"evals/cases/systematic-debugging/SKILL.md": "2d130b11d680d43f02ef126a77b462c585101e4df1632b89873597acdd7b189b",
"evals/cases/systematic-debugging/workflow.ts": "86aaacc7d1fab4e320e5ce0fab7bd9972c137d6b17e8ebddc6d118299d2a1a80",
"evals/cases/vercel-deploy/README.md": "e585f09b81b3a3a2f23624f1d8e7c816af5b50d88f4204513072128bdef39b0c",
"evals/cases/vercel-deploy/SKILL.md": "4ac80ad71f42c327f63685530b6b98ce3e287459527d2c0bafb5e1c716c313da",
"evals/cases/vercel-deploy/workflow.ts": "78667cea9b4b4d3dfc54ef6944ca9aa75d837b0170a4d3811b8f55172f63d6f2",
"evals/package-lock.json": "22c099aa5f2d9959084d6703e4a94ff7fa34216cac21b9d81137e2dd155c09d5",
"evals/package.json": "3361b8d579664bf9d6642eaeca6a906a1bca43fd0b1b543b00b2d5e0aa021b17",
"evals/results/README.md": "8eda7b901660466e60e5e32609cb2a3dcca7d9d512a93f2f379090b16da15525",
"evals/results/latest.json": "339a15841921fa270c668aa5c55cd33a271493735b4b85beaeaba1b25807216a",
"evals/scripts/measure-source.mjs": "064a62073422c067e12b0f2b77cc65a38f7d3440efb1a91020595a941470db7a",
"evals/scripts/validate.mjs": "2616fdf2ebbfd940dff7e06bcb33f4a362e063484a4beb62f54ef2856747abd7",
"evals/README.md": "7b0a92908d4eff84f1716756a00536ed79b31ef542928a2fd0c17c234161bba6",
"evals/package-lock.json": "cfbd68d590e92b94233a807fd774e6667e9474096ab76e1c5cd037acfb4e3200",
"evals/package.json": "aa97f75cadbf6d1802dcf03f4b34dcbbc4aea8c2b5ef18130af26a9e63ecac51",
"evals/results/README.md": "65b74bfab83dfc51fc5b8a3a4824c37b722fb3a358a5dcfb217dc2af199f4801",
"evals/results/latest.json": "cc9346c0db8e8395464a7348ea440a891fad0080350ac5d677b1a2875047e4d4",
"evals/scripts/run.mjs": "119248cd4d383326d396f4459df28c51ce7c31088d97639851f2f5ab4c329892",
"evals/scripts/validate.mjs": "7ac2a3239cd1b52a604c8a8a52a1af4ac55c4007ac2751d4fbfe20d0e1f568dc",
"examples/convert-skill/SKILL.md": "e6376f34365d4ac030d316db55e91f0c606a501668099f4ecf1e27d43ec806a2",
"examples/convert-skill/fixtures/responses.json": "5b5f27b0ae5962360ac7b0e779992f430c655f352a38da37debb3505c32cd60a",
"examples/convert-skill/main.go": "eb987ed1ce9139e6d0d40a2fae20c532c2fba3b27357e31555db2f0f52db4e39",
Expand Down Expand Up @@ -286,6 +266,7 @@
"release-notes/2026-08-01-evaluation-case-guides.md": "570bc996eb4d2892456a02d938fb6299d106c431bb842ce726db1df550f779d7",
"release-notes/2026-08-01-evaluation-surface.md": "64fb5e2fdad3ccd41967e028cf4675c45939174000281c443f4b28d96dc07549",
"release-notes/2026-08-01-example-library.md": "14d6ca40529a6aeeb72872295e57bb0d7dcc7824df31e029e497602060ce4c97",
"release-notes/2026-08-01-first-party-evaluations.md": "9e74c48112343c409a88358138db80d731004f87900f1b23073f16c415163fcf",
"release-notes/2026-08-01-initial-projection.md": "d38f0832b5552fb97237b30d19bb63369442ddde69e678b752ac07eedeb7ba3d",
"release-notes/2026-08-01-multi-language-and-converter.md": "d0cf62d191e6a58e827f3b35d442f88b83be4dde98aa25ad54e4ae3ff787a453",
"release-notes/2026-08-01-remove-stray-analysis-traces.md": "0567f78ee97ffd23b3f26b5c39606e9ff6659c50a3fdef04ed9b3aa86cfa99af",
Expand All @@ -302,7 +283,7 @@
"generator": "operatorstack/yield:project",
"schema_version": 1,
"source": {
"commit": "381aea9c6d153503e68ad96e82ac236a354db432",
"commit": "4c338abc0d14f017eb3501885adca7594bd5b6f4",
"path": "labs/22-yield",
"repository": "operatorstack/intelligence-flow"
}
Expand Down
16 changes: 7 additions & 9 deletions docs/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,14 +61,12 @@ When adapting an example, change the repository-specific commands and model
instructions. Keep stable operation IDs for existing steps so saved runs can
replay them.

## Conversion cases used by evaluations
## Evaluated examples

The [evaluation cases](../evals/cases/) are different from the example library.
Each one pins a public third-party skill to an exact commit and digest, then
shows the smaller model-facing `SKILL.md` beside the TypeScript workflow that
owns its order, commands, gates, and completion.
The [first-party evaluation suite](../evals/) runs every example-library
pattern through TypeScript, Python, Go, and Rust. It also checks resume, replay,
changed behavior, blocking, and changed source.

The source-size harness and early summaries live in [`evals/`](../evals/).
Large transcripts and temporary repositories are external artifacts, never
committed source. A behavioral result is publishable only when its summary
binds the exact artifact URI and SHA-256.
These results show that the tested Yield version runs the workflow steps in
code. They do not compare Yield with prose or test whether an agent's judgment
is correct.
70 changes: 43 additions & 27 deletions evals/README.md
Original file line number Diff line number Diff line change
@@ -1,39 +1,55 @@
# Yield evaluations

This directory contains the public, reviewable part of Yield's evaluation
system: case definitions, pinned source identities, conversion programs,
measurement code, validation rules, and small result summaries.
These evaluations test Yield itself. They do not compare Yield with another
tool, company, prompt, or skill.

Raw agent transcripts, temporary repositories, command logs, and model
responses do not belong in Git. A full campaign uploads those files as one
immutable artifact bundle and records its URI and SHA-256 in the published
summary. Until that bundle exists, the summary must say `unpublished`.
The suite answers two questions:

## Layout
1. Can each checked-in example workflow reach its expected final result
through every supported SDK?
2. Does the runtime behave correctly when a run resumes, replays, blocks, or
encounters changed code?

- `cases/` — pinned public source identities plus the thin skill and Yield
program used for each conversion.
- `scripts/measure-source.mjs` — reproduces the source-size comparison from
pinned upstream files.
- `scripts/validate.mjs` — fail-closed validation for cases and summaries.
- `results/latest.json` — small website-safe summary. It is not raw evidence.
- `runs/` — local or CI output; ignored by Git and projected releases.
## Current coverage

## Run
- 10 workflow patterns written by this project.
- 4 SDKs: TypeScript, Python, Go, and Rust.
- 40 end-to-end workflow tests.
- 5 runtime checks: resume and complete, repeat the same saved step, stop when
behavior changes, block when a rule fails, and require approval for changed
source.

Run the exact suite and refresh the checked-in result:

```bash
cd evals
npm run eval
```

Check that the published result still matches the current source:

```bash
npm install
npm run validate
npm run measure
npm test
```

`npm run measure` writes a fresh summary under `runs/`. Publishing that result
requires a separate promotion step that binds the raw artifact digest, the
exact Yield commit, model identity, harness version, and case-set digest.
## What a passing result proves

A passing result proves that the tested Yield revision:

- executes each owned workflow test to `completed`;
- runs command steps rather than asking the model to invent their outputs;
- presents requests in the program-defined order;
- resumes from recorded responses;
- returns to the same saved step during replay;
- stops on changed behavior or failed requirements.

## What it does not prove

## Claim boundary
This suite does not prove that Yield is better than prose, that an agent's
judgment is correct, or that illustrative commands are production-safe. The
Fixed test data supplies agent and human responses so the suite can test only
the code-controlled workflow layer.

Source-size measurements show how much model-facing text and workflow source
the prototypes contain. They do not prove behavioral equivalence. Behavioral
claims require executable fixtures, held-out oracles, repeated model runs, and
the immutable raw artifact named by the result summary.
`results/latest.json` is a compact, website-safe result. Its source hash is
computed from the CLI, engine, protocol, SDKs, example workflows, fixtures, and
evaluation harness. CI reruns the suite instead of trusting that file alone.
15 changes: 0 additions & 15 deletions evals/cases/README.md

This file was deleted.

34 changes: 0 additions & 34 deletions evals/cases/actions-auditor/README.md

This file was deleted.

17 changes: 0 additions & 17 deletions evals/cases/actions-auditor/SKILL.md

This file was deleted.

13 changes: 0 additions & 13 deletions evals/cases/actions-auditor/workflow.ts

This file was deleted.

34 changes: 0 additions & 34 deletions evals/cases/doc-coauthoring/README.md

This file was deleted.

17 changes: 0 additions & 17 deletions evals/cases/doc-coauthoring/SKILL.md

This file was deleted.

15 changes: 0 additions & 15 deletions evals/cases/doc-coauthoring/workflow.ts

This file was deleted.

Loading
Loading