Release v0.7.0: build, evaluate, and improve decision-model systems - #91
Conversation
Claude Code adversarial review of 02f379aRequested Fable 5.1 / xhigh / permission bypass; observed Claude Code model label Reviewed commit and checksReviewed Checks actually run:
I independently re-derived the Hoeffding radius, the regret bound, the union bound, the XOR side-information example, the zero-error exact bound, and the Jensen claim. All match the runtime text. FindingsNo blocker. The four prior Astra defects are fixed and covered by tests that I could not break. Notes follow, most material first.
Composition-law review: the laws are stated with their preconditions and I found no counterexample within those preconditions. One sentence would help: the "free observation cannot raise Bayes risk" statement holds for a single decision-maker with a correct joint model; estimated policies in finite samples, misspecified models, and strategic settings can lose value from extra information. The text already limits scope with "common joint model," so this is a clarity note. Statistical design: fixed-sample one-sided Hoeffding with Bonferroni over a declared family is correct as written and conservative. Leakage, adaptive reuse, family declaration, causal identification, and label validity are explicitly delegated to the caller in the docstring, the reference, and the output Process items: the shell collectors contain no git, gh, push or publish verbs; Limits, scope, and identityNot verified by me: Scope and claims: a minor bump is justified. New executable capability, a new reference section, new schema, no breaking valid-input change. Public language limits itself to capability and disclaims measured superiority, deployment gain, growth, and native activation. The "build requests produce working adapters" wording rests on one synthetic fixture build, which the notes state. I found no overclaim. Requested versus observable identity: the user requested Claude Code Fable 5.1 at xhigh effort with permission bypass. My environment reports model ID CLI completion metadata{
"type": "result",
"subtype": "success",
"is_error": false,
"duration_ms": 760825,
"num_turns": 48,
"total_cost_usd": 7.5074775,
"usage": {
"input_tokens": 354,
"cache_creation_input_tokens": 211790,
"cache_read_input_tokens": 1114750,
"output_tokens": 59789,
"output_tokens_details": {
"thinking_tokens": 38076
},
"server_tool_use": {
"web_search_requests": 0,
"web_fetch_requests": 0
},
"service_tier": "standard",
"cache_creation": {
"ephemeral_1h_input_tokens": 211790,
"ephemeral_5m_input_tokens": 0
},
"inference_geo": "not_available",
"iterations": [
{
"input_tokens": 32,
"output_tokens": 3893,
"cache_read_input_tokens": 213241,
"cache_creation_input_tokens": 2506,
"cache_creation": {
"ephemeral_5m_input_tokens": 0,
"ephemeral_1h_input_tokens": 2506
},
"type": "message"
}
],
"speed": "standard"
},
"modelUsage": {
"claude-fable-5-1": {
"inputTokens": 354,
"outputTokens": 59789,
"cacheReadInputTokens": 1114750,
"cacheCreationInputTokens": 211790,
"webSearchRequests": 0,
"costUSD": 7.5074775,
"contextWindow": 1000000,
"maxOutputTokens": 64000,
"thinkingTokens": 38076,
"canonicalModel": "claude-fable-5-1",
"provider": "firstParty",
"costBasis": "list"
}
},
"session_id": "69cefb52-f4c1-404b-b3ae-57c31d34f565",
"permission_denials": []
} |
Final Fable acceptance — addf6c5Requested Claude Code Fable 5.1 / xhigh / permission bypass. Observed startup Reviewed commit: Requested vs observable model/effort: the session reports model What I ran
Independent probes, run on all three interpreters:
Fix-by-fix assessment
FindingsNo material defect. Two low, non-blocking observations:
Public claims in README, CHANGELOG, release notes, marketplace and citation are version-aligned at 0.7.0 and match what I measured. The README clone command and release-note links point at a Not run by meJekyll/Pages build and ASTRA REVIEW Coordinator acceptanceAccepted for release after personally inspecting the full accumulated changes and corrections; running 106 tests on Python 3.11.16, 3.12.13 and 3.14.7; matched Pages build/rendered checks; plugin validation; isolated exact-candidate installation and one-skill/no-other-components read-back; installed helper execution from The remaining Frozen candidate: CLI completion metadata{
"type": "result",
"subtype": "success",
"is_error": false,
"duration_ms": 473134,
"num_turns": 32,
"total_cost_usd": 4.347564500000001,
"usage": {
"input_tokens": 450,
"cache_creation_input_tokens": 122723,
"cache_read_input_tokens": 818818,
"output_tokens": 33678,
"output_tokens_details": {
"thinking_tokens": 15921
},
"server_tool_use": {
"web_search_requests": 0,
"web_fetch_requests": 0
},
"service_tier": "standard",
"cache_creation": {
"ephemeral_1h_input_tokens": 122723,
"ephemeral_5m_input_tokens": 0
},
"inference_geo": "not_available",
"iterations": [
{
"input_tokens": 32,
"output_tokens": 2602,
"cache_read_input_tokens": 126604,
"cache_creation_input_tokens": 2863,
"cache_creation": {
"ephemeral_5m_input_tokens": 0,
"ephemeral_1h_input_tokens": 2863
},
"type": "message"
}
],
"speed": "standard"
},
"modelUsage": {
"claude-fable-5-1": {
"inputTokens": 450,
"outputTokens": 33678,
"cacheReadInputTokens": 818818,
"cacheCreationInputTokens": 122723,
"webSearchRequests": 0,
"costUSD": 4.347564500000001,
"contextWindow": 1000000,
"maxOutputTokens": 64000,
"thinkingTokens": 15921,
"canonicalModel": "claude-fable-5-1",
"provider": "firstParty",
"costBasis": "list"
}
},
"session_id": "4d96b9da-9369-4faa-a428-f4cfac40610d",
"permission_denials": []
}API-EQUIVALENT COST RECEIPT: the two disjoint Fable CLI sessions report USD 11.855042 at their reported list-price basis. This is delegated-review-only, not the whole task, subscription charges or all-Astra savings. Native parent/Astra usage and supported Fable/cache-write repricing are unavailable. No same-token comparison is claimed. |
v0.7.0 publication receipt — 2026-09-22Published Augustus v0.7.0 Exact revisions and acceptance
The earlier delegated-scope replacement ledger is complete; historical reports, Verification
Published website and discovery surfacesPages deployment GitHub About read-back matches the find/build/evaluate/improve mission; The user has now requested a separate Stitch/OpenDesign-led website redesign; API-EQUIVALENT COST RECEIPT: two disjoint Fable review sessions reported |
Why 0.7.0
Augustus equips agents to find, build, evaluate and iteratively improve decision-model systems, including Software 3.0 workflows and prompt/program hill climbing. This release makes that capability explicit and operational instead of treating the skill as a survey or design-only guide. The earlier 0.6.1 patch plan was not published.
Changes
Evidence already run
make check: 106 tests plus repository, numerical/fingerprint and shell checks; local Python 3.11.16, 3.12.13 and 3.14.7.9b4d86d(requested GPT-6 Astra/ultra) returned fix-first: finite workflow mean overflow/underflow, selective-cost overflow, and log-loss cancellation. All three corrected with eight additional regression tests, including exact-ratio randomized comparisons, schema mutations and fixed-sample null enumeration. Fresh independent approval of the corrected revision is still required.977059ffor early normalization erasing paired deltas. Exact paired accumulation and two further tests corrected it in02f379a, then reviewed by Claude Code Fable 5.1/xhigh/YOLO at the user's request. Startup reportedclaude-fable-5-1andbypassPermissions; effective effort is unobservable.addf6c5includes eight more regressions. A fresh final Fable 5.1/xhigh/YOLO review returned ship on that exact candidate, with independent numerical probes and all three Python-version suites passing.Acceptance and release gates
Completed: coordinator integrated-diff inspection and checks, fresh final Fable ship verdict, exact-candidate/tag isolated installations and Skills CLI discovery, intended-head/main CI, merge/tag/release and deployed-commit/content read-back. See the publication receipt for exact revisions and limits. Published v0.7.0 at merge
ef8e035; the subsequent user-requested website redesign is separate work.Research/acceptance:
research/decision-model-review-2026-09-22.md,research/decision-engine-2026-09-22.md, andresearch/audits/2026-09-22-refresh-acceptance.md.Maintainer prompt:
research/prompts/maintainer.md.No third-party benchmark reproduction, external-directory submission, or claimed deployment gain. Model-based code review is separate from provider evaluation. The external research scheduler remains unverified. Whole-session API-equivalent cost unavailable: native tools did not expose complete usage.