One communication method: the consulting pyramid as the mandatory top layer, per-surface derivation additive-only (closes #832) - #833
Conversation
…s product speaks #830 deleted the obligation whitelist and its post-merge rerun (PR #831) moved first-answer length by less than 5% in all four frozen scenes. Deleting obligations vacates space; nothing positive said what an answer *is*, so the space refilled with discretionary elaboration. Meanwhile the answer-first principle existed five times, written five different ways, which is drift by construction. expression-contract.md gains section 3, the one statement of the shape every user-visible answer takes: one-sentence answer on top, an increment-gated middle (delete a block; if the decision does not change, delete it), the rest of the computed inventory behind a single offer, one caliber block at the end, and one paragraph of voice. Four named bans are encoded with the slugs the exemplar corpus references them by. Derivation by a surface is additive-only and an empty derivation is the default. V, D and C are frozen for shape and length: no new ID for "answers are too long" or "lead with X", because a sixth phrasing with an ID on it is still a sixth. V1 keeps its ID as the failure class its fixtures and cross-host rulings cite, marks its own definition superseded, and routes its shape half to the mother chapter. V10 stays unallocated. No character-count cap: #543's ceiling was deleted by #827 and stays deleted. Length is the shape's consequence, not its rule. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ivation each Every surface that stated the answer shape in its own words now derives it from the mother chapter and keeps only what it genuinely adds: - the review card is the document incarnation — keynote = top floor, the three middle blocks = the middle, the Block-1 footnote = the end block. Its structure does not change; the derivation is recorded so it stops reading as an independent statement of answer-first. - `consider` adds exactly two parameters: which fact wins the top sentence (lead selection) and which blocks the middle floor may hold (answer slots). Its reader's-question-chain section, the sixth phrasing the audit found, becomes that derivation. - no recorded book adds two: with no book the top sentence is a research-backed baseline, and the strategy-class map is a middle-floor block set. - the weekly market read adds one: its optional question comes after the complete brief. - freeform answers add nothing, and say so. An empty derivation is valid and is the expected case; text-first is a latency default, never a shape. - SKILL.md keeps the shape inline because it is always loaded, and is labelled as the mother chapter's projection rather than a second wording. It lands at 8,512 bytes, below its previous 8,522, and the always-loaded pair stays inside its budget. A funding shortfall now outranks every other lead candidate on `consider` (#778): a negative post-trade cash balance is not a portfolio consequence, it says the trade cannot be done out of the recorded book. The two numbers are support; the decision they imply is the answer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… them The binding statement of the answer pyramid is the corpus, not the prose. `tests/agent/expression-witnesses.json` becomes schema 2: 14 canonical exemplars across the four conversational surfaces (three to five each, gated), each declaring the one-sentence answer it leads with and the increment every one of its blocks adds. The owner-approved acceptance templates from #830 are three of them; every issuer is fictional (WDGT, GRDC, FABR, ACME) and nothing is derived from a user record. Two new assertions, both honest about their half: - E-7 fails a scene whose declared core is not in its opening block. Its negative witness is a complete, anchored, obligation-discharging answer to the same call template 1 answers, which never says which candidate to buy. - E-8 fails a scene whose declared blocks are not all present in order, or which declares no increment for a block, or the same increment twice. Its negative witnesses are the manufactured all-in-one-name simulation (a block with no increment) and a closing summary that is the opening judgment in a second form. The two bans nothing mechanical reaches — a system default explained as insight, and a hedging couplet — get `counter` scenes that must PASS every assertion. The coverage boundary is asserted rather than promised: the day an oracle can catch one of them, that scene is what says so. #778's delivered answer joins as a second E-7 witness: the two cash numbers stated, the funding decision never, the opening spent on the boundary and hedging. `tests/test_expression_contract.py` gains the grep-checkable acceptance — zero retired answer-first phrasings remain, every surface declares a derivation, every named ban is defined in the mother chapter, the registry freeze is recorded in both the contract and the maintainer route, and no character-count cap came back. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`test_research_priors.py` asserted the literal sentence "lead with the bounded value already supported", which is one of the six answer-first phrasings #832 replaces with a derivation. Restoring the sentence would defeat the change; dropping the assertion would lose a real protection (#597/#598: the user sees the bounded value before any intake question). So it pins the protection harder instead. The route's own block order must run baseline -> strategy-class map -> question, in that order — the question being last is now asserted rather than implied — and the section must declare itself a derivation of the shape rather than a second statement of it. The numeric question-cap regression still reddens. Recorded on the #832 mirrored-surfaces row, since a test that pins prose by literal is exactly the kind of reader a shape change has to carry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge live measurement #2 (sanitized): the length curve finally bends — and the one regression pinpoints the residual forceRe-ran the same four frozen scenes against merged main First-answer / conversation-total zh chars, loosened arm → deletion round → this round:
What visibly worked: the named bans bite (no manufactured all-in table in the cash scene — the same insight now costs one sentence; hedging couplets largely gone; coverage scope stated up front in discovery). Writes were fully clean this round: exploration 0 rows everywhere, discovery's final pick 1 row only after the user explicitly delegated the choice (disclosed in-answer), selection scene exact-1. The regression is diagnostic: the only scene that got longer is the no-book freeform one — exactly the surface with an explicitly empty derivation and no engine payload anchoring the answer, and the only one that opened with a process line ("以下是我的回覆。"). Everywhere a concrete derivation + engine anchor exists, the pyramid holds; where guidance is most abstract, the model free-writes. Residual force, per #832's stop condition (no new rules; owner decides): the exemplars are wired into QC (expression-witnesses) but never enter the generation path — and examples only steer output when the model sees them at generation time. Live answers run 3–7× the ~300-char acceptance templates. Options for the owner: (a) put one ~300-char mini-exemplar per surface into the always-loaded layer (direct fix, costs bytes), (b) accept current scale and judge by real usage, (c) add the demoted reading-budget oracle as a deterministic backstop. |
What this is
One communication method.
docs/expression-contract.mdgains a mother chapter(§3) that states the shape of every user-visible answer once, and the six
surfaces that each stated it in their own words now derive from it instead.
Why now. #830 deleted the obligation whitelist, and its post-merge rerun
(receipts in PR #831) fixed composition while leaving first-answer length
within ±5% of the pre-trim arm in all four frozen scenes. Deleting obligations
vacates space; nothing positive said what an answer is, so the space refilled
with discretionary elaboration. And the norm that could have said it existed
five times, written five different ways — which is drift by construction.
The mother law (§3)
NEW decision-relevant fact or judgment. The test is not "is this true" or "is
this owed" (under the whitelist era everything printed was both); it is
delete this block — does the decision change?
request. The answer says once that expansion is available.
Four named bans, each observed owner-live in the #827/#830 runs, each carrying
the slug the exemplar corpus references it by:
manufactured_scenario,default_as_insight,hedging_couplet,restated_point.Derivation is additive-only, and empty is the default
consideroutput-voice.mdV1Registry freeze
V, D and C take no new ID for a shape or length concern. A style fix has
two lanes: amend §3 (owner ruling, logged) or add an exemplar/counter-exemplar
(day-to-day). V10 stays unallocated. Recorded in §3.4, in
output-voice.md,and in the maintainer guide's new #832 mirrored-surfaces row.
No character-count cap. #543's ceiling was deleted by #827 and stays
deleted;
test_no_character_count_cap_came_backkeeps it that way.Exemplars are the spec
tests/agent/expression-witnesses.json(schema 2) carries 14 canonicalexemplars across the four conversational surfaces — three to five each, gated —
each declaring the one-sentence answer it leads with and the increment every
one of its blocks adds. The three owner-approved acceptance templates from #830
are three of them. All issuers fictional (WDGT, GRDC, FABR, ACME); nothing is
derived from a user record.
Two new assertions in
tests/agent/check_expression.py, both honest abouttheir half:
negative witness
bloat_no_coreis a complete, anchored, obligation-discharging answer to the same call template 1 answers — and it never says
which candidate to buy.
which declares no increment for a block, or the same increment twice.
Negative witnesses: the all-in-one-name simulation the owner called
meaningless (a block with no increment), and a closing summary that is the
opening judgment in a second form.
The coverage boundary is asserted, not promised.
default_as_insightandhedging_coupletare well-formed blocks carrying true content; nothingmechanical reaches them. Their
counterscenes must PASS every assertion,so the day an honest oracle for one of them exists, that scene is what says the
boundary moved.
Issue coverage
closes #832. Every acceptance box:
mother chapter exists and all surfaces reference it, with zero local
restatements (grep-gated); increment gate and the four bans encoded, with a
no-core witness and a manufactured-scenario counter-exemplar that both fail;
exemplar sets per surface, fictional only, oracles derived from them; integrity
gates untouched (engine-owned numbers, provenance, canonical writes, execution
truth, privacy — no engine behaviour changed, one comment aside); no
character-count cap; both suites green.
#579 — NOT closed. Verified rather than assumed. Its three layers:
deterministic
rule_effectderivation, product-safe projection, agentrealization. Layers 1 and 2 shipped before this PR —
consequence.classify_rule_effect,RULE_EFFECTS, themust_convey/must_not_conveysemantic slots, both schema enums, andimproved_but_still_over/resolved_existing_breachhandling intrade-consequence.md. This PR strengthens layer 3 only: everyrule_effectsentry is middle-floor content that is never traded away, said in both SKILL.md
and the
considerderivation. #579's remaining acceptance items are anowner-live
considerrerun confirming the distinction is immediatelyunderstandable, plus independent review — neither is something a PR can close.
#778 — NOT closed. Three of its four acceptance criteria are delivered:
the reasoning that the balance and its weight are support and the decision
they imply is the answer;
recites the balance without naming the funding decision
(
funding_shortfall_recited_not_decidedfails E-7, paired with the positiveconsider_funding_shortfall). It lives in the expression corpus rather thanas a
computed_considervoice fixture becausecheck_voice.classify_failureis English-keyword-only (PR feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831 follow-up item 4), so a zh-TW witness cannot
be classified there;
still stated once, beside the claim it qualifies, never as the lead, and
nothing assumes the user does or does not have other cash;
owed item (
topic: "funding"). That is an engine change toevaluation_challenge.build_challengeplus themust_statetopic enum,TOPICS, andtests/test_evaluation_challenge.py's ordered equality. [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832scoped itself to the expression layer and left the obligation floor exactly
where [design·M1] Owner-live: answers exceed the reading budget — verbosity is the next usefulness bottleneck after #827 #830 put it, so widening
must_statebelongs in its own change.#590 — cross-reference, not closed. The increment gate is the judging
criterion it asked for and is now stated where a rubric can cite it. Its own
eval work (a fifth judge axis, or the
may_state_totalreceipt — PR #831follow-up items 1 and 2) is not done here.
#676 — cross-reference, not closed. V1 is re-scoped to a failure class and
routes its shape half to §3; the registry is otherwise unchanged. #676's
verdict remains its own owner-live acceptance.
#829 — untouched and out of scope, as #832 directed. The still-live
sentence is unchanged.
Spec interpretations
trade-consequence.md's "reader'squestion chain" section — added by [design·M1] Owner-live: answers exceed the reading budget — verbosity is the next usefulness bottleneck after #827 #830, after [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832's list was written — is
an independent phrasing of the same idea. Leaving it would have failed the
issue's own "zero remaining restatements", so it became the
considerderivation.
reference would put the shape behind a file the always-loaded layer does not
load. It states the four floors in the mother chapter's own terms and says
it is a projection, not a second wording. Net bytes fell: 8,522 → 8,512, and
the always-loaded pair is 16,320 against its 16,384 budget.
fixtures and cross-host rulings address it by ID and [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832 asks that
historical rows survive.
block with no increment, and a restated point is one increment declared
twice — both reachable by E-8, which is what makes [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832's "a
manufactured-scenario counter-exemplar fails" true. I did not build a
keyword oracle for the other two; that is exactly the English-only
fragility PR feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831 flagged.
spec, so its oracle reads the corpus. It decides that the declared core
really leads and that the declared decomposition is faithful and
non-repeating; it decides nothing about whether the core is the right call.
The CLI path still runs E-5/E-6 against any answer and prints which
assertions it did not run.
test_research_priors.pywas re-pinned, not restored. It asserted theliteral sentence "lead with the bounded value already supported". Restoring
it would defeat the change; dropping the assertion would lose a real
protection. It now pins the protection harder — baseline → map → question,
in that order, with the question last asserted rather than implied, plus the
derivation declaration.
Tests
Follow-ups (owner's call, not opened here)
fundingowed item that would close [bug·voice·consider·P2] A negative post-trade cash balance is owed as two numbers but not as the decision it implies — the answer leads with the boundary instead of "this needs funding or a sale" #778.check_voice.classify_failureis English-keyword-only, so the voice oraclecannot classify a zh-TW answer at all (PR feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831 follow-up item 4). This PR
routed around it; the gap is still there.
concentration family and the unchecked list (PR feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831 follow-up item 3) — now
inconsistent with the pyramid as well as with [design·M1] Owner-live: answers exceed the reading budget — verbosity is the next usefulness bottleneck after #827 #830's shape.
exemplar scale with the mother law in place, do not add rules — report
which residual force drives length and return to the owner.
🤖 Generated with Claude Code