Skip to content

One communication method: the consulting pyramid as the mandatory top layer, per-surface derivation additive-only (closes #832) - #833

Merged
atomchung merged 4 commits into
mainfrom
claude/issue-832-one-communication-method
Aug 22, 2026
Merged

One communication method: the consulting pyramid as the mandatory top layer, per-surface derivation additive-only (closes #832)#833
atomchung merged 4 commits into
mainfrom
claude/issue-832-one-communication-method

Conversation

@atomchung

Copy link
Copy Markdown
Owner

What this is

One communication method. docs/expression-contract.md gains a mother chapter
(§3) that states the shape of every user-visible answer once, and the six
surfaces that each stated it in their own words now derive from it instead.

Why now. #830 deleted the obligation whitelist, and its post-merge rerun
(receipts in PR #831) fixed composition while leaving first-answer length
within ±5% of the pre-trim arm in all four frozen scenes. Deleting obligations
vacates space; nothing positive said what an answer is, so the space refilled
with discretionary elaboration. And the norm that could have said it existed
five times, written five different ways — which is drift by construction.

The mother law (§3)

  • Top — one sentence: the stance and the reason that decides it.
  • Middle — only blocks that pass the increment gate: every block adds a
    NEW decision-relevant fact or judgment. The test is not "is this true" or "is
    this owed" (under the whitelist era everything printed was both); it is
    delete this block — does the decision change?
  • Bottom — evidence and expansions stay in the data layer and surface on
    request. The answer says once that expansion is available.
  • End — one compact caliber block (D7, unchanged).
  • Voice — one paragraph: a senior analyst speaking to their own principal.

Four named bans, each observed owner-live in the #827/#830 runs, each carrying
the slug the exemplar corpus references it by: manufactured_scenario,
default_as_insight, hedging_couplet, restated_point.

Derivation is additive-only, and empty is the default

Surface What it adds to §3
Review card The document incarnation — keynote = top floor, three middle blocks = the middle, the Block-1 footnote = the end block. Its structure does not change.
consider Exactly two parameters: lead selection, and which blocks the middle floor may hold.
No recorded book Two: with no book the top sentence is a research-backed baseline; the strategy-class map is a middle-floor block set.
Weekly market read One: its optional question comes after the complete brief.
Freeform answers Nothing — and it says so. Text-first is a latency default, never a shape.
output-voice.md V1 Keeps its ID as the failure class its fixtures and cross-host rulings cite; its own definition is marked superseded and its shape half routes to §3.

Registry freeze

V, D and C take no new ID for a shape or length concern. A style fix has
two lanes: amend §3 (owner ruling, logged) or add an exemplar/counter-exemplar
(day-to-day). V10 stays unallocated. Recorded in §3.4, in output-voice.md,
and in the maintainer guide's new #832 mirrored-surfaces row.

No character-count cap. #543's ceiling was deleted by #827 and stays
deleted; test_no_character_count_cap_came_back keeps it that way.

Exemplars are the spec

tests/agent/expression-witnesses.json (schema 2) carries 14 canonical
exemplars across the four conversational surfaces — three to five each, gated —
each declaring the one-sentence answer it leads with and the increment every
one of its blocks adds. The three owner-approved acceptance templates from #830
are three of them. All issuers fictional (WDGT, GRDC, FABR, ACME); nothing is
derived from a user record.

Two new assertions in tests/agent/check_expression.py, both honest about
their half:

  • E-7 fails a scene whose declared core is not in its opening block. Its
    negative witness bloat_no_core is a complete, anchored, obligation-
    discharging answer to the same call template 1 answers — and it never says
    which candidate to buy.
  • E-8 fails a scene whose declared blocks are not all present in order, or
    which declares no increment for a block, or the same increment twice.
    Negative witnesses: the all-in-one-name simulation the owner called
    meaningless (a block with no increment), and a closing summary that is the
    opening judgment in a second form.

The coverage boundary is asserted, not promised. default_as_insight and
hedging_couplet are well-formed blocks carrying true content; nothing
mechanical reaches them. Their counter scenes must PASS every assertion,
so the day an honest oracle for one of them exists, that scene is what says the
boundary moved.

Issue coverage

closes #832. Every acceptance box:
mother chapter exists and all surfaces reference it, with zero local
restatements (grep-gated); increment gate and the four bans encoded, with a
no-core witness and a manufactured-scenario counter-exemplar that both fail;
exemplar sets per surface, fictional only, oracles derived from them; integrity
gates untouched (engine-owned numbers, provenance, canonical writes, execution
truth, privacy — no engine behaviour changed, one comment aside); no
character-count cap; both suites green.

#579 — NOT closed. Verified rather than assumed. Its three layers:
deterministic rule_effect derivation, product-safe projection, agent
realization. Layers 1 and 2 shipped before this PR —
consequence.classify_rule_effect, RULE_EFFECTS, the must_convey /
must_not_convey semantic slots, both schema enums, and
improved_but_still_over / resolved_existing_breach handling in
trade-consequence.md. This PR strengthens layer 3 only: every rule_effects
entry is middle-floor content that is never traded away, said in both SKILL.md
and the consider derivation. #579's remaining acceptance items are an
owner-live consider rerun confirming the distinction is immediately
understandable, plus independent review — neither is something a PR can close.

#778 — NOT closed. Three of its four acceptance criteria are delivered:

#590 — cross-reference, not closed. The increment gate is the judging
criterion it asked for and is now stated where a rubric can cite it. Its own
eval work (a fifth judge axis, or the may_state_total receipt — PR #831
follow-up items 1 and 2) is not done here.

#676 — cross-reference, not closed. V1 is re-scoped to a failure class and
routes its shape half to §3; the registry is otherwise unchanged. #676's
verdict remains its own owner-live acceptance.

#829 — untouched and out of scope, as #832 directed. The still-live
sentence is unchanged.

Spec interpretations

  1. Five phrasings turned out to be six. trade-consequence.md's "reader's
    question chain" section — added by [design·M1] Owner-live: answers exceed the reading budget — verbosity is the next usefulness bottleneck after #827 #830, after [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832's list was written — is
    an independent phrasing of the same idea. Leaving it would have failed the
    issue's own "zero remaining restatements", so it became the consider
    derivation.
  2. SKILL.md keeps the shape inline, labelled as a projection. A bare
    reference would put the shape behind a file the always-loaded layer does not
    load. It states the four floors in the mother chapter's own terms and says
    it is a projection, not a second wording. Net bytes fell: 8,522 → 8,512, and
    the always-loaded pair is 16,320 against its 16,384 budget.
  3. V1 keeps its ID and its old sentence, marked superseded, because
    fixtures and cross-host rulings address it by ID and [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832 asks that
    historical rows survive.
  4. Which bans get an oracle. A manufactured scenario is structurally a
    block with no increment, and a restated point is one increment declared
    twice — both reachable by E-8, which is what makes [design·M1] One communication method: the consulting pyramid as the mandatory top layer; per-surface derivation optional and additive-only #832's "a
    manufactured-scenario counter-exemplar fails" true. I did not build a
    keyword oracle for the other two; that is exactly the English-only
    fragility PR feat(challenge): delete the obligation whitelist, keep the seven facts that earn their place (closes #830) #831 flagged.
  5. E-7/E-8 read a declared scene, not arbitrary prose. The corpus is the
    spec, so its oracle reads the corpus. It decides that the declared core
    really leads and that the declared decomposition is faithful and
    non-repeating; it decides nothing about whether the core is the right call.
    The CLI path still runs E-5/E-6 against any answer and prints which
    assertions it did not run.
  6. test_research_priors.py was re-pinned, not restored. It asserted the
    literal sentence "lead with the bounded value already supported". Restoring
    it would defeat the change; dropping the assertion would lose a real
    protection. It now pins the protection harder — baseline → map → question,
    in that order, with the question last asserted rather than implied, plus the
    derivation declaration.

Tests

python3.12 tests/run_all.py --group product   → PASS: all 48 suites passed
python3.12 tests/run_all.py --group qa-eval   → PASS: all 11 suites passed

Follow-ups (owner's call, not opened here)

🤖 Generated with Claude Code

test and others added 4 commits August 22, 2026 10:27
…s product speaks

#830 deleted the obligation whitelist and its post-merge rerun (PR #831) moved
first-answer length by less than 5% in all four frozen scenes. Deleting
obligations vacates space; nothing positive said what an answer *is*, so the
space refilled with discretionary elaboration. Meanwhile the answer-first
principle existed five times, written five different ways, which is drift by
construction.

expression-contract.md gains section 3, the one statement of the shape every
user-visible answer takes: one-sentence answer on top, an increment-gated
middle (delete a block; if the decision does not change, delete it), the rest
of the computed inventory behind a single offer, one caliber block at the end,
and one paragraph of voice. Four named bans are encoded with the slugs the
exemplar corpus references them by. Derivation by a surface is additive-only
and an empty derivation is the default.

V, D and C are frozen for shape and length: no new ID for "answers are too
long" or "lead with X", because a sixth phrasing with an ID on it is still a
sixth. V1 keeps its ID as the failure class its fixtures and cross-host
rulings cite, marks its own definition superseded, and routes its shape half
to the mother chapter. V10 stays unallocated.

No character-count cap: #543's ceiling was deleted by #827 and stays deleted.
Length is the shape's consequence, not its rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ivation each

Every surface that stated the answer shape in its own words now derives it
from the mother chapter and keeps only what it genuinely adds:

- the review card is the document incarnation — keynote = top floor, the three
  middle blocks = the middle, the Block-1 footnote = the end block. Its
  structure does not change; the derivation is recorded so it stops reading as
  an independent statement of answer-first.
- `consider` adds exactly two parameters: which fact wins the top sentence
  (lead selection) and which blocks the middle floor may hold (answer slots).
  Its reader's-question-chain section, the sixth phrasing the audit found,
  becomes that derivation.
- no recorded book adds two: with no book the top sentence is a
  research-backed baseline, and the strategy-class map is a middle-floor block
  set.
- the weekly market read adds one: its optional question comes after the
  complete brief.
- freeform answers add nothing, and say so. An empty derivation is valid and
  is the expected case; text-first is a latency default, never a shape.
- SKILL.md keeps the shape inline because it is always loaded, and is labelled
  as the mother chapter's projection rather than a second wording. It lands at
  8,512 bytes, below its previous 8,522, and the always-loaded pair stays
  inside its budget.

A funding shortfall now outranks every other lead candidate on `consider`
(#778): a negative post-trade cash balance is not a portfolio consequence, it
says the trade cannot be done out of the recorded book. The two numbers are
support; the decision they imply is the answer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… them

The binding statement of the answer pyramid is the corpus, not the prose.
`tests/agent/expression-witnesses.json` becomes schema 2: 14 canonical
exemplars across the four conversational surfaces (three to five each, gated),
each declaring the one-sentence answer it leads with and the increment every
one of its blocks adds. The owner-approved acceptance templates from #830 are
three of them; every issuer is fictional (WDGT, GRDC, FABR, ACME) and nothing
is derived from a user record.

Two new assertions, both honest about their half:

- E-7 fails a scene whose declared core is not in its opening block. Its
  negative witness is a complete, anchored, obligation-discharging answer to
  the same call template 1 answers, which never says which candidate to buy.
- E-8 fails a scene whose declared blocks are not all present in order, or
  which declares no increment for a block, or the same increment twice. Its
  negative witnesses are the manufactured all-in-one-name simulation (a block
  with no increment) and a closing summary that is the opening judgment in a
  second form.

The two bans nothing mechanical reaches — a system default explained as
insight, and a hedging couplet — get `counter` scenes that must PASS every
assertion. The coverage boundary is asserted rather than promised: the day an
oracle can catch one of them, that scene is what says so.

#778's delivered answer joins as a second E-7 witness: the two cash numbers
stated, the funding decision never, the opening spent on the boundary and
hedging.

`tests/test_expression_contract.py` gains the grep-checkable acceptance —
zero retired answer-first phrasings remain, every surface declares a
derivation, every named ban is defined in the mother chapter, the registry
freeze is recorded in both the contract and the maintainer route, and no
character-count cap came back.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`test_research_priors.py` asserted the literal sentence "lead with the bounded
value already supported", which is one of the six answer-first phrasings #832
replaces with a derivation. Restoring the sentence would defeat the change;
dropping the assertion would lose a real protection (#597/#598: the user sees
the bounded value before any intake question).

So it pins the protection harder instead. The route's own block order must run
baseline -> strategy-class map -> question, in that order — the question being
last is now asserted rather than implied — and the section must declare itself
a derivation of the shape rather than a second statement of it. The numeric
question-cap regression still reddens.

Recorded on the #832 mirrored-surfaces row, since a test that pins prose by
literal is exactly the kind of reader a shape change has to carry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@atomchung
atomchung merged commit 6aabbcd into main Aug 22, 2026
5 checks passed
@atomchung

Copy link
Copy Markdown
Owner Author

Post-merge live measurement #2 (sanitized): the length curve finally bends — and the one regression pinpoints the residual force

Re-ran the same four frozen scenes against merged main 6aabbcd, same model and inputs as every prior round.

First-answer / conversation-total zh chars, loosened arm → deletion round → this round:

What visibly worked: the named bans bite (no manufactured all-in table in the cash scene — the same insight now costs one sentence; hedging couplets largely gone; coverage scope stated up front in discovery). Writes were fully clean this round: exploration 0 rows everywhere, discovery's final pick 1 row only after the user explicitly delegated the choice (disclosed in-answer), selection scene exact-1.

The regression is diagnostic: the only scene that got longer is the no-book freeform one — exactly the surface with an explicitly empty derivation and no engine payload anchoring the answer, and the only one that opened with a process line ("以下是我的回覆。"). Everywhere a concrete derivation + engine anchor exists, the pyramid holds; where guidance is most abstract, the model free-writes.

Residual force, per #832's stop condition (no new rules; owner decides): the exemplars are wired into QC (expression-witnesses) but never enter the generation path — and examples only steer output when the model sees them at generation time. Live answers run 3–7× the ~300-char acceptance templates. Options for the owner: (a) put one ~300-char mini-exemplar per surface into the always-loaded layer (direct fix, costs bytes), (b) accept current scale and judge by real usage, (c) add the demoted reading-budget oracle as a deterministic backstop.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment