Repository navigation
Prompt building Tier 0 — echo-don't-compute improvements (tracking) #29
Copy link
Copy link
Open
Labels
enhancementNew feature or requestNew feature or request
Description
Activity
Coordination update after reviewing #30 (the prompt-refactor/server-compose PR):
- The "whichever builder survives" question is answered: Real composition behind POST /api/compose: bundle the SDK, key stays server-side #30 deletes the app-side six-section prompt contract (
packages/app/src/lib/compose/prompt.ts,grammar.ts,templates.ts). The SDK'sbuildComposeMessages(packages/arbitration-sdk/src/compose.ts) is the one surviving prompt builder — every integration note in Prompt T0.1 — deterministic risk-appetite classifier, rendered as one prompt line #24–Prompt T0.5 (spike) — baseline-example anchoring: measure compliance gain vs parroting #28 now points there. - Prompt T0.5 (spike) — baseline-example anchoring: measure compliance gain vs parroting #28's "after the prompt-refactor PR" gate = after Real composition behind POST /api/compose: bundle the SDK, key stays server-side #30 merges. Bonus: Real composition behind POST /api/compose: bundle the SDK, key stays server-side #30 returns the exact last-attempt
messagesin the API response, which is precisely the observability the spike needs. - Prompt T0.3 — three risk-tiered strategies on one pair, band width as the risk axis #26's request plumbing is essentially done — the route already sends
maxStrategies: 3as server policy. - New follow-up item this set inherits: Real composition behind POST /api/compose: bundle the SDK, key stays server-side #30 orphans
PROMPT_VERSION(the files its comment references are deleted, and the surviving SDK builder carries no version — F2 §9 requires one in every trace). ReintroducepromptVersioninbuildComposeMessages/ServerComposeResult; until then, the "bump PROMPT_VERSION" notes in Prompt T0.1 — deterministic risk-appetite classifier, rendered as one prompt line #24–Prompt T0.3 — three risk-tiered strategies on one pair, band width as the risk axis #26 and Prompt T0.5 (spike) — baseline-example anchoring: measure compliance gain vs parroting #28 read as "add it, then bump it".
- The "whichever builder survives" question is answered: Real composition behind POST /api/compose: bundle the SDK, key stays server-side #30 deletes the app-side six-section prompt contract (
Tier 0 status:
- Prompt T0.4 — rejection feedback that states the fix, not just the rule #27 → Validator feedback names the fix, not just the rule (T0.4) #38 (validator feedback names the fix)
- Prompt T0.1 — deterministic risk-appetite classifier, rendered as one prompt line #24, Prompt T0.2 — precompute pairing arithmetic in the prompt; the model echoes amounts, never derives them #25, Prompt T0.3 — three risk-tiered strategies on one pair, band width as the risk axis #26 → Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 (appetite, reference pairing, band tiers;
PROMPT_VERSION→sluice.compose/3) - Prompt T0.5 (spike) — baseline-example anchoring: measure compliance gain vs parroting #28 → still open, blocked only on a funded Galileo key. Notes added there on what Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 changes about the method.
The
PROMPT_VERSIONfollow-up this tracking issue inherited from #30 is closed — it was reintroduced in the SDK builder by #32 and is bumped by #39, so "add it, then bump it" is done.Two things surfaced while building that are worth carrying forward rather than losing in a PR body:
- Volatility units are ambiguous and it matters ~3.5x. Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 reads
realizedVol7dPctas annualised (the other reading gives ±87%–±99% bands — full-range in all but name). This must be confirmed when F3 Open Q2 lands a real volatility source; the field may want renaming then. - Off-mid pricing has no invariant. Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 prevents it at the prompt level by precomputing the pairing, but nothing rejects a strategy whose
virtualAmountsratio is far from mid. That is the Tier 2 invariant (I17 in the original plan) and it is worth its own issue whenever multi-pair work starts.
Tier 0 is complete except for the measurement.
issue state #24 T0.1 appetite ✅ #39 #25 T0.2 pairing ✅ #39 #26 T0.3 band tiers ✅ #39 #27 T0.4 feedback names the fix ✅ #38 #28 T0.5 spike 🟡 harness in #40; numbers need a funded key The
PROMPT_VERSIONfollow-up inherited from #30 is closed: reintroduced by #32, nowsluice.compose/3.Three things Tier 0 surfaced that outlive it, none of which belong to this tracking issue but all of which would otherwise be lost in PR bodies:
- Band tiers are invisible in the UI. Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 computes wide/mid/tight server-side, but
ServerComposeResultdoesn't carry the tier — so the screen shows three band percentages with no framing, andfrom-server.tsstill fabricates a risk chip from band presence while the app README says risk ratings are unavailable until Gate 2. Worth one issue covering both. - Volatility units are ambiguous and it matters ~3.5x — Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 reads
realizedVol7dPctas annualised; confirm when F3 Q2 lands a real source. - Off-mid pricing still has no invariant. Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 makes it unlikely at the prompt level; only a validator rule makes it impossible (the Tier 2 I17).
I'd close this tracking issue once #28 has its numbers.
- Band tiers are invisible in the UI. Prompt Tier 0: derive risk appetite, pairing and band tiers (T0.1–T0.3) #39 computes wide/mid/tight server-side, but
Metadata
Metadata
Assignees
Labels
enhancementNew feature or requestNew feature or request
Prompt building Tier 0 — echo-don't-compute improvements (tracking)
Tier 0 of the prompt-building improvement ladder: everything here is prompt copy + pure deterministic modules — zero new data dependencies, zero extra inference round trips, no validator-invariant changes, no schema changes. Later tiers (deterministic APR-per-tier modeling; multi-pair candidates from held tokens; fine-tune follow-up prompts) build on these.
The constraint that shapes all of it
The composer model is
qwen/qwen2.5-omni-7b— a 7B model withMAX_COMPOSE_ATTEMPTS = 2and a user watching a spinner (latency is user-facing, F2 Q3). It has already been caught fabricating chainIds and deadlines until the prompt said "copy these EXACTLY" (compose.ts). So the tier's principle: deterministic code computes (classification, arithmetic, rankings, corrective values); the model picks from menus and echoes. Every issue below is that principle applied to one seam.Issues
Suggested order: #27 first (fully independent, validator-side), #24/#25 in parallel (pure modules), then #26 (composes them), #28 last.
Coordination
A PR refactoring prompt building is in flight. All pure modules (appetite, pairing math, band tiers, validator messages) are safe to build immediately; the prompt-rendering integration in each issue targets whichever builder that PR leaves standing (
packages/arbitration-sdk/src/compose.tsvs the F2 §9 six-section contract inpackages/app/src/lib/compose/prompt.ts). BumpPROMPT_VERSIONon any prompt-contract change, per F2 §9.Per repo rule (CLAUDE.md): decisions made while building — the appetite lexicon, the band-tier rubric, the T0.5 go/no-go — go back to the Notion feature pages (F2 §9, F3 band logic), not only into the PRs.
References