Non-autoregressive decision engine. Typed decisions over text in a single parallel forward pass, returned as JSON probabilities. No LLM, no text generation: nothing to parse, nothing to hallucinate.
Full design, FLOPs budget and per-modality SLAs: docs/architecture.md.
uv venv && uv pip install -e ".[dev]"
python -m pytest -qfrom opendecision import Decider
d = Decider() # random weights until you train or load a checkpoint
r = d.predict({
"state": "Hi, we were billed twice for March. Refund the duplicate or we cancel.",
"questions": {
"department": {"type": "choice", "instructions": "Which department?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs", "other": "else"}},
"urgency": {"type": "score", "instructions": "How urgent?", "criteria": ["low", "medium", "high"]},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to leave?"},
},
})
print(r["answers"]["department"]["probabilities"])Response (same shape as Laya POST /v1/systemone):
{"model": "opendecision-0.1",
"answers": {"department": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.91, "technical": 0.04, "other": 0.05},
"unknown": 0.0, "confidence": 0.74, "answer_confidence": 0.91},
"churn_risk": {"type": "noul", "noul": 0.89, "probabilities": {"no": 0.11, "yes": 0.89}}},
"usage": {"input_tokens": 31, "output_tokens": 0}}| Stage | Function | Purpose |
|---|---|---|
| 0 Pretrain | train.mlm_step |
masked-token CPT; CLM->MLM 25/75 schedule for text (arXiv 2507.00994) |
| 1 Distill | train.distill_step |
offline-cached teacher distributions (Apache teachers only) |
| 2 Fine-tune | train.finetune_step |
strictly proper loss (log / Brier / spherical, RPS for ordinal) + optional coherence loss |
| 3 RL | train.rl_step |
only with outcome feedback: reward r = c - p_a, leave-one-out baseline (unbiased half-Brier gradient, tested) |
| 4 Calibrate | calibrate.fit_temperature, conformal_threshold |
per-type temperature on a disjoint split; conformal as an abstain gate, not "calibrated probabilities" |
Details: docs/training.md.
Pinned baseline version, same items and criteria, zero-shot and fine-tuned reported separately, raw and post-temperature rows, per-K, paired cluster bootstrap, frozen test split, contamination flags. docs/evaluation.md.
architecture · training · evaluation · serving · economics · long context · world knowledge · datasets & licences · research evidence
Builds on ideas from Laya (Apache-2.0), LAVOIR, eve-rlcd, ModernBERT, SigLIP 2, Perceiver IO. License: Apache-2.0.
