feat: Add the log_prob return mode and move the envelope to schema v4 - #17
Conversation
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.
| " history,\n", | ||
| " query_times=SCORED_QUERY_TIMES,\n", | ||
| " query_rows=pd.DataFrame([outcome], columns=EXPECTED_COLUMNS),\n", | ||
| " requested_columns=READABLE_COLUMNS,\n", |
There was a problem hiding this comment.
Notebook scores with old projection
Medium Severity
The new scoring cell passes READABLE_COLUMNS, which drops the pinned column. v4 log_prob validation requires every declared column in declared order, so this call raises before it reaches the service. The same leftover readable-column rule is still in the forecast_log_prob docstring.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.


Two changes to the served contract, both breaking for existing callers, so the
envelope moves to
schema_versionv4in one step.1.
log_probreturn modeThe one return mode that answers a question about values the caller already
has.
JointFMClient.forecast_log_prob(...)takesquery_rows— one observedrow per entry of
query_times, each carrying a value for every declaredcolumn — and returns the model's log density of those values under its joint
at that position.
LogProbResult/LogProbScorescarry per-query-timevaluesandnll_valuesplustotal,mean,nll_totalandnll_mean. The summariesare derivable from
values, so the client recomputes and checks them ratherthan trusting them (
DERIVED_SCORE_TOLERANCE, relative and absolute at1e-6, because the service sums in the model's own precision).ConditionBlock: the score is taken under the conditional atthe position the block names, and the deployment refuses a row that
contradicts the condition — a pinned column given another value, or a
bounded column outside its range — instead of scoring it.
query_rowsgoes through the same column encoding as the history(
adapters.py), so a categorical label is mapped the same way in both.query_rowsmust appear together, rows must align one-to-one withquery_times, and every declared column needs a finite observed value.2.
schema_versionv4 — one projection rule for every query moderequested_columnsis now the only thing that decides what a responsecarries, and it is independent of roles and of conditions.
requested_columnsomittedlog_probprojection ruleThe point is comparability: a conditioned answer and an unconditioned forecast
over the same schema now line up column for column without either request
restating the projection, so a scenario can be read against its base case
directly. A pinned column answers with the request's own constant, which is
what keeps the two frames aligned; a caller who does not want it simply does
not name it.
This moves the version because a response's column set changes while no field
changes shape — no client could detect it by parsing, and the service compares
schema_versionfor exact equality ahead of all other validation.Breaking changes
SCHEMA_VERSIONis"v4". This client is refused by a v3 deployment onevery request,
forecastincluded, and vice versa. Server and clientmust roll together.
as well. Positional consumers of the column axis must either pass
requested_columnsexplicitly or index by label viaresult.requested_columns.log_probrequest that previously omitted its pinned column fromrequested_columnsmust now list every declared column — or omit the field,which resolves to exactly that.
Testing
task pre-commitgreen: ruff, ty, typos, nbstripout, and the suite withcoverage. New
tests/test_log_prob_mode.py; request/response fixture pairs forboth
forecastandconditionscoring, exercised bytest_fixture_compatibility.py;tests/test_condition_mode.pyextended forthe projection rules (a pin readable in any position, omitting the projection
states nothing on the wire).
Notes for the reviewer
forecast_conditionnotebook gains a scoring section and thecandidate-condition ranking;
README.mdanddocs/api-reference.mdtrackthe new rules.
must be released before the examples repo can move off its local checkout —
four of the five examples now declare
jointfm-client>=0.9.0from PyPI.Note
High Risk
Breaking schema bump (v3 ↔ v4 hard mismatch) plus changed default
requested_columnssemantics; server and SDK must deploy together or forecasts and column alignment break silently for callers that assumed v3 behavior.Overview
Bumps the JointFM API envelope to
schema_version="v4"and adds first-classlog_probscoring in the Python SDK, with docs, samples, and tests aligned to the new contract.log_prob: NewJointFMClient.forecast_log_prob(...)(plusquery_rowsonforecast/ dataframe builders) sends observed rows and returnsLogProbResult/LogProbScores. Request validation tiesreturn_mode="log_prob"toquery_rows, enforces one row perquery_times, and requiresrequested_columnsto list every declared column in order. Responses parseoutputs.log_proband re-check summary fields againstvalues(DERIVED_SCORE_TOLERANCE). Works withConditionBlock; equality pins may appear in projections like other columns.v4 projection rules: Default omitted
requested_columnsis documented as all declared columns in declared order (not targets-only). Equality-pinned columns may be named inrequested_columns(client no longer rejects them). Config defaults (JOINTFM_SCHEMA_VERSION,config.sample.yaml, fixtures) and compatibility checks now expect v4 only.Docs / examples: README and
docs/api-reference.mddescribelog_prob, plausibility chaining, and v4 rules;forecast_condition.ipynbadds unconditional vs conditional comparison, draws, quantile bands, scenario ranking, and conditional log-prob scoring.Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.