Skip to content

feat: Add the log_prob return mode and move the envelope to schema v4 - #17

Merged
shackmann merged 6 commits into
mainfrom
stefan/conditioning
Sep 23, 2026
Merged

shackmann merged 6 commits into
mainfrom
stefan/conditioning

Conversation

@shackmann

@shackmann shackmann commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Two changes to the served contract, both breaking for existing callers, so the
envelope moves to schema_version v4 in one step.

1. log_prob return mode

The one return mode that answers a question about values the caller already
has. JointFMClient.forecast_log_prob(...) takes query_rows — one observed
row per entry of query_times, each carrying a value for every declared
column — and returns the model's log density of those values under its joint
at that position.

  • LogProbResult / LogProbScores carry per-query-time values and
    nll_values plus total, mean, nll_total and nll_mean. The summaries
    are derivable from values, so the client recomputes and checks them rather
    than trusting them (DERIVED_SCORE_TOLERANCE, relative and absolute at
    1e-6, because the service sums in the model's own precision).
  • Works under a ConditionBlock: the score is taken under the conditional at
    the position the block names, and the deployment refuses a row that
    contradicts the condition — a pinned column given another value, or a
    bounded column outside its range — instead of scoring it.
  • query_rows goes through the same column encoding as the history
    (adapters.py), so a categorical label is mapped the same way in both.
  • Request-side validation fails before any round trip: the mode and
    query_rows must appear together, rows must align one-to-one with
    query_times, and every declared column needs a finite observed value.

2. schema_version v4 — one projection rule for every query mode

requested_columns is now the only thing that decides what a response
carries, and it is independent of roles and of conditions.

v3 v4
requested_columns omitted target columns only every declared column, declared order
equality-pinned column listed rejected allowed; reads back the pinned value
log_prob projection rule every readable column, pinned excused every declared column, nothing excused

The point is comparability: a conditioned answer and an unconditioned forecast
over the same schema now line up column for column without either request
restating the projection, so a scenario can be read against its base case
directly. A pinned column answers with the request's own constant, which is
what keeps the two frames aligned; a caller who does not want it simply does
not name it.

This moves the version because a response's column set changes while no field
changes shape — no client could detect it by parsing, and the service compares
schema_version for exact equality ahead of all other validation.

Breaking changes

  • SCHEMA_VERSION is "v4". This client is refused by a v3 deployment on
    every request, forecast included, and vice versa. Server and client
    must roll together.
  • Callers that relied on the targets-only default now receive feature columns
    as well. Positional consumers of the column axis must either pass
    requested_columns explicitly or index by label via
    result.requested_columns.
  • A log_prob request that previously omitted its pinned column from
    requested_columns must now list every declared column — or omit the field,
    which resolves to exactly that.

Testing

task pre-commit green: ruff, ty, typos, nbstripout, and the suite with
coverage. New tests/test_log_prob_mode.py; request/response fixture pairs for
both forecast and condition scoring, exercised by
test_fixture_compatibility.py; tests/test_condition_mode.py extended for
the projection rules (a pin readable in any position, omitting the projection
states nothing on the wire).

Notes for the reviewer

  • The forecast_condition notebook gains a scoring section and the
    candidate-condition ranking; README.md and docs/api-reference.md track
    the new rules.
  • This client pairs with the matching server change in the JointFM repo. It
    must be released before the examples repo can move off its local checkout —
    four of the five examples now declare jointfm-client>=0.9.0 from PyPI.

Note

High Risk
Breaking schema bump (v3 ↔ v4 hard mismatch) plus changed default requested_columns semantics; server and SDK must deploy together or forecasts and column alignment break silently for callers that assumed v3 behavior.

Overview
Bumps the JointFM API envelope to schema_version="v4" and adds first-class log_prob scoring in the Python SDK, with docs, samples, and tests aligned to the new contract.

log_prob: New JointFMClient.forecast_log_prob(...) (plus query_rows on forecast / dataframe builders) sends observed rows and returns LogProbResult / LogProbScores. Request validation ties return_mode="log_prob" to query_rows, enforces one row per query_times, and requires requested_columns to list every declared column in order. Responses parse outputs.log_prob and re-check summary fields against values (DERIVED_SCORE_TOLERANCE). Works with ConditionBlock; equality pins may appear in projections like other columns.

v4 projection rules: Default omitted requested_columns is documented as all declared columns in declared order (not targets-only). Equality-pinned columns may be named in requested_columns (client no longer rejects them). Config defaults (JOINTFM_SCHEMA_VERSION, config.sample.yaml, fixtures) and compatibility checks now expect v4 only.

Docs / examples: README and docs/api-reference.md describe log_prob, plausibility chaining, and v4 rules; forecast_condition.ipynb adds unconditional vs conditional comparison, draws, quantile bands, scenario ranking, and conditional log-prob scoring.

Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.

" history,\n",
" query_times=SCORED_QUERY_TIMES,\n",
" query_rows=pd.DataFrame([outcome], columns=EXPECTED_COLUMNS),\n",
" requested_columns=READABLE_COLUMNS,\n",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Notebook scores with old projection

Medium Severity

The new scoring cell passes READABLE_COLUMNS, which drops the pinned column. v4 log_prob validation requires every declared column in declared order, so this call raises before it reaches the service. The same leftover readable-column rule is still in the forecast_log_prob docstring.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6d78b66. Configure here.

@shackmann
shackmann merged commit 99987f8 into main Sep 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant