Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions docs/features/additional-histories.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,45 @@ trajectory = Trajectory(
)
```

## Migrating captured histories with different token prefixes

Tokenization now rejects a known change to the complete token prefix of a
captured sampled response unless exact-source reconstruction preserves its
conditioning. Earlier versions could carry the response's sampled logprobs into
a different rendered prefix. Matching decoded text is not enough: the token
sequence determines the conditioning. Explicit tokenizer or template overrides
and `reconcile_text_equivalent_tokenizations=True` do not waive this check.

When captured turns have different authoritative prefixes, keep separate
histories where the protocol supports them:

```python
tokenized = trajectory.tokenize(
multi_history=True,
reconcile_text_equivalent_tokenizations=False,
)
```

This preserves separate histories that the protocol already produces, with each
history's original prompt, sampled outputs and logprobs. It does not automatically
split every incompatible captured history. Every resulting history must still
satisfy the conditioning checks; incomplete or inconsistent source evidence is
not made valid. Remove an override that changes sampled conditioning, or supply
matching source evidence.
If you deliberately want to train on edited or approximate text, construct
source-less/manual histories or use the [SFT message format](/fundamentals/sft-training#data-format)
as a separate supervised dataset. Do not carry the old sampled logprobs or source
references into that approximation. This is not an opt-in to recondition
captured RL samples or a new mode that substitutes NaN logprobs.

The `Exact source prefix mismatch` error reports only structural coordinates:
the source's zero-based message index (or `None` if unavailable), assembled and
recorded prefix lengths, and the first differing token offset. If the prefixes
differ only in length, the offset is the shorter length. It contains no token IDs, message
text, request IDs or prefix hashes. Use those coordinates to locate the source
and rendering override in your own data. Missing prompt metadata still follows
the existing fallback; absence of this error is not proof of alignment.

## How It Works

### Tokenization Process
Expand Down
16 changes: 16 additions & 0 deletions src/art/trajectories/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -876,6 +876,22 @@ def tokenize(
chat_template: str | None = None,
chat_template_kwargs: Mapping[str, object] | None = None,
) -> TokenizedTrajectory | TokenizedMultiHistoryTrajectory:
"""Tokenize histories while retaining their sampled-source evidence.

A recorded sampled logprob belongs to its complete original token
prefix. A known different prefix raises ValueError unless exact-source
reconstruction can preserve every sampled source's conditioning.
Explicit template/tokenizer overrides and text-equivalent reconciliation
do not permit carrying sampled logprobs into different conditioning.

For divergent captured prompts, use ``multi_history=True`` with the
default ``reconcile_text_equivalent_tokenizations=False`` to retain
separate authoritative histories. This preserves existing separate
histories; it does not force a split, and each history must still pass
the conditioning checks. This does not recondition samples or discard
their evidence. Missing prompt metadata follows the existing
fallback and is not certified aligned by this known-mismatch check.
"""
from ._tokenize import tokenize_trajectory

return tokenize_trajectory(
Expand Down
Loading
Loading