Skip to content

Add an explicit native representation for sampled tokenization - #976

Closed
bradhilton wants to merge 2 commits into
hayek/explicit-sampled-stop-api-20260925from
hayek/native-sampled-representation-20260926
Closed

bradhilton wants to merge 2 commits into
hayek/explicit-sampled-stop-api-20260925from
hayek/native-sampled-representation-20260926

Conversation

@bradhilton

@bradhilton bradhilton commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Recorded tool-call completions can make template-based tokenization fail even when the original prompts and output tokens are complete. This adds representation="native" to art.tokenize_sampled(...), letting callers build histories directly from the recorded evidence.

Native mode verifies complete conditioning, token IDs, logprobs, source order and STOP flags. It joins compatible consecutive generations and keeps others separate. representation="rendered" remains the default, and generic art.tokenize is unchanged. This PR depends on #973.

Public regressions and full CI pass. Independent Astra and Fable source reviews found no blockers. One bounded invocation on a retained failure input also passed all 30 sources’ token, logprob, STOP and ownership checks on the current head.

This opt-in representation supports the SAMPLED first-occurrence loss contract. It does not establish numerical-training equivalence or interchangeable behavior for OUTPUT/SFT losses or arbitrary per-history weighting.

Validation details

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant