Skip to content

[bitnet] Maturity gate: golden-token parity, model-gated smoke, entry-point parity + CLI knobs #360

Description

@michalharakal

Problem

#346's maturity gate stands at 0/3 for BitNet:

  1. No llama.cpp golden-token parity test.
  2. No tests/smoke/smoke-models.json row — PR SKaiNET-transformers → engine 0.51: full engine-loader migration, MAPPED default, BitNet fold-in #353's "3/3 pass … BitNet-2B4T" run was manual, nothing gates regressions.
  3. Entry-point parity 1/3 — only the unified skainet-cli; no chat-model facade, no family CLI (the facade arrives with [bitnet] Full #346 conformance: BitNetWeightLoader/RuntimeWeights naming + llm-runtime/kbitnet facade #359's llm-runtime/kbitnet).

Plus CLI rough edges on the BitNet path:

  • i2sLayout hardcoded to GROUP_128 at the call site — the I2sGgufLayout knob (GROUP_64 / SEQUENTIAL) is not exposed, and a wrong flavor only "fails fast on code 3 where possible".
  • --context silently ignored (Main.kt:246-248).
  • The explicit NativeTernaryF32GemvKernel.install / NativeTernaryLmheadKernel.install at Main.kt:249-250 are the only kernel bootstraps that survived the 0.52 sweep — load-bearing (the ternary packs are not in the engine's self-healing SPI, Ternary kernel packs are missing from the self-healing dispatch SPI — consumers silently get the int8 fallback SKaiNET#1240) but uncommented, so they look like a missed cleanup.

Tasks

Verification

  • Gate: parity + smoke jobs green in CI with the model present, skipped-not-failed without it.
  • --i2s-layout sequential loads a NeoGPU-converted GGUF; wrong-layout run fails with a clear message.

Refs: #346 (maturity gate), #335, #359, SKaiNET-developers/SKaiNET#1240.

Activity

  1. michalharakal commented on Aug 31, 2026

    @michalharakal
    ContributorAuthor

    Status after PR #364: parity probe, smoke row, and CLI knobs are in. Remaining here:

    • Entry-point parity — the llm-runtime/kbitnet facade landed in refactor(bitnet): full #346 conformance — BitNetWeightLoader, RuntimeWeights, kbitnet facade (#359) #363; a chat-facade/family-CLI entry point is still open.
    • Drop the explicit ternary kernel installs in Main.kt once SKaiNET#1240 (PR SKaiNET#1241) ships in a consumed BOM.
    • BF16 arbitration of the greedy divergence found while building the oracle — bitnet.cpp (pinned working build: microsoft/BitNet @ de371b7, fork @ 5eb47b721; current HEAD does not compile) and SKaiNET genuinely disagree from generated token 4 on the standard prompt. The oracle quantizes activations to int8 per matmul; SKaiNET's LUT path is exact FP32×ternary — so the divergence is expected in some direction, but which implementation tracks the BF16 reference (microsoft/bitnet-b1.58-2B-4T safetensors) is unverified. Proposal: an eval-callback-style probe against HF transformers BF16 (the gemma 8a7f1f1 method), gated like the parity test. Details and the oracle's top-5 at the divergence position are in the committed fixture header.
    • README family table + BitNet docs page (still pending, with 0.49.0 adoption E: docs sweep + load diagnostics (explainPlacements, TraceSession.identify) #344's guide rewrite).
  2. added a commit that references this issue on Aug 31, 2026
  3. michalharakal commented on Aug 31, 2026

    @michalharakal
    ContributorAuthor

    E5 (BF16 arbitration) is resolved — and it found a real bug. The HF BF16 reference sided with bitnet.cpp: SKaiNET was the diverging implementation, and the culprit was the RoPE pairing (INTERLEAVED instead of the NEOX-style SPLIT_HALF bitnet.cpp assigns all three BITNET arches). Fixed in PR #365; with it, SKaiNET, bitnet.cpp, and the BF16 reference agree token-for-token on all 32 greedy tokens of the parity prompt, and the committed fixture now gates on that full equality. Details, culprit-isolation steps, and the one documented dense-head near-tie are in the PR and fixture header.

    Remaining on this issue: entry-point parity (E3) and dropping the CLI's explicit kernel installs once a BOM containing SKaiNET#1241 is consumed.

  4. michalharakal commented on Sep 1, 2026

    @michalharakal
    ContributorAuthor

    E3 done in PR #371: a real BitNet chat provider (BitNetChatTemplate per the HF checkpoint's authoritative format — the GGUF's embedded template is a conversion artifact — plus stopTokenStrings() so the loop stops at <|eot_id|> despite the GGUF declaring 128001 as EOS). Verified end-to-end on 2B4T: clean auto-detected chat, correct answer, proper stop.

    E4's last piece (dropping the CLI's explicit kernel installs) is in PR #370 (the 0.52.1 BOM bump — merge after the engine release publishes). With #370 + #371 merged, everything on this issue is done and it can close.

  5. added a commit that references this issue on Sep 1, 2026
  6. michalharakal commented on Sep 2, 2026

    @michalharakal
    ContributorAuthor

    Closing — every item is in:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions