Skip to content

Common structure and naming convention for model family modules #346

Description

@michalharakal

Motivation

Each model family module grew its own structure organically. With the 0.49.0 memory-model adoption (#338–#344) rewriting every family's loading path, now is the moment to fix one common scheme so every family looks the same and adding a new one is mechanical.

De-facto convention (Apertus/Voxtral/Gemma follow it most closely)

For a family <F> (module llm-inference/<f>):

Concern Name File
DSL network definition <f>Network() builder fn <F>NetworkDef.kt
End-to-end module loader <F>NetworkLoader (fromGguf / fromSafeTensors / fromWeights) <F>NetworkLoader.kt
GGUF weight materialization <F>WeightLoader (fromSource / fromRandomAccess, loadToMap) <F>WeightLoader.kt
Weight containers + tensor names <F>Weights, <F>RuntimeWeights, <F>TensorNames <F>RuntimeWeights.kt
HF config parsing <F>ConfigParser <F>ConfigParser.kt
Family-specific layers/ops bare descriptive names in the family module (ApertusXIELU.kt, AltUp.kt); shared ones live in transformer-core / llm-core —
Runtime facade <F>Ingestion in llm-runtime/k<f> —

Weight materialization is delegated to the engine (StreamingGgufParametersLoader / WeightForm from sk.ainet.core:skainet-io-gguf); no per-family quant/packing/converter code (*QuantLayout, *PackedWeights, *MemSegConverter, QuantizedTensorFactory are legacy and retire with the 0.49.0 migration).

Known drift to converge

  • llama/qwen/smollm2: shared DecoderGgufWeightLoader instead of family prefix; QwenGgufWeightSource naming; three different tensor-name patterns (LlamaGgufTensorNames vs Gemma3nGgufTensorNames vs ApertusTensorNames vs VoxtralTensorNames).
  • gemma: versioned + unversioned prefixes mixed in one module (Gemma4WeightLoader, Gemma3nWeightLoader, GemmaNetworkLoader).
  • Per-family quant machinery (llama QuantizedTensorFactory/LlamaQuantLayout, gemma GemmaQuantLayout/BlockQuantPacking use, apertus ApertusMemSegConverter — the latter already removed in 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339).

Plan

  1. 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339 (Apertus) establishes the reference shape — done as part of the B1 migration.
  2. 0.49.0 adoption B2: llama/qwen/smollm2 onto the engine loader (retire QuantizedTensorFactory) #340/0.49.0 adoption B3: gemma4/gemma3n onto the engine loader (retire BlockQuantPacking + capability gate) #341 align llama/qwen/smollm2 and gemma to this skeleton while migrating them onto the engine loader (rename where cheap; note deliberate exceptions here when a rename is not worth the churn).
  3. The BitNet module ([bitnet] T1: BitNet architecture — ModelFamily, network def, packed-weight loader #336) is the first greenfield instantiation of the template.
  4. 0.49.0 adoption E: docs sweep + load diagnostics (explainPlacements, TraceSession.identify) #344 documents the resulting scheme as the "how to add a model family" guide.

This issue is the single reference for the convention; deviations should be justified in PRs against it.

Activity

  1. michalharakal commented on Aug 30, 2026

    @michalharakal
    ContributorAuthor

    Convention update: what the Gemma 4 GGUF failure taught us (case study + new rules)

    The Gemma 4 E2B Q4_K_M failure (degenerate <turn|> output at ~0.04 tok/s, while Llama 3.2 ran fine through the identical runtime) turned out to be a stack of five defects — every single one an instance of drift this issue already names. Recording it here as the motivating case study, plus the convention rows it forces.

    Case study: each drift item had a measurable cost

    Convention rule (this issue) The Gemma 4 violation Cost
    Weight materialization delegated to the engine, no per-family quant code Gemma4WeightLoader forced token_embd → GEMMA_DEQUANTIZE_ALL (dense FP32, ~2.4 GB) while llama's shared loader keeps it packed and rewraps via PackedRowDequantTensorData — with tied lm_head, the vocab-sized matmul ran dense off the packed kernel chain every step dominant perf hit + heap blowup
    Shared decoderMetadataFromGguf parser gemma kept a private duplicated parser that never read tokenizer.ggml.bos_token_id / eos_token_id / layer_norm_rms_epsilon — BOS defaulted to 1, which is Gemma's <eos> (Gemma BOS = 2), so every generation was prefilled with an end-of-sequence token garbage logits from step 0
    Family template completeness kgemma CLI never installed the kernel packs (KernelPacks.install() + FfmRowMajorKernelPack.install()) — kllama got exactly this fix in #354, kgemma was skipped ~1000× slowdown via the silent decoding-reference-kernel fallback
    — (new rule needed) stop tokens: Gemma 4 ships three stop ids ([1, 106, 50] in generation_config.json); the GGUF path stopped on {1} only, and the CLI plain mode had no stop check at all the literal `<turn
    — (new rule needed) engine registries: TokenizerFactory allowlist and ModelArchitecture.ggufIdMap had no gemma4, so the unified skainet-cli route threw UnsupportedTokenizerException before generating a single token whole entry point dead

    Fixes are on feat/0.51-migration (transformers) + engine branch fix/gemma4-tokenizer-and-registry (tokenizer registry, SpecialTokenSplitter.decodeToken streaming-space fix, toIntFlexible for token_type UInt arrays, observable KernelDispatch fallback with a one-time loud warning). Also relevant to #325 (parity gate verifies) and #330 (the two-tokenizer disagreement — three of the engine fixes close real divergences).

    New convention rows (proposed additions to the table above)

    Concern Rule
    Kernel bootstrap Every runtime facade and CLI MUST install the kernel packs before the first forward. Target shape: one shared RuntimeBootstrap.ensureInstalled() in llm-core — per-runtime copies (KLlamaJava's, kgemma's new KgemmaKernels) are transitional and should be hoisted
    Token ids bos/eos/pad MUST be parsed from GGUF metadata and threaded into the runtime — never rely on bosToken = 1-style defaults (that literally means <eos> on Gemma)
    Stop tokens Stop handling takes a Set of ids (llm-core generateUntilStop(eosTokenIds: Set<Int>) now exists); family modules provide a <F>StopTokens.resolve(tokenizer) that resolves from the vocab with documented fallbacks
    Entry-point parity A family is not "done" until all three entry points work: chat-model facade, family CLI, and the unified skainet-cli route
    Engine registries Landing a family includes adding its general.architecture id to the engine's TokenizerFactory allowlist and ModelArchitecture.ggufIdMap — checklist item, not an afterthought

    Maturity gate (proposed; from the architecture review)

    A family's README/status row may claim "verified" ONLY with, at minimum:

    1. Golden-token parity vs a llama.cpp reference (greedy N tokens from pinned ids — Gemma4ReferenceParityDiagnostic is the pattern),
    2. a model-gated smoke test with a time bound (e.g. 16 tokens < 30 s on the reference checkpoint),
    3. tool-calling e2e if tool calling is claimed,
    4. compiled-path export green if a compiled path is claimed.

    "Code exists" and "verified" are different columns. Proposal: CI-gate the verified column as a follow-up issue, and until then keep the experimental banner prominent.

    Freeze + consolidation direction (proposed)

    • Freeze new llm-inference/<family> modules until the shared decoder core (common TransformerDecoder + config-driven weight map + one OptimizedLLMRuntime surface) is solid; new models enter as thin adapters over this issue's template.
    • Consolidation order suggested by the Gemma 4 evidence: (1) metadata parser unification (the private-parser bug class), (2) stop-token resolution + chat-template registry, (3) kernel bootstrap hoist, (4) skainet-cli as the one consumer-facing CLI with family CLIs demoted to dev tools.
    • Engine↔transformers contract: propose a pin-bump regression suite (quant formats × packed/dequant × eager/traced) as its own issue — the Gemma 4 kernel-pack gap shipped precisely because nothing tripped on it.

    This comment turns the Gemma 4 debugging session into convention text; deviations should be justified in PRs against it, per this issue's charter.

  2. michalharakal commented on Aug 31, 2026

    @michalharakal
    ContributorAuthor

    BitNet — this convention's first greenfield instantiation (plan item 3) — now conforms fully via PR #363: BitNetWeightLoader (renamed), BitNetRuntimeWeights.kt with BitNetTensorNames, and the llm-runtime/kbitnet BitNetIngestion facade. Two deviations documented per this issue's rule (in the PR body and class KDoc): the runtime-weights container is the name→tensor map the DSL runtime actually consumes rather than Apertus-style typed per-layer fields (BitNet has no hand-rolled loop), and the ConfigParser row is satisfied by the shared decoderMetadataFromGguf. The maturity-gate items for BitNet remain tracked in #360.

  3. added 2 commits that reference this issue on Aug 31, 2026
  4. michalharakal commented on Sep 1, 2026

    @michalharakal
    ContributorAuthor

    Per-family conformance close-out — the llama+gemma arc is complete (2026-09-01)

    The arc this issue chartered is done: #372/#373/#374/#375/#376 all delivered and closed (PRs #378, #379, #381, #382, #383, #384, #385, on the engine 0.52.0 base). Updated table:

    Family Template conformance Maturity gate
    Apertus ✅ reference shape (unchanged) ⚠️ no parity probe / smoke row yet — the gate predates it; retrofit is the natural next candidate
    BitNet ✅ full (the greenfield instantiation) ✅ 3/3 — three-way token-equality parity (bitnet.cpp + HF BF16), smoke row, all entry points incl. chat
    Llama ✅ — the shared decoder machinery now lives in llm-core/sk.ainet.lang.nn.dsl.decoder (#378: GgufDecoderMetadata et al., the family no longer lends its name to shared types), LlamaWeightLoader family row added (#379) ✅ — LlamaGoldenTokenParityTest asserts full cross-implementation greedy text equality vs mainline llama.cpp b10621 (the strongest gate in the repo); the smoke rows certify the shipping DSL path since kllama-cli's GGUF leg retired the deprecated LlamaRuntime (#382, closing #354's GGUF half — measured: DSL 32 tok/s vs legacy DNF-in-10-min)
    Gemma (3 + 4) ✅ — unversioned names with one-release shims, ~535 LOC dead runtime deleted, KgemmaKernels on the self-healing dispatch, androidNativeArm32 repaired (#381); dtype policy truthful and actually threaded, one metadata parser instead of hand-synced twins (#383); the SafeTensors materialization-policy remainder is engine-gap SKaiNET#1246 ✅ — golden-token + timed-smoke tests CI-wired into the smoke-reference tier with a gemma4_gguf_url staging input, E2B smoke row added, gemma3 checkpoints route to the DSL lane instead of the 3n runtime, skainet-cli refuses gemma3n/gemma2 loudly (#385); #325 closed with the gate's evidence
    Qwen ⏸ postponed to the next release #352 — plus a concrete lead from the resolver dedup: the deleted llm-core fork carried qwen2 attn q/k/v/o bias mapping rules the engine resolver lacks
    Gemma 3n ⏸ postponed — #377 (module split + packed loading + gate) 0/5; deliberately untouched by the sweep

    Also hardened along the way: the gemma DSL quant-parity tests pin the reference kernels (#384) — the 0.52.0 self-healing dispatch had started serving platform int8-path Q4_K kernels inside un-bootstrapped test JVMs (SKaiNET#944's numerics), which is worth knowing for any future tight-tolerance parity test.

    This issue stays open as the living convention reference; the durable how-to is docs/how-to/add-model.adoc (#368). Tracker with the full evidence trail: LLAMA-GEMMA-CONFORMANCE-TRACKER.md in the maintainer's working notes.

  5. added a commit that references this issue on Sep 1, 2026
  6. michalharakal commented on Sep 1, 2026

    @michalharakal
    ContributorAuthor

    Qwen column update — the postponed family is now fully conformant (2026-09-01)

    The "postponed to the next release" row above is resolved. The qwen lane delivered in two PRs on the 0.52.0 base:

    Family Template conformance Maturity gate
    Qwen ✅ — QwenWeightLoader family row (thin wrapper over the llm-core decoder loader, public QWEN_ARCHITECTURES), QwenGgufWeightSource naming drift resolved (QwenTensorNames, one-release shim), QwenGGUFNameResolver carries the family's bias rules (#389) ✅ — #352's root cause fixed in #387: the attention projection biases (blk.N.attn_{q,k,v}.bias, enormous — blk.0's K bias moves the channel sum 11→507) were dropped by the decoder loader's .weight-only wanted-set and unmapped by the resolver; both fixed, plus a loud failure if a file bias ever goes unbound again. QwenGoldenTokenParityTest asserts full 32-step greedy text equality vs mainline llama.cpp b10621 for both variants — Qwen2.5-0.5B Q8_0 (biases, no QK-norm) and Qwen3-1.7B Q8_0 (QK-norm, no biases) — CI-wired into the smoke-reference tier (qwen25_gguf_url staging input; the qwen3 download feeds both its smoke test and the gate). Smoke lane gains a Qwen2.5-0.5B chat row certifying the fixed path end-to-end (kllama, 9.6 tok/s locally)

    Deliberate non-rows, recorded in #389: no family-typed QwenRuntimeWeights container (the DSL path consumes DecoderGgufWeights directly — a wrapper would be artificial), and the runtime facade remains kllama (it serves the whole llama-compatible architecture set; a separate kqwen would duplicate it).

    Worth knowing for future gates: the qwen3 fixture prompt is chosen for greedy decisiveness (min top-1/top-2 gap 1.66 nats across 32 steps). "The capital of France is" hits a measured 0.06-nat three-way tie at step 6 on Qwen3-1.7B that Q8_0 cross-implementation noise legitimately flips — full-text-equality fixtures should check their oracle's per-step margins (n_probs) before pinning a prompt.

    Remaining family work is tracked elsewhere: #118 (HF SafeTensors validation), gemma3n #377, Apertus gate retrofit.

  7. michalharakal commented on Sep 1, 2026

    @michalharakal
    ContributorAuthor

    Apertus row update — the gate retrofit is done, and it caught the family broken (2026-09-02)

    The "⚠️ no parity probe / smoke row yet" cell above is resolved by PR #390 — and the retrofit immediately justified itself: Apertus decode was broken in production (real Apertus-8B-Instruct GGUFs decoded to <unk> noise on every entry point). Two defects, both found and pinned by the new gate:

    1. XIELUActivation used exp() as a "simplified softplus approx" — the real model's per-layer alpha_p values reach 174, exp(174) is Inf, and the branch-mask multiply turned every logit NaN. Now exact guarded softplus, computed host-side on the frozen per-layer scalars.
    2. apertusNetwork() built RoPE on the DSL defaults (INTERLEAVED, base 10 000) — Apertus is an HF rotate-half model llama.cpp runs as NEOX with rope_theta = 12M; the metadata carried ropeTheta but the builder never passed it.
    Family Template conformance Maturity gate
    Apertus ✅ reference shape (unchanged) ✅ — ApertusGoldenTokenParityTest asserts full 32-step greedy text equality vs mainline llama.cpp b10621 on Apertus-8B-Instruct-2509 Q4_K_S, on the DSL path the CLI ships (QK-norm + per-layer xIELU + ungated FFN end-to-end). CI-wired into the smoke-reference tier (apertus_gguf_url staging input + 12g test-heap arg); smoke-models.json gains an Apertus row on the skainet-cli runner

    With this, every shipped generative family except Gemma 3n (#377) has a reference-implementation parity gate. The pattern held for the third time running (BitNet RoPE pairing, Qwen2 bias loading, now Apertus xIELU+RoPE): each family "declared end-to-end" without a golden-token gate turned out to have a silent structural defect the gate found within hours. Docs now say only what the gates prove: README + antora index carry a per-family verified-against matrix, and reference/architecture.adoc is rewritten around the DSL-centric decoder core and this issue's template.

  8. michalharakal commented on Sep 2, 2026

    @michalharakal
    ContributorAuthor

    Closing the tracking issue: the arc it chartered is complete — llama, gemma (#372–#376), qwen (#352 lane), apertus (#390) and BitNet (#363) all conform to the template and carry a reference-implementation parity gate; the SafeTensors half of "no per-family quant code" landed with 0.53.0 (Gemma #398, shared decoder #400, Apertus + Gemma 3n #401 on the engine's ShardedSafeTensorsParametersLoader, SKaiNET#1246). The convention itself — the table in the description plus the Gemma 4 case-study rules in the comments — stays the reference; link here from CONTRIBUTING/family docs rather than reopening. The one family without a full maturity gate is Gemma 3n, tracked in #377; Voxtral's SafeTensors collapse waits on SKaiNET#1256.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions