Repository navigation
Common structure and naming convention for model family modules #346
Description
Activity
Convention update: what the Gemma 4 GGUF failure taught us (case study + new rules)
The Gemma 4 E2B Q4_K_M failure (degenerate
<turn|>output at ~0.04 tok/s, while Llama 3.2 ran fine through the identical runtime) turned out to be a stack of five defects — every single one an instance of drift this issue already names. Recording it here as the motivating case study, plus the convention rows it forces.Case study: each drift item had a measurable cost
Convention rule (this issue) The Gemma 4 violation Cost Weight materialization delegated to the engine, no per-family quant code Gemma4WeightLoaderforcedtoken_embd→GEMMA_DEQUANTIZE_ALL(dense FP32, ~2.4 GB) while llama's shared loader keeps it packed and rewraps viaPackedRowDequantTensorData— with tied lm_head, the vocab-sized matmul ran dense off the packed kernel chain every stepdominant perf hit + heap blowup Shared decoderMetadataFromGgufparsergemma kept a private duplicated parser that never read tokenizer.ggml.bos_token_id/eos_token_id/layer_norm_rms_epsilon— BOS defaulted to 1, which is Gemma's<eos>(Gemma BOS = 2), so every generation was prefilled with an end-of-sequence tokengarbage logits from step 0 Family template completeness kgemma CLI never installed the kernel packs ( KernelPacks.install()+FfmRowMajorKernelPack.install()) — kllama got exactly this fix in #354, kgemma was skipped~1000× slowdown via the silent decoding-reference-kernel fallback — (new rule needed) stop tokens: Gemma 4 ships three stop ids ( [1, 106, 50]ingeneration_config.json); the GGUF path stopped on{1}only, and the CLI plain mode had no stop check at allthe literal `<turn — (new rule needed) engine registries: TokenizerFactoryallowlist andModelArchitecture.ggufIdMaphad nogemma4, so the unifiedskainet-cliroute threwUnsupportedTokenizerExceptionbefore generating a single tokenwhole entry point dead Fixes are on
feat/0.51-migration(transformers) + engine branchfix/gemma4-tokenizer-and-registry(tokenizer registry,SpecialTokenSplitter.decodeTokenstreaming-space fix,toIntFlexiblefortoken_typeUInt arrays, observableKernelDispatchfallback with a one-time loud warning). Also relevant to #325 (parity gate verifies) and #330 (the two-tokenizer disagreement — three of the engine fixes close real divergences).New convention rows (proposed additions to the table above)
Concern Rule Kernel bootstrap Every runtime facade and CLI MUST install the kernel packs before the first forward. Target shape: one shared RuntimeBootstrap.ensureInstalled()inllm-core— per-runtime copies (KLlamaJava's, kgemma's newKgemmaKernels) are transitional and should be hoistedToken ids bos/eos/pad MUST be parsed from GGUF metadata and threaded into the runtime — never rely on bosToken = 1-style defaults (that literally means<eos>on Gemma)Stop tokens Stop handling takes a Set of ids (llm-core generateUntilStop(eosTokenIds: Set<Int>)now exists); family modules provide a<F>StopTokens.resolve(tokenizer)that resolves from the vocab with documented fallbacksEntry-point parity A family is not "done" until all three entry points work: chat-model facade, family CLI, and the unified skainet-clirouteEngine registries Landing a family includes adding its general.architectureid to the engine'sTokenizerFactoryallowlist andModelArchitecture.ggufIdMap— checklist item, not an afterthoughtMaturity gate (proposed; from the architecture review)
A family's README/status row may claim "verified" ONLY with, at minimum:
- Golden-token parity vs a llama.cpp reference (greedy N tokens from pinned ids —
Gemma4ReferenceParityDiagnosticis the pattern), - a model-gated smoke test with a time bound (e.g. 16 tokens < 30 s on the reference checkpoint),
- tool-calling e2e if tool calling is claimed,
- compiled-path export green if a compiled path is claimed.
"Code exists" and "verified" are different columns. Proposal: CI-gate the verified column as a follow-up issue, and until then keep the experimental banner prominent.
Freeze + consolidation direction (proposed)
- Freeze new
llm-inference/<family>modules until the shared decoder core (commonTransformerDecoder+ config-driven weight map + oneOptimizedLLMRuntimesurface) is solid; new models enter as thin adapters over this issue's template. - Consolidation order suggested by the Gemma 4 evidence: (1) metadata parser unification (the private-parser bug class), (2) stop-token resolution + chat-template registry, (3) kernel bootstrap hoist, (4)
skainet-clias the one consumer-facing CLI with family CLIs demoted to dev tools. - Engine↔transformers contract: propose a pin-bump regression suite (quant formats × packed/dequant × eager/traced) as its own issue — the Gemma 4 kernel-pack gap shipped precisely because nothing tripped on it.
This comment turns the Gemma 4 debugging session into convention text; deviations should be justified in PRs against it, per this issue's charter.
- Golden-token parity vs a llama.cpp reference (greedy N tokens from pinned ids —
BitNet — this convention's first greenfield instantiation (plan item 3) — now conforms fully via PR #363:
BitNetWeightLoader(renamed),BitNetRuntimeWeights.ktwithBitNetTensorNames, and thellm-runtime/kbitnetBitNetIngestionfacade. Two deviations documented per this issue's rule (in the PR body and class KDoc): the runtime-weights container is the name→tensor map the DSL runtime actually consumes rather than Apertus-style typed per-layer fields (BitNet has no hand-rolled loop), and the ConfigParser row is satisfied by the shareddecoderMetadataFromGguf. The maturity-gate items for BitNet remain tracked in #360.Per-family conformance close-out — the llama+gemma arc is complete (2026-09-01)
The arc this issue chartered is done: #372/#373/#374/#375/#376 all delivered and closed (PRs #378, #379, #381, #382, #383, #384, #385, on the engine 0.52.0 base). Updated table:
Family Template conformance Maturity gate Apertus ✅ reference shape (unchanged) ⚠️ no parity probe / smoke row yet — the gate predates it; retrofit is the natural next candidateBitNet ✅ full (the greenfield instantiation) ✅ 3/3 — three-way token-equality parity (bitnet.cpp + HF BF16), smoke row, all entry points incl. chat Llama ✅ — the shared decoder machinery now lives in llm-core/sk.ainet.lang.nn.dsl.decoder(#378:GgufDecoderMetadataet al., the family no longer lends its name to shared types),LlamaWeightLoaderfamily row added (#379)✅ — LlamaGoldenTokenParityTestasserts full cross-implementation greedy text equality vs mainline llama.cpp b10621 (the strongest gate in the repo); the smoke rows certify the shipping DSL path since kllama-cli's GGUF leg retired the deprecatedLlamaRuntime(#382, closing #354's GGUF half — measured: DSL 32 tok/s vs legacy DNF-in-10-min)Gemma (3 + 4) ✅ — unversioned names with one-release shims, ~535 LOC dead runtime deleted, KgemmaKernelson the self-healing dispatch, androidNativeArm32 repaired (#381); dtype policy truthful and actually threaded, one metadata parser instead of hand-synced twins (#383); the SafeTensors materialization-policy remainder is engine-gap SKaiNET#1246✅ — golden-token + timed-smoke tests CI-wired into the smoke-reference tier with a gemma4_gguf_urlstaging input, E2B smoke row added, gemma3 checkpoints route to the DSL lane instead of the 3n runtime, skainet-cli refuses gemma3n/gemma2 loudly (#385); #325 closed with the gate's evidenceQwen ⏸ postponed to the next release #352 — plus a concrete lead from the resolver dedup: the deleted llm-core fork carried qwen2 attn q/k/v/o bias mapping rules the engine resolver lacks Gemma 3n ⏸ postponed — #377 (module split + packed loading + gate) 0/5; deliberately untouched by the sweep Also hardened along the way: the gemma DSL quant-parity tests pin the reference kernels (#384) — the 0.52.0 self-healing dispatch had started serving platform int8-path Q4_K kernels inside un-bootstrapped test JVMs (SKaiNET#944's numerics), which is worth knowing for any future tight-tolerance parity test.
This issue stays open as the living convention reference; the durable how-to is
docs/how-to/add-model.adoc(#368). Tracker with the full evidence trail:LLAMA-GEMMA-CONFORMANCE-TRACKER.mdin the maintainer's working notes.- added a commit that references this issue
on Sep 1, 2026 Qwen column update — the postponed family is now fully conformant (2026-09-01)
The "postponed to the next release" row above is resolved. The qwen lane delivered in two PRs on the 0.52.0 base:
Family Template conformance Maturity gate Qwen ✅ — QwenWeightLoaderfamily row (thin wrapper over thellm-coredecoder loader, publicQWEN_ARCHITECTURES),QwenGgufWeightSourcenaming drift resolved (QwenTensorNames, one-release shim),QwenGGUFNameResolvercarries the family's bias rules (#389)✅ — #352's root cause fixed in #387: the attention projection biases ( blk.N.attn_{q,k,v}.bias, enormous — blk.0's K bias moves the channel sum11→507) were dropped by the decoder loader's.weight-only wanted-set and unmapped by the resolver; both fixed, plus a loud failure if a file bias ever goes unbound again.QwenGoldenTokenParityTestasserts full 32-step greedy text equality vs mainline llama.cpp b10621 for both variants — Qwen2.5-0.5B Q8_0 (biases, no QK-norm) and Qwen3-1.7B Q8_0 (QK-norm, no biases) — CI-wired into the smoke-reference tier (qwen25_gguf_urlstaging input; the qwen3 download feeds both its smoke test and the gate). Smoke lane gains a Qwen2.5-0.5B chat row certifying the fixed path end-to-end (kllama, 9.6 tok/s locally)Deliberate non-rows, recorded in #389: no family-typed
QwenRuntimeWeightscontainer (the DSL path consumesDecoderGgufWeightsdirectly — a wrapper would be artificial), and the runtime facade remainskllama(it serves the whole llama-compatible architecture set; a separatekqwenwould duplicate it).Worth knowing for future gates: the qwen3 fixture prompt is chosen for greedy decisiveness (min top-1/top-2 gap 1.66 nats across 32 steps). "The capital of France is" hits a measured 0.06-nat three-way tie at step 6 on Qwen3-1.7B that Q8_0 cross-implementation noise legitimately flips — full-text-equality fixtures should check their oracle's per-step margins (
n_probs) before pinning a prompt.Remaining family work is tracked elsewhere: #118 (HF SafeTensors validation), gemma3n #377, Apertus gate retrofit.
Apertus row update — the gate retrofit is done, and it caught the family broken (2026-09-02)
The "
⚠️ no parity probe / smoke row yet" cell above is resolved by PR #390 — and the retrofit immediately justified itself: Apertus decode was broken in production (real Apertus-8B-Instruct GGUFs decoded to<unk>noise on every entry point). Two defects, both found and pinned by the new gate:XIELUActivationusedexp()as a "simplified softplus approx" — the real model's per-layeralpha_pvalues reach 174,exp(174)isInf, and the branch-mask multiply turned every logitNaN. Now exact guarded softplus, computed host-side on the frozen per-layer scalars.apertusNetwork()built RoPE on the DSL defaults (INTERLEAVED, base 10 000) — Apertus is an HF rotate-half model llama.cpp runs as NEOX withrope_theta = 12M; the metadata carriedropeThetabut the builder never passed it.
Family Template conformance Maturity gate Apertus ✅ reference shape (unchanged) ✅ — ApertusGoldenTokenParityTestasserts full 32-step greedy text equality vs mainline llama.cpp b10621 on Apertus-8B-Instruct-2509 Q4_K_S, on the DSL path the CLI ships (QK-norm + per-layer xIELU + ungated FFN end-to-end). CI-wired into the smoke-reference tier (apertus_gguf_urlstaging input + 12g test-heap arg);smoke-models.jsongains an Apertus row on the skainet-cli runnerWith this, every shipped generative family except Gemma 3n (#377) has a reference-implementation parity gate. The pattern held for the third time running (BitNet RoPE pairing, Qwen2 bias loading, now Apertus xIELU+RoPE): each family "declared end-to-end" without a golden-token gate turned out to have a silent structural defect the gate found within hours. Docs now say only what the gates prove: README + antora index carry a per-family verified-against matrix, and
reference/architecture.adocis rewritten around the DSL-centric decoder core and this issue's template.Closing the tracking issue: the arc it chartered is complete — llama, gemma (#372–#376), qwen (#352 lane), apertus (#390) and BitNet (#363) all conform to the template and carry a reference-implementation parity gate; the SafeTensors half of "no per-family quant code" landed with 0.53.0 (Gemma #398, shared decoder #400, Apertus + Gemma 3n #401 on the engine's
ShardedSafeTensorsParametersLoader, SKaiNET#1246). The convention itself — the table in the description plus the Gemma 4 case-study rules in the comments — stays the reference; link here from CONTRIBUTING/family docs rather than reopening. The one family without a full maturity gate is Gemma 3n, tracked in #377; Voxtral's SafeTensors collapse waits on SKaiNET#1256.
Motivation
Each model family module grew its own structure organically. With the 0.49.0 memory-model adoption (#338–#344) rewriting every family's loading path, now is the moment to fix one common scheme so every family looks the same and adding a new one is mechanical.
De-facto convention (Apertus/Voxtral/Gemma follow it most closely)
For a family
<F>(modulellm-inference/<f>):<f>Network()builder fn<F>NetworkDef.kt<F>NetworkLoader(fromGguf/fromSafeTensors/fromWeights)<F>NetworkLoader.kt<F>WeightLoader(fromSource/fromRandomAccess,loadToMap)<F>WeightLoader.kt<F>Weights,<F>RuntimeWeights,<F>TensorNames<F>RuntimeWeights.kt<F>ConfigParser<F>ConfigParser.ktApertusXIELU.kt,AltUp.kt); shared ones live intransformer-core/llm-core<F>Ingestioninllm-runtime/k<f>Weight materialization is delegated to the engine (
StreamingGgufParametersLoader/WeightFormfromsk.ainet.core:skainet-io-gguf); no per-family quant/packing/converter code (*QuantLayout,*PackedWeights,*MemSegConverter,QuantizedTensorFactoryare legacy and retire with the 0.49.0 migration).Known drift to converge
DecoderGgufWeightLoaderinstead of family prefix;QwenGgufWeightSourcenaming; three different tensor-name patterns (LlamaGgufTensorNamesvsGemma3nGgufTensorNamesvsApertusTensorNamesvsVoxtralTensorNames).Gemma4WeightLoader,Gemma3nWeightLoader,GemmaNetworkLoader).QuantizedTensorFactory/LlamaQuantLayout, gemmaGemmaQuantLayout/BlockQuantPackinguse, apertusApertusMemSegConverter— the latter already removed in 0.49.0 adoption B1: Apertus onto the engine loader (fixes #100) #339).Plan
This issue is the single reference for the convention; deviations should be justified in PRs against it.