fix: correct tensor memory and architecture estimates - #52
rupayon123 wants to merge 12 commits into
Conversation
How to test: PYTHONPATH=src pytest tests/test_scalar_tensors.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_explicit_head_dimension.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_gguf_native_names.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_gpt2_projection_layout.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_primitive_dtype_sizes.py; PYTHONPATH=src pytest
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Not up to standards ⛔🔴 Issues
|
| Category | Results |
|---|---|
| BestPractice | 1 medium |
🟢 Metrics 18 complexity · 0 duplication
Metric Results Complexity 18 Duplication 0
NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.
|
Two additional memory-estimation corrections:
CI for the latest head is still being checked. AI-assisted changes reviewed and verified locally before publishing. |
|
Additional fixes pushed to this branch:
All 108 local tests pass. TQ1_0, TQ2_0 and MXFP4 identifiers/block sizes were checked against official GGML definitions. Invalid tensor dimensions now raise ValueError rather than producing invalid footprints or incidental exceptions. Prepared with AI assistance. Changes are submitted for review; this is not a claim of maintainer approval. |
|
September 16 verified update: Nested multimodal text configuration, lazy dtype estimates and the documented GGUF attention-head default now have regression coverage. The final local suite has 114 passing tests. Lazy dtype inference remains an estimate for mixed or quantized weights. Commits: Exact current head: AI-assisted with OpenAI Codex; changes and validation were inspected before submission. |
Summary
Correct five independently reproduced errors in tensor memory and KV-cache estimates: scalar tensors were skipped, BOOL/U8 buffers were charged two bytes per element, an explicit head_dim was ignored, native GGUF blk/attn_k/attn_qkv names were not recognized, and GPT-2 Conv1D projections used the wrong output axis.
The changes preserve missing-shape handling, zero-length tensors, normal linear QKV layouts, and the existing config fallback when head_dim is absent. Each fix includes a focused regression that failed before its change.
Motivation and references
Type of change
Validation
Python 3.11, PYTHONPATH=src pytest -q: 76 passed. All five new test files pass ruff; git diff --check passes. No model weights were downloaded. These are parser/calculator fixture checks, not actual GPU inference or a claim that all architectures are supported.
Each conventional commit includes How to test instructions. Prepared with AI assistance.