Skip to content

fix: correct tensor memory and architecture estimates - #52

Open
rupayon123 wants to merge 12 commits into
pipe1os:mainfrom
rupayon123:contribution/tensor-memory-correctness-20260913
Open

rupayon123 wants to merge 12 commits into
pipe1os:mainfrom
rupayon123:contribution/tensor-memory-correctness-20260913

Conversation

@rupayon123

Copy link
Copy Markdown
Contributor

Summary

Correct five independently reproduced errors in tensor memory and KV-cache estimates: scalar tensors were skipped, BOOL/U8 buffers were charged two bytes per element, an explicit head_dim was ignored, native GGUF blk/attn_k/attn_qkv names were not recognized, and GPT-2 Conv1D projections used the wrong output axis.

The changes preserve missing-shape handling, zero-length tensors, normal linear QKV layouts, and the existing config fallback when head_dim is absent. Each fix includes a focused regression that failed before its change.

Motivation and references

Type of change

  • Bug fixes with regression tests

Validation

Python 3.11, PYTHONPATH=src pytest -q: 76 passed. All five new test files pass ruff; git diff --check passes. No model weights were downloaded. These are parser/calculator fixture checks, not actual GPU inference or a claim that all architectures are supported.

Each conventional commit includes How to test instructions. Prepared with AI assistance.

How to test: PYTHONPATH=src pytest tests/test_scalar_tensors.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_explicit_head_dimension.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_gguf_native_names.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_gpt2_projection_layout.py; PYTHONPATH=src pytest
How to test: PYTHONPATH=src pytest tests/test_primitive_dtype_sizes.py; PYTHONPATH=src pytest
@coderabbitai

coderabbitai Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 5ddb46a9-7677-47ab-962c-b941e61068ad


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codacy-production

codacy-production Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Not up to standards ⛔

🔴 Issues 1 medium

Alerts:
⚠ 1 issue (≤ 0 issues of at least minor severity)

Results:
1 new issue

Category Results
BestPractice 1 medium

View in Codacy

🟢 Metrics 18 complexity · 0 duplication

Metric Results
Complexity 18
Duplication 0

View in Codacy

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

@rupayon123

Copy link
Copy Markdown
Contributor Author

Two additional memory-estimation corrections:

  • 7659ead reads GPT-2's serialized n_layer, n_head, and n_embd aliases, preserving precedence for generic config keys. The regression previously returned zero layers/KV dimensions; 78 local tests passed after the fix. Reference: https://github.com/huggingface/transformers/blob/main/src/transformers/models/gpt2/configuration_gpt2.py
  • dd7cc5f corrects 12 quantized storage multipliers. Expected block sizes were independently checked by compiling sizeof(block_...) against GGML's current ggml-common.h: https://github.com/ggml-org/ggml/blob/master/src/ggml-common.h . Of 21 packed-storage regression cases, 12 failed before the fix; the full local suite now passes 99 tests. Two older expectations encoded the incorrect Q4_K and IQ2_XXS values and were corrected to the independently verified sizes. These are packed tensor storage estimates, excluding GGUF file metadata/alignment.

CI for the latest head is still being checked. AI-assisted changes reviewed and verified locally before publishing.

@rupayon123

Copy link
Copy Markdown
Contributor Author

Additional fixes pushed to this branch:

  • bd7e9eb: Recognize TQ1_0, TQ2_0 and MXFP4 GGUF storage.
  • 3157f91: Reject negative, fractional and malformed tensor shapes.

All 108 local tests pass. TQ1_0, TQ2_0 and MXFP4 identifiers/block sizes were checked against official GGML definitions. Invalid tensor dimensions now raise ValueError rather than producing invalid footprints or incidental exceptions.

Prepared with AI assistance. Changes are submitted for review; this is not a claim of maintainer approval.

@rupayon123

Copy link
Copy Markdown
Contributor Author

September 16 verified update:

Nested multimodal text configuration, lazy dtype estimates and the documented GGUF attention-head default now have regression coverage. The final local suite has 114 passing tests. Lazy dtype inference remains an estimate for mixed or quantized weights.

Commits: 769f892b10b47211a8bcb4605bb9383ad1583c41, 2cd130e42134834dc62cb529af873a042ced8f15, b956af9f957232dacf133dae0a58d7c11cc212b3.

Exact current head: b956af9f957232dacf133dae0a58d7c11cc212b3. GitHub checks at this read: Test on ubuntu-latest with Python 3.10: SUCCESS, Test on ubuntu-latest with Python 3.11: SUCCESS, Test on ubuntu-latest with Python 3.12: SUCCESS, Test on macos-latest with Python 3.10: SUCCESS, Test on macos-latest with Python 3.11: SUCCESS, Test on macos-latest with Python 3.12: SUCCESS, Test on windows-latest with Python 3.10: SUCCESS, Test on windows-latest with Python 3.11: SUCCESS, Test on windows-latest with Python 3.12: SUCCESS, Codacy Static Code Analysis: ACTION_REQUIRED, CodeRabbit: SUCCESS. Local results do not imply maintainer approval.

AI-assisted with OpenAI Codex; changes and validation were inspected before submission.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant