Skip to content

Repository files navigation

inference-visually

Hosted demo: https://arthur031221.github.io/inference-visually/

Interactive explainers of the LLM serving stack that you drive yourself: prefill and decode, the KV cache, paged attention, continuous batching, prefix caching, speculative decoding, sampling and quantization. Each page pairs a simulator with a short numpy reference implementation and a script that measures the same effect on your own Mac.

Interactive walkthrough of KV cache, paged attention, continuous batching, prefix caching, and speculative decoding

On an M5 MacBook Air, Qwen3-0.6B-4bit decodes at 276 tokens/s after a 32-token prompt and 83 tokens/s after an 8,192-token prompt. Same model, same weights. At 8k context the KV cache read per generated token is 896 MiB against 320 MiB of weights, so 74 percent of the memory traffic per token is cache, not model.1

CI License Version

KV cache calculator

Paged attention block pool

Sampling pipeline

Quantization error histogram

Why

Every visual explainer of transformers stops at the forward pass. The questions people actually hit when they serve a model start after it: why is decode 35 tokens/s on a laptop when prefill is 5,000, how much memory does a 32k context cost, why does vLLM page the cache, what does continuous batching change, how much does a shared system prompt save, what does an acceptance rate of 0.8 buy. The answers are formulas with three or four inputs. This site puts sliders on those inputs, shows the chart, and gives you the 80-line numpy file and the mlx-lm command to check it on your own machine.

Pages

# Page Simulator Reference Measured on this Mac
1 Prefill vs decode Roofline chart, decode throughput versus batch size, chip presets with Apple's bandwidth figures py/roofline.py measure/decode_tps.py
2 KV cache memory Paste a config.json, detects MHA, GQA, MQA, MLA and sliding window, memory versus context chart py/kv_cache.py measure/kv_cache_growth.py
3 Paged attention Block pool and block tables versus contiguous reservations, step by step, fragmentation visible py/paged_attention.py numpy only
4 Continuous batching Static, continuous and chunked-prefill schedulers on one random workload, Gantt timeline, TTFT and throughput py/continuous_batching.py numpy only, costs from decode_tps
5 Prefix caching Radix tree of shared prompts, hit rate, LRU eviction, prefill time and memory saved py/prefix_cache.py measure/prefix_cache_ttft.py
6 Speculative decoding Expected tokens per step and speedup curves, a toy-vocabulary sampler that verifies the output distribution py/speculative_decoding.py numpy only
7 Sampling Temperature, top-k, top-p, min-p on a small logit vector, every pipeline stage in a table, a 200-draw histogram against the theoretical distribution py/sampling.py numpy only
8 Quantization Symmetric and affine group quantization, exact GGUF Q4_0 and Q8_0, a 41-bin error histogram, the Q4_K_M tensor mix on a Llama 3.1 8B shape py/quantization.py numpy only

The landing page has a "why is my Mac 40 tokens per second" calculator: chip, model size, weight format, context, expected decode speed. There is also a 40-question interview page with answers sourced from the other pages' own formulas and measurements.

Install and run locally

git clone https://github.com/Arthur031221/inference-visually
cd inference-visually
npm install
npm run dev

Open http://localhost:5173/inference-visually/.

Quick start with the references

uv sync --group dev
uv run python py/kv_cache.py path/to/config.json 8192 1   # bytes per token and total at 8k context
uv run python py/roofline.py                                # ceiling for an 8B 4-bit model on an M5
uv run pytest                                               # 23 hand-computed checks

Measure it on your Mac

Apple Silicon only. Installs mlx-lm and downloads mlx-community/Qwen3-0.6B-4bit (about 330 MB). The three scripts take about a minute in total.

uv sync --group measure
uv run python measure/decode_tps.py          # prefill and decode tok/s at 32 to 8192 prompt tokens
uv run python measure/kv_cache_growth.py     # cache bytes against 2 * layers * kv_heads * head_dim * 2
uv run python measure/prefix_cache_ttft.py   # time to first token with and without a cached prefix

Each writes measure/results/<name>.json and .txt. The site imports those files, so rerunning the scripts and rebuilding puts your numbers on your copy of the pages.

Results from the machine this repository was built on:1

 prompt  prefill tok/s  decode tok/s  ceiling tok/s  of ceiling  peak GB
     32         1443.2         275.8          441.7        62%     0.44
    256         4173.2         258.9          411.2        63%     0.60
   1024         8802.0         221.1          332.5        66%     1.11
   4096         6905.8         125.8          188.3        67%     1.35
   8192         5497.7          82.5          119.3        69%     1.91
 prefix  cold TTFT ms  warm TTFT ms  speedup  fill once ms
    512         115.8          70.8     1.6x          62.7
   2048         309.2          77.3     4.0x         243.6
   8192        1596.1         118.6    13.5x        1497.0

How it works

  • src/models/ holds one pure TypeScript module per page (no DOM). Every formula in it has a vitest case with a value computed by hand, for example Llama 3.1 8B at 2 x 32 x 8 x 128 x 2 = 131,072 bytes per token and 1 GiB at 8,192 tokens.
  • py/ holds the same computations in numpy, each under 120 lines, shown verbatim on the page's reference tab and tested in py/test_references.py against the same hand values.
  • measure/ holds the mlx-lm scripts. They record the chip, memory, macOS, mlx and mlx-lm versions, model and date next to every number.
  • src/pages/ builds the simulators with D3 and plain DOM. No framework. Hash routing so GitHub Pages needs no rewrite rules.
  • Chip bandwidth presets were checked against apple.com tech specs pages (M4, M5 families) and Wikipedia (M1 to M3) on 2026-09-29. GGUF block layouts were checked against ggml-common.h, MLX quantization against the mlx.core.quantize documentation.

Comparison

Project What it covers What it lacks for this use
poloclub/transformer-explainer GPT-2 forward pass, attention, in the browser Nothing after the forward pass: no KV cache, batching, paging, caching or decoding strategies
bbycroft/llm-viz 3D walkthrough of a small GPT forward pass Same scope, forward pass only
GeeeekExplorer/nano-vllm A readable vLLM in about 1,200 lines of CUDA-side Python Code, not visuals, and needs an NVIDIA GPU to run
skyzh/tiny-llm Text course building an inference engine on MLX Reading and coding exercises, no interactive simulators or measurements
Baseten inferenceengineering.tech Interactive serving explainers Closed source, datacenter GPUs only, no local measurement
inference-visually Eight interactive pages after the forward pass, numpy references, Apple Silicon measurements, 40-question interview page No NVIDIA measurements

Reference

npm run dev        vite dev server
npm run build      typecheck and build to dist/
npm test           vitest (tests/*.test.ts)
npm run lint       eslint
uv run pytest      numpy reference tests
uv run ruff check py measure
uv run python py/<name>.py                 run a reference
uv run python measure/<name>.py [--json]   run a measurement (Apple Silicon)

Deploys to GitHub Pages from main through .github/workflows/pages.yml (Vite base is /inference-visually/).

Limits and FAQ

Is the roofline exact? No. It counts matmul FLOPs (2 per parameter per token) and ignores attention FLOPs, which matter at very long context. It treats the KV cache as fully read every step. Peak compute is a slider because Apple publishes bandwidth but not GPU TFLOP/s.

Why is the measured decode speed 62 to 69 percent of the ceiling? Runtime overhead, attention, sampling and the parts of a step that are not a weight read. Bigger models get closer to the ceiling because the weight read dominates.

Why a 0.6B model? It is small enough to keep total measurement time under two minutes on a 24 GB laptop and shows every effect: its KV cache is 112 KiB per token, so the cache overtakes the weights at about 3k tokens.

Does the paged attention page measure anything? No. mlx-lm serves one conversation with one cache, so there is nothing to fragment. The numpy file shows the block-table gather and checks it against dense attention.

What about NVIDIA GPUs? The simulators are hardware agnostic (enter your bandwidth and compute). The measurement scripts are mlx-lm only.

Related projects

  • llm-doctor: Diagnoses the local serving setup these explainers assume is already working.
  • gpuwho: Shows the GPU activity behind the throughput numbers inference-visually's calculators use.
  • gpuwait: Measures the idle time between requests that the continuous batching page covers in theory.

Contributing and license

See CONTRIBUTING.md. MIT, copyright 2026 Arthur.

Footnotes

  1. MacBook Air, Apple M5, 24 GB, macOS 26.6, mlx 0.32.3, mlx-lm 0.31.3, mlx-community/Qwen3-0.6B-4bit, 2026-09-30. Decode: median of 3 runs of 128 greedy tokens per prompt length, measure/decode_tps.py. Ceiling: 153 GB/s (Apple's M5 figure) divided by weight bytes (319.8 MiB) plus KV cache bytes at the average context of the run. KV: cache arrays read after prefilling the prompt, measure/kv_cache_growth.py. Prefix: median of 5 trials, measure/prefix_cache_ttft.py. ↩ ↩2

About

Interactive explainers of the LLM serving stack: KV cache, paged attention, continuous batching, prefix caching, speculative decoding, measured on your Mac

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages