Hosted demo: https://arthur031221.github.io/inference-visually/
Interactive explainers of the LLM serving stack that you drive yourself: prefill and decode, the KV cache, paged attention, continuous batching, prefix caching, speculative decoding, sampling and quantization. Each page pairs a simulator with a short numpy reference implementation and a script that measures the same effect on your own Mac.
On an M5 MacBook Air, Qwen3-0.6B-4bit decodes at 276 tokens/s after a 32-token prompt and 83 tokens/s after an 8,192-token prompt. Same model, same weights. At 8k context the KV cache read per generated token is 896 MiB against 320 MiB of weights, so 74 percent of the memory traffic per token is cache, not model.1
Every visual explainer of transformers stops at the forward pass. The questions people actually hit when they serve a model start after it: why is decode 35 tokens/s on a laptop when prefill is 5,000, how much memory does a 32k context cost, why does vLLM page the cache, what does continuous batching change, how much does a shared system prompt save, what does an acceptance rate of 0.8 buy. The answers are formulas with three or four inputs. This site puts sliders on those inputs, shows the chart, and gives you the 80-line numpy file and the mlx-lm command to check it on your own machine.
| # | Page | Simulator | Reference | Measured on this Mac |
|---|---|---|---|---|
| 1 | Prefill vs decode | Roofline chart, decode throughput versus batch size, chip presets with Apple's bandwidth figures | py/roofline.py |
measure/decode_tps.py |
| 2 | KV cache memory | Paste a config.json, detects MHA, GQA, MQA, MLA and sliding window, memory versus context chart | py/kv_cache.py |
measure/kv_cache_growth.py |
| 3 | Paged attention | Block pool and block tables versus contiguous reservations, step by step, fragmentation visible | py/paged_attention.py |
numpy only |
| 4 | Continuous batching | Static, continuous and chunked-prefill schedulers on one random workload, Gantt timeline, TTFT and throughput | py/continuous_batching.py |
numpy only, costs from decode_tps |
| 5 | Prefix caching | Radix tree of shared prompts, hit rate, LRU eviction, prefill time and memory saved | py/prefix_cache.py |
measure/prefix_cache_ttft.py |
| 6 | Speculative decoding | Expected tokens per step and speedup curves, a toy-vocabulary sampler that verifies the output distribution | py/speculative_decoding.py |
numpy only |
| 7 | Sampling | Temperature, top-k, top-p, min-p on a small logit vector, every pipeline stage in a table, a 200-draw histogram against the theoretical distribution | py/sampling.py |
numpy only |
| 8 | Quantization | Symmetric and affine group quantization, exact GGUF Q4_0 and Q8_0, a 41-bin error histogram, the Q4_K_M tensor mix on a Llama 3.1 8B shape | py/quantization.py |
numpy only |
The landing page has a "why is my Mac 40 tokens per second" calculator: chip, model size, weight format, context, expected decode speed. There is also a 40-question interview page with answers sourced from the other pages' own formulas and measurements.
git clone https://github.com/Arthur031221/inference-visually
cd inference-visually
npm install
npm run devOpen http://localhost:5173/inference-visually/.
uv sync --group dev
uv run python py/kv_cache.py path/to/config.json 8192 1 # bytes per token and total at 8k context
uv run python py/roofline.py # ceiling for an 8B 4-bit model on an M5
uv run pytest # 23 hand-computed checksApple Silicon only. Installs mlx-lm and downloads mlx-community/Qwen3-0.6B-4bit (about 330 MB). The three scripts take about a minute in total.
uv sync --group measure
uv run python measure/decode_tps.py # prefill and decode tok/s at 32 to 8192 prompt tokens
uv run python measure/kv_cache_growth.py # cache bytes against 2 * layers * kv_heads * head_dim * 2
uv run python measure/prefix_cache_ttft.py # time to first token with and without a cached prefixEach writes measure/results/<name>.json and .txt. The site imports those files, so rerunning the scripts and rebuilding puts your numbers on your copy of the pages.
Results from the machine this repository was built on:1
prompt prefill tok/s decode tok/s ceiling tok/s of ceiling peak GB
32 1443.2 275.8 441.7 62% 0.44
256 4173.2 258.9 411.2 63% 0.60
1024 8802.0 221.1 332.5 66% 1.11
4096 6905.8 125.8 188.3 67% 1.35
8192 5497.7 82.5 119.3 69% 1.91
prefix cold TTFT ms warm TTFT ms speedup fill once ms
512 115.8 70.8 1.6x 62.7
2048 309.2 77.3 4.0x 243.6
8192 1596.1 118.6 13.5x 1497.0
src/models/holds one pure TypeScript module per page (no DOM). Every formula in it has a vitest case with a value computed by hand, for example Llama 3.1 8B at 2 x 32 x 8 x 128 x 2 = 131,072 bytes per token and 1 GiB at 8,192 tokens.py/holds the same computations in numpy, each under 120 lines, shown verbatim on the page's reference tab and tested inpy/test_references.pyagainst the same hand values.measure/holds the mlx-lm scripts. They record the chip, memory, macOS, mlx and mlx-lm versions, model and date next to every number.src/pages/builds the simulators with D3 and plain DOM. No framework. Hash routing so GitHub Pages needs no rewrite rules.- Chip bandwidth presets were checked against apple.com tech specs pages (M4, M5 families) and Wikipedia (M1 to M3) on 2026-09-29. GGUF block layouts were checked against
ggml-common.h, MLX quantization against the mlx.core.quantize documentation.
| Project | What it covers | What it lacks for this use |
|---|---|---|
| poloclub/transformer-explainer | GPT-2 forward pass, attention, in the browser | Nothing after the forward pass: no KV cache, batching, paging, caching or decoding strategies |
| bbycroft/llm-viz | 3D walkthrough of a small GPT forward pass | Same scope, forward pass only |
| GeeeekExplorer/nano-vllm | A readable vLLM in about 1,200 lines of CUDA-side Python | Code, not visuals, and needs an NVIDIA GPU to run |
| skyzh/tiny-llm | Text course building an inference engine on MLX | Reading and coding exercises, no interactive simulators or measurements |
| Baseten inferenceengineering.tech | Interactive serving explainers | Closed source, datacenter GPUs only, no local measurement |
| inference-visually | Eight interactive pages after the forward pass, numpy references, Apple Silicon measurements, 40-question interview page | No NVIDIA measurements |
npm run dev vite dev server
npm run build typecheck and build to dist/
npm test vitest (tests/*.test.ts)
npm run lint eslint
uv run pytest numpy reference tests
uv run ruff check py measure
uv run python py/<name>.py run a reference
uv run python measure/<name>.py [--json] run a measurement (Apple Silicon)
Deploys to GitHub Pages from main through .github/workflows/pages.yml (Vite base is /inference-visually/).
Is the roofline exact? No. It counts matmul FLOPs (2 per parameter per token) and ignores attention FLOPs, which matter at very long context. It treats the KV cache as fully read every step. Peak compute is a slider because Apple publishes bandwidth but not GPU TFLOP/s.
Why is the measured decode speed 62 to 69 percent of the ceiling? Runtime overhead, attention, sampling and the parts of a step that are not a weight read. Bigger models get closer to the ceiling because the weight read dominates.
Why a 0.6B model? It is small enough to keep total measurement time under two minutes on a 24 GB laptop and shows every effect: its KV cache is 112 KiB per token, so the cache overtakes the weights at about 3k tokens.
Does the paged attention page measure anything? No. mlx-lm serves one conversation with one cache, so there is nothing to fragment. The numpy file shows the block-table gather and checks it against dense attention.
What about NVIDIA GPUs? The simulators are hardware agnostic (enter your bandwidth and compute). The measurement scripts are mlx-lm only.
- llm-doctor: Diagnoses the local serving setup these explainers assume is already working.
- gpuwho: Shows the GPU activity behind the throughput numbers inference-visually's calculators use.
- gpuwait: Measures the idle time between requests that the continuous batching page covers in theory.
See CONTRIBUTING.md. MIT, copyright 2026 Arthur.
Footnotes
-
MacBook Air, Apple M5, 24 GB, macOS 26.6, mlx 0.32.3, mlx-lm 0.31.3, mlx-community/Qwen3-0.6B-4bit, 2026-09-30. Decode: median of 3 runs of 128 greedy tokens per prompt length,
measure/decode_tps.py. Ceiling: 153 GB/s (Apple's M5 figure) divided by weight bytes (319.8 MiB) plus KV cache bytes at the average context of the run. KV: cache arrays read after prefilling the prompt,measure/kv_cache_growth.py. Prefix: median of 5 trials,measure/prefix_cache_ttft.py. ↩ ↩2




