diff --git a/README.md b/README.md index 44741c99d..923f75c60 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@
Same tokens as vLLM. Same throughput. 140x less to install.
- Continuous batching, paged KV, 37 registered architectures, CUDA / CPU / Metal / Vulkan. No Python anywhere.
+ Continuous batching, paged KV, 40 registered architectures, CUDA / CPU / Metal / Vulkan. No Python anywhere.
@@ -37,15 +37,15 @@
## News
+- **2026-08** **MiniMax-Music3 generates music through the OpenAI-compatible server.** Every model
+ stage and a real `POST /v1/audio/speech` request are gated. The device path is partial, and no
+ reference speed number is published yet.
- **2026-08** **v0.0.2 ships eight server archives.** Download CPU, CUDA, Vulkan, Metal, and MLX
builds from [GitHub Releases](https://github.com/mudler/vllm.cpp/releases/tag/v0.0.2).
- **2026-08** **MiniMax-H3 generates video with audio.** All tasks run through `POST /v1/videos`;
use Q4_K_M.
- **2026-08** **MXFP4 holds parity with vLLM.** Qwen3-8B MXFP4 runs W4A16 Marlin by default,
matches the vLLM oracle token for token, and decodes **45.45 vs 41.94 tok/s**.
-- **2026-08** **Vulkan matches llama.cpp on a 27B.** Qwen3.6-27B decodes **4.36 vs llama.cpp Vulkan
- 4.35 tok/s** on GB10, up from 2.40. A narrow pass: the 0.69% leg spread is the noise floor.
- Prefill **21.5x**, a self-ratio. Denominator SUPERSEDED.
vllm.cpp is a from-scratch C++20 inference engine chasing three things at once: be the
**smallest** thing you can deploy, be the **fastest** on the hardware you already own, and still
@@ -77,9 +77,9 @@ Where that stands today:
ahead at all six concurrencies but only c1 outside our noise band. Also **1.18x llama.cpp's
prefill** on the same GGUF file (denominator SUPERSEDED, see below), and **ahead of MLX-LM on
prefill** on Apple Silicon. Most other architectures are speed-pending, and say so.
-- **Everything.** 37 registered architectures, 36 tool-parser families, structured output including
- GBNF, three speculative decoders, image and video and audio input, external KV offload, Prometheus
- metrics, and the SGLang knobs, all in a library you can `dlopen`.
+- **Everything.** 40 registered architectures, 38 tool-parser families, structured output including
+ GBNF, three speculative decoders, multimodal input, music generation, external KV offload,
+ Prometheus metrics, and the SGLang knobs, all in a library you can `dlopen`.
## Performance
@@ -187,7 +187,7 @@ configs, token-for-token the same output. Switching to it should be boring. Ever
you get on top, most of it borrowed from whichever engine does it best:
- **One 66 MiB binary instead of a 9.1 GiB install.** A flat, exception-free, llama.cpp-style C ABI
- ([`include/vllm.h`](include/vllm.h), ABI v19, 36 functions) for C, C++, Go, or Rust. No Python
+ ([`include/vllm.h`](include/vllm.h), ABI v21) for C, C++, Go, or Rust. No Python
interpreter in the process.
- **GGUF as a first-class citizen.** Load the same quantized files llama.cpp uses, and on CPU
**compute directly on the compressed blocks** (Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K) with no BF16
@@ -217,7 +217,7 @@ you get on top, most of it borrowed from whichever engine does it best:
sample logprobs.
- **Structured output.** JSON schema, JSON object, regex, choice, and GBNF grammar, enforced in the
engine with a per-step logits bitmask.
-- **Tool calling and reasoning.** 36 tool-parser families (40 accepted names) and 12 reasoning
+- **Tool calling and reasoning.** 38 tool-parser families (42 accepted names) and 12 reasoning
parser names, streaming, selectable with `--tool-call-parser` / `--reasoning-parser`. Chat templates
render through the vendored google/minja engine, the same renderer llama.cpp ships.
- **Multimodal.** Image, video, and audio to text, correctness-complete. Image chat requests are
@@ -250,7 +250,7 @@ InternLM2/3, MiniCPM and MiniCPM3, Yi, OPT, plus Qwen3-VL and Qwen3.6-27B vision
and Voxtral (audio).
The full architecture matrix (37 registered architectures grouped by family)
+The full architecture matrix (40 registered architectures grouped by family)
| Architecture | Example checkpoint | GGUF | Correctness | Speed |
|---|---|:---:|---|---|
@@ -293,7 +293,7 @@ sampler, no logits); upstream is `vllm-project/vllm-omni`. Five conditioning mod
Compressed-tensors NVFP4A16 (W4A16) dense weights also load and compute natively
(RedHatAI/Qwen3-32B-NVFP4A16). Long-context RoPE (YaRN, Llama-3, LongRoPE, dynamic-NTK) and
sliding-window attention are gated feature-positive. The authoritative per-architecture list, bound
-to the C++ registry (all 37 registered architectures with their tested checkpoint and gate, plus the
+to the C++ registry (all 40 registered architectures with their tested checkpoint and gate, plus the
standalone audio/diffusion lanes and the inventoried-but-blocked archs), is in
[docs/FEATURES.md](docs/FEATURES.md); family-by-family lifecycle detail, including what is
hardware-blocked and why, is in [docs/STATUS.md](docs/STATUS.md).