Skip to content
#

multi-turboquant

Here are 2 public repositories matching this topic...

Native Windows vLLM 0.27.1 wheels: Python 3.13, PyTorch 2.13 + CUDA 13.0, SM 7.5-12.0 for RTX 20/30/40/50, OpenAI-compatible serving, FlashAttention/Rust, 10 KV formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload - no WSL or Docker.

  • Updated Aug 21, 2026
  • Python

Add this topic to your repo

To associate your repository with the multi-turboquant topic, visit your repo's landing page and select "manage topics."

Learn more