GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
-
Updated
Sep 4, 2026 - Python
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights
LLM inference server for ExLlamaV3 / EXL3, with OpenAI- and Anthropic-compatible APIs optimized for Agent workloads.
High-performance runtime extensions for vLLM.
DeepSeek-V4-Flash-Vision-Exp (EXL3 MixedK, 256 experts, uncensored) on one NVIDIA DGX Spark with vLLM + sparkinfer: 245,760 context, vision + DSpark speculative decoding, CUDA graphs. Recipe, overlay patches, benchmarks, receipts.
Containerized private AI lab: TabbyAPI EXL3 + SillyTavern + Open WebUI + Ollama + SearXNG
GLM-5.2 753B MoE at EXL3-TR3 3.0bpw on 4x RTX PRO 6000 (SM120): digest-pinned serving stack, validation suite, results. Upstream: vllm#139 / sparkinfer#49
Deploy GLM-5.3 Flash with DFlash2 speculative decoding on dual RTX PRO 6000 Blackwell 96GB GPUs, enabling one-million-token context and 16-image prompts via an OpenAI-compatible API.
GLM-5.3-Flash EXL3 4bpw on 2x NVIDIA DGX Spark (GB10): production recipe, boot ladder, quality gate, benchmarks and lessons learned — reproducible from CLAUDE.md/AGENTS.md
Run GLM-5.3-Flash, a 320B-parameter model, on two NVIDIA DGX Spark desktops as a private OpenAI-compatible API with 1.3M-token context.
To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."