Skip to content
@vLLM-HUST

vLLM-HUST

Upstream-compatible vLLM fork organization for domestic hardware enablement and AGI4S serving

vLLM-HUST

vLLM-HUST 面向国产算力维护 vLLM 推理运行时、Ascend 适配和配套工程工具。研究项目以成果仓库的形式保存核心实现、论文、实验和运行时插件。

vLLM-HUST maintains an upstream-compatible vLLM stack for domestic hardware, together with research artifacts, runtime plugins, benchmarks, and developer tooling.

快速入口 仓库
推理运行时 vllm-hust · vllm-ascend-hust
插件生态 插件目录与路线图
编译与模型工具 triton-ascend-hust · vllm-ascend-quant-hust
成果仓库 BidKV · DiffSpec
开发与验证 Dev Hub · Benchmark · Performance Analyzer
文档与展示 Docs · Website · Workstation

生态边界 / Ecosystem Boundary

SAGE(Streaming-Augmented Generative Execution)是 IntelliStream 研究生态的共同旗舰产品,以流计算思维赋能大模型推理与智能体执行。RIDE Lab 负责其核心仓库的主要维护,Sage Mate 是基于 SAGE 构建的应用。这些调用方使用 vLLM-HUST 完成模型执行。

SAGE — Streaming-Augmented Generative Execution — is the IntelliStream research ecosystem's shared flagship product, applying streaming-computing principles to LLM inference and agent execution. RIDE Lab is the principal steward of its core repositories, while Sage Mate is an application built with SAGE. These caller-side systems use vLLM-HUST for model execution.

vLLM-HUST 是独立的推理底座,拥有模型执行、KV-cache、解码调度、编译、算子与硬件后端;RIDE Lab 和 SAGE 不是 vLLM-HUST 内部的运行时层。

成果仓库 / Research Outcomes

每个成果仓库对应一项可独立使用的研究成果,集中维护核心实现、论文、实验和运行时集成。

Repository Core technique Publication Team Runtime integration
BidKV Utility-guided KV-cache victim selection BidKV: Utility-Guided Preemption Scheduling for KV-Pressure LLM Serving, SC 2026 主要作者:陈彦博、王明琪
指导老师:张书豪
vllm-hust · vllm-ascend-hust · vLLM / SGLang adapters
DiffSpec Differential speculative decoding DiffSpec: Accelerating Long Sequence Generation with Differential Speculative Decoding 主要作者:杜忠承
指导老师:黄禹
vllm-hust · vllm-ascend-hust · vLLM adapters

仓库索引 / Repository Index

类别 仓库 用途
核心运行时 vllm-hust · vllm-ascend-hust vLLM 服务与 Ascend 硬件插件
编译与算子基础设施 triton-ascend-hust Triton Ascend 编译后端;为运行时插件提供基础设施,但本身不是插件
成果仓库 vllm-ascend-hust-bidkv 核心技术、论文、实验与运行时插件
成果仓库 vllm-ascend-hust-diffspec 核心技术、论文、实验与运行时集成
Ascend 工具 vllm-ascend-quant-hust · ascend-runtime-manager 模型量化准备、环境诊断与运行时维护;不作为 vLLM 运行时插件展示
开发与验证 vllm-hust-dev-hub · vllm-hust-benchmark · vllm-hust-perf-analyzer · claude-code-hust 多仓工作区、Benchmark、Profiler 分析与开发工具
产品与应用 vllm-hust-website · vllm-hust-workstation · EvoScientist 官网、Web 工作台与科研智能体应用
文档与社区 vllm-hust-docs · .github · vllm-hust.github.io 文档、组织级社区配置与 Pages 入口
论文与活动 CCCF 综述 · fcs-domestic-chip-llm-recsys(链接未公开) · StateSys 2026 论文仓库与学术活动

Fork Status

核心 fork 持续跟踪上游;精确基线以各仓库的 upstream_version.json 为准。

Repository Upstream
vllm-hust vllm-project/vllm
vllm-ascend-hust vllm-project/vllm-ascend
triton-ascend-hust triton-lang/triton-ascend
EvoScientist EvoScientist/EvoScientist

Publications

Paper Venue Repository
BidKV: Utility-Guided Preemption Scheduling for KV-Pressure LLM Serving SC 2026 BidKV
DiffSpec: Accelerating Long Sequence Generation with Differential Speculative Decoding SC 2026 DiffSpec
国产算力推理引擎综述 CCCF 通讯专刊 cccf-domestic-inference-engine-survey
LLM-Powered Recommendation Systems on Domestic AI Chips Frontiers of Computer Science fcs-domestic-chip-llm-recsys(链接未公开)

贡献者

贡献者名单、身份合并规则与统计方法见 CONTRIBUTORS.md。

Contributing

Popular repositories Loading

  1. vllm-hust-dev-hub vllm-hust-dev-hub Public

    C++ 3 3

  2. vllm-hust-perf-analyzer vllm-hust-perf-analyzer Public

    A performance analyzer for vllm-hust.

    C++ 3 1

  3. vllm-hust-website vllm-hust-website Public

    vllm-hust official website and benchmark-driven serving showcase

    Python 1 1

  4. vllm-hust-docs vllm-hust-docs Public

    TeX 1

  5. EvoScientist EvoScientist Public

    Forked from EvoScientist/EvoScientist

    🔬 Harness Vibe Research with Self-evolving AI Scientists

    Python 1

  6. vllm-hust-benchmark vllm-hust-benchmark Public

    Python 1 3

Repositories

Showing 10 of 55 repositories
  • vllm-hust-kv-materialization-arrival-control Public

    vLLM-HUST MOD for arrival-time KV materialization policy research

    vLLM-HUST/vllm-hust-kv-materialization-arrival-control's past year of commit activity
    Python 0 0 3 2 Updated Sep 28, 2026
  • vllm-hust Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vLLM-HUST/vllm-hust's past year of commit activity
    Python 0 Apache-2.0 22,993 3 5 Updated Sep 28, 2026
  • vLLM-HUST/ascend-distributed-metadata's past year of commit activity
    Python 0 0 0 0 Updated Sep 28, 2026
  • sglang-hust Public Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    vLLM-HUST/sglang-hust's past year of commit activity
    Python 0 Apache-2.0 9,245 0 0 Updated Sep 28, 2026
  • vllm-ascend-pyramidkv-hust Public

    PyramidKV compression provider for the vLLM-HUST Ascend stack

    vLLM-HUST/vllm-ascend-pyramidkv-hust's past year of commit activity
    Python 0 Apache-2.0 0 2 1 Updated Sep 28, 2026
  • vllm-ascend-kvcompress-hust Public

    An extensible KV-cache compression plugin for the HUST-maintained vLLM Ascend stack. It connects compression methods to vLLM-HUST's transactional KV-cache lifecycle without modifying the vLLM-HUST or vLLM-Ascend-HUST source trees.

    vLLM-HUST/vllm-ascend-kvcompress-hust's past year of commit activity
    Python 0 Apache-2.0 0 3 0 Updated Sep 28, 2026
  • vllm-ascend-hust Public Forked from vllm-project/vllm-ascend

    Community maintained hardware plugin for vLLM on Ascend

    vLLM-HUST/vllm-ascend-hust's past year of commit activity
    Python 0 Apache-2.0 2,540 5 4 Updated Sep 28, 2026
  • vllm-hust-website Public

    vllm-hust official website and benchmark-driven serving showcase

    vLLM-HUST/vllm-hust-website's past year of commit activity
    Python 1 1 1 4 Updated Sep 28, 2026
  • BetterScale Public

    Incremental DeepSeek V4 improvements on pinned vLLM and vLLM-Ascend.

    vLLM-HUST/BetterScale's past year of commit activity
    Python 0 Apache-2.0 1 3 0 Updated Sep 28, 2026
  • vllm-metal-hust Public Forked from vllm-project/vllm-metal

    Community maintained hardware plugin for vLLM on Apple Silicon

    vLLM-HUST/vllm-metal-hust's past year of commit activity
    Python 0 Apache-2.0 274 0 0 Updated Sep 27, 2026