Skip to content

perf(gb300): Add more AgentX vLLM MTP aggregate and P/D recipes / 增加更多 GB300 AgentX vLLM MTP 聚合与 P/D 配方 - #2665

Open
ivanium wants to merge 25 commits into
mainfrom
agent/update-gb300-agentx-vllm-mtp-pd
Open

perf(gb300): Add more AgentX vLLM MTP aggregate and P/D recipes / 增加更多 GB300 AgentX vLLM MTP 聚合与 P/D 配方#2665
ivanium wants to merge 25 commits into
mainfrom
agent/update-gb300-agentx-vllm-mtp-pd

Conversation

@ivanium

@ivanium ivanium commented Aug 19, 2026

Copy link
Copy Markdown
Collaborator

Summary / 摘要

English

  • Add a few TP4 agg and DEP32 disagg recipes.

中文

  • 增加一些额外的 TP4 agg 和 DEP32 disagg 配方。

Validation / 验证

  • 19 generated configs, 232 matrix tests, YAML/topology/launcher/diff checks passed. / 生成 19 个配置,232 项矩阵测试及 YAML、拓扑、启动器和差异检查均通过。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@ivanium
ivanium force-pushed the agent/update-gb300-agentx-vllm-mtp-pd branch from 185f2e3 to ad2c1a8 Compare August 19, 2026 02:02
@ivanium ivanium changed the title perf(gb300): refresh AgentX vLLM MTP P/D recipes / 刷新 GB300 AgentX vLLM MTP P/D 配方 perf(gb300): refresh AgentX vLLM MTP aggregate and P/D recipes / 刷新 GB300 AgentX vLLM MTP 聚合与 P/D 配方 Aug 19, 2026
@ivanium
ivanium marked this pull request as ready for review August 19, 2026 06:29
@ivanium
ivanium requested a review from a team August 19, 2026 06:29
@ivanium
ivanium marked this pull request as draft August 19, 2026 06:42
@ivanium
ivanium marked this pull request as ready for review August 19, 2026 06:46

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline finding on the checked-in synthetic MTP config, I also checked the new 1P3D GB300 recipe (disagg-gb300-1p3d-dep4-tp8-c3-mtp-agentic.yaml) for missing multi-node NCCL/UCX env vars on the cross-node TP8 decode workers — this was examined and ruled out.

Extended reasoning...

Bugs were found in this run and are already posted as inline comments; this note only records an additional item that was examined and ruled out, not a full re-review.

@ivanium
ivanium force-pushed the agent/update-gb300-agentx-vllm-mtp-pd branch from be4e8ef to 023bf49 Compare August 19, 2026 06:57
@ivanium ivanium added agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled labels Aug 19, 2026
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

ivanium and others added 25 commits August 20, 2026 05:59
刷新 GB300 AgentX vLLM MTP P/D 配方。
新增 GB300 AgentX 并发度 512 的 P/D 测试点。
恢复 GB300 AgentX 聚合 TP4/TP8 扫描,并在高并发点启用 vLLM Simple CPU KV 卸载。
将全部 GB300 AgentX 配方回退到 426e59f 镜像,禁用前缀缓存保留间隔,并对齐聚合 TP 的低延迟归约配置。
Restore the GB300 DeepSeek-V4-Pro aggregate TP4 and TP8 recipes to 0.94 GPU memory utilization. The 0.92 setting belongs to the separate low-latency 1P3D decode topology.

将 GB300 DeepSeek-V4-Pro 聚合式 TP4 和 TP8 配方的 GPU 显存利用率恢复为 0.94。0.92 设置仅适用于独立的低延迟 1P3D 解码拓扑。
Remove checked-in synthetic acceptance fields from the four GB300 aggregate recipes. Throughput jobs inject acceptance length 2.49 from nvidia-master.yaml, while eval-only runs retain real target-model verification.

从四个 GB300 聚合式配方中移除检入的合成接受率字段。吞吐任务由 nvidia-master.yaml 注入 2.49 的接受长度,而仅评测任务保留真实目标模型验证。
为聚合 TP4/TP8 配方启用 Mooncake Store,并调整 Simple CPU Offload 调度容量。
将聚合 Simple CPU Offload 的主机 DRAM 预算提高到 80%。
将 TP4 并发 6、8 和 TP8 并发 12 调整为无 CPU 卸载运行。
将聚合高并发卸载路径从 vLLM Simple CPU Offload 更正为 Mooncake Store。
在两个聚合式 Mooncake Store 配置中固定 PYTHONHASHSEED,以确保跨进程和节点的前缀块键保持一致。
将聚合式 TP4 和 TP8 扫描上限统一为并发 16,并添加经过当前分支运行时配置适配的 1P1D DEP8/DEP32 并发 388 P/D 点。
将 c388 DEP32 解码 worker 的 Mooncake 接收线程数从 20 降至 4,并保持 DEP8 预填充 worker 为 20。
将 GB300 AgentX 的 1P1D DEP8/DEP32 数据点从并发 388 调整为并发 384,并同步配方名称、路径和 JIT 缓存标识。
将 GB300 AgentX 的 1P1D DEP8/DEP32 数据点调整到并发 448,并将解码侧 max-num-seqs 提高到 8、max-num-batched-tokens 提高到 32。
添加 GB300 AgentX DEP8 预填充和 DEP32 解码的 c256 配置点,并复用 c448 运行时调优。
保留现有聚合与解聚配置不变,仅新增 TP4 并发点、1P4D c4 以及 DEP16/DEP32 配置,并将性能变更日志标记为 append-only。
@ivanium
ivanium force-pushed the agent/update-gb300-agentx-vllm-mtp-pd branch from 2454f22 to 2a8bf7b Compare August 20, 2026 06:03
@github-actions

Copy link
Copy Markdown
Contributor

@ivanium ivanium changed the title perf(gb300): refresh AgentX vLLM MTP aggregate and P/D recipes / 刷新 GB300 AgentX vLLM MTP 聚合与 P/D 配方 perf(gb300): Add more AgentX vLLM MTP aggregate and P/D recipes / 增加更多 GB300 AgentX vLLM MTP 聚合与 P/D 配方 Aug 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants