Skip to content

perf(kernel): optimize SM120 FP8 GEMM with FP16 accumulation - #1479

Open
STwangyingrui wants to merge 6 commits into
mainfrom
yr/h3-fp8-f16-accum
Open

perf(kernel): optimize SM120 FP8 GEMM with FP16 accumulation#1479
STwangyingrui wants to merge 6 commits into
mainfrom
yr/h3-fp8-f16-accum

Conversation

@STwangyingrui

Copy link
Copy Markdown
Contributor

Optimizes MiniMax-H3 inference on RTX 5090/SM120 with an FP8 GEMM path using FP16 accumulation. It includes exact-shape CUTLASS autotuning with persistent caching, qmax-aware checkpoint conversion and validation, and safe FP8-SGL fallback. The optimized path covers selected DiT and Video VAE decoder projections, with configuration and documentation included.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant