FlashAttention-2 forward and Flash-Decoding kernels written from scratch in Triton, with grouped-query attention (GQA) and causal masking, benchmarked against PyTorch SDPA and a naive eager implementation, then run end-to-end inside a from-scratch Llama-3.2-1B.
benchmarking flash kernel decoding speed triton autotuning profiling sdpa gqa kv-cache fa2 flash-attention flashdecoding
-
Updated
Sep 21, 2026 - Jupyter Notebook