[ExecuTorch][WebGPU] Add Gemma 4 plain runtime and guarded routes - #21672
[ExecuTorch][WebGPU] Add Gemma 4 plain runtime and guarded routes#21672JCNTH wants to merge 1 commit into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21672
Note: Links to docs will display an error until the docs builds have been completed. ✅ You can merge normally! (2 Unrelated Failures)As of commit bb78e08 with merge base ceca90f ( FLAKY - The following job failed but was likely due to flakiness present on trunk:
BROKEN TRUNK - The following job failed but was present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
|
Hi @JCNTH! Thank you for your pull request. We require contributors to sign our Contributor License Agreement, and yours needs attention. You currently have a record in our system, but the CLA is no longer valid, and will need to be resubmitted. ProcessIn order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA. Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks! |
Stack from ghstack (oldest at bottom):
Enable the plain Gemma 4 WebGPU runtime at long context with paired 2D Slice correctness and guarded routes. This extends attention, quantized linear and embedding, RMSNorm, RoPE, symbolic selection, and dynamic dispatch support while preserving generic fallbacks.
Key changes:
Slice.cppandslice.wgslland folded X/Y dispatch and flattening atomically, including both dimensions on resize.WebGPUGraphowns and clears Slice, CQP, and RMS fusion state across every build lifecycle path.Plain Gemma remains independently landable before MTP. The historical raw-FP32 BK32 requirement is superseded and is neither implemented nor claimed here; final-source GPU correctness and performance remain external gates.
Co-authored-with: Claude Code.
Differential Revision: D115234078