Skip to content

feat(inference): opt-in fused Snake CUDA decoder - #538

Open
zain1806481 wants to merge 1 commit into
remsky:masterfrom
zain1806481:feat/opt-in-decoder-snake-fusion
Open

zain1806481 wants to merge 1 commit into
remsky:masterfrom
zain1806481:feat/opt-in-decoder-snake-fusion

Conversation

@zain1806481

Copy link
Copy Markdown

Summary

  • Opt-in via KOKORO_DECODER_FUSION=1 after the model is on CUDA
  • Freezes weight-norm modules and fuses Snake activations into one contiguous FP32 CUDA kernel (NVRTC)
  • Fails closed when CUDA/NVRTC/FP32 requirements are not met; never enables a CPU fallback

Measured (Quadro K620, sm_50)

Alternating original/fused on the same frozen-weight model, fixed seed:

  • ~8.5–8.8% wall-time reduction
  • Exact waveform and duration match on every paired trial

Other GPUs may work (kernel targets the current device arch) but were not the validation target.

Test plan

  • Boot API without the env flag (unchanged path)
  • Boot with KOKORO_DECODER_FUSION=1 on CUDA FP32 and confirm log line Experimental decoder fusion enabled
  • Compare a short synthesis against stock for audible/waveform parity on your GPU
  • Confirm missing NVRTC fails closed with a clear error

Made with Cursor

Add KOKORO_DECODER_FUSION=1 path that freezes weight-norm layers and
replaces AdaIN residual Snake activations with one NVRTC FP32 kernel.
Fails closed without CUDA/NVRTC; no CPU fallback. Measured ~8.5% faster
with exact audio on Quadro K620 (sm_50).

Co-authored-by: Cursor <cursoragent@cursor.com>
@zain1806481

Copy link
Copy Markdown
Author

Supporting notes for this PR:

Fusion remains opt-in via KOKORO_DECODER_FUSION=1 and fails closed without CUDA/NVRTC.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant