Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -108,3 +108,4 @@ test-results/
ROCM_TRIM.md
Kokoro-FastAPI.code-workspace
experiments/*
api/src/models/v1_0/inno_tuner/
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,10 @@ Per-PR attribution and contributor credits are published automatically on the co
### Added
- `normalization_options.remove_emoji` drops emoji before synthesis instead of reading them by name, any language (#353). Off by default.
- `normalization_options.caps_normalization` reads all-caps headers and names (`ARNE SAKNUSSEMM`, `TODO_LIST`) as words instead of letter by letter. On by default. Short acronyms (`FBI`, `US GDP`) are still spelled.
- `POST /dev/tune`: tune a voice from a short reference clip and speak with it in one request, via [inno-kokoro](https://github.com/remsky/inno-kokoro). Off by default, `ENABLE_INNO_TUNER=true` turns it on. See [docs/inno-tune.md](docs/inno-tune.md).
- `return_voice_pack=true` returns the tuned `.pt` instead of audio; `save_voice=<name>` keeps it in `VOICES_DIR` as `<name>_tuned`, behind `ALLOW_LOCAL_VOICE_SAVING`.
- Four tuned voices bundled with the server, `_inno` suffix: `af_amelia_inno`, `af_goodall_inno`, `am_price_inno`, `bm_atten_inno`.
- Web player: Tune tab, record or upload a clip and generate with it, download the pack, or save it to the server.

### Changed
- Text normalization refactored towards multi-language support: `Normalizer` base class w/ neutral passes, `EnglishNormalizer` implements the rest, registry keyed by lang code. Adding a language is a subclass + table test, see CONTRIBUTING.md.
Expand Down
39 changes: 29 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,16 +10,14 @@
[![Misaki](https://img.shields.io/badge/misaki-0.9.4-B8860B)](https://github.com/hexgrad/misaki)
[![Tested at Model Commit](https://img.shields.io/badge/model-1.0::41e5892-blue)](https://huggingface.co/hexgrad/Kokoro-82M/commit/41e5892b9d8b43e56fc560f892312a328a410973)

[![Try on Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Try%20on-Spaces-blue)](https://huggingface.co/spaces/Remsky/FastKoko) [![Downloads](https://img.shields.io/badge/downloads-2.5M%2B-2496ED?logo=docker&logoColor=white)](https://github.com/remsky?tab=packages&repo_name=Kokoro-FastAPI)
[![Try on Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Try%20on-Spaces-blue)](https://huggingface.co/spaces/Remsky/FastKoko) [![Downloads](https://img.shields.io/badge/downloads-2.6M%2B-2496ED?logo=docker&logoColor=white)](https://github.com/remsky?tab=packages&repo_name=Kokoro-FastAPI)


Dockerized FastAPI wrapper for [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) text-to-speech model. Generate hours of high quality speech in minutes.

> [!NOTE]
> Looking for custom voices? Try the [Inno Clone-Tuner](https://github.com/remsky/inno-kokoro)
- OpenAI-compatible Speech endpoint, multi-language support
- English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, Mandarin Chinese
- Custom voicepack generation via [Inno Clone-Tuner](https://github.com/remsky/inno-kokoro)
- Optional integrated WebUI; read-along long-generation
- Inline multi-speaker generation & voice mixing + aliasing weighted combinations, SSML support
- Per-word, or per-chunk timestamped caption generation
Expand All @@ -38,11 +36,11 @@ Dockerized FastAPI wrapper for [Kokoro-82M](https://huggingface.co/hexgrad/Kokor

Community projects that use, recommend, or enable Kokoro-FastAPI as a backend:

- Home Assistant: [wyoming_openai](https://github.com/roryeckel/wyoming_openai), [openai_tts](https://github.com/sfortis/openai_tts), [Kokoro-TTS](https://github.com/beecho01/Kokoro-TTS)
- App stores and templates: [Umbrel](https://github.com/getumbrel/umbrel-apps/tree/master/kokoro), [Unraid Community Apps](https://github.com/nwithan8/unraid_templates), [GPUStack](https://github.com/gpustack/gpustack), [jetson-containers](https://github.com/dusty-nv/jetson-containers/tree/master/packages/speech/kokoro-tts)
- Readers and audiobooks: [openreader](https://github.com/richardr1126/openreader), [epub_to_audiobook](https://github.com/p0n1/epub_to_audiobook), [audiobook-creator](https://github.com/prakharsr/audiobook-creator), [Zotero-TTS](https://github.com/xujialiu/Zotero-TTS)
- Assistants and agents: [xiaozhi-esp32-server](https://github.com/xinnan-tech/xiaozhi-esp32-server), [call-me](https://github.com/ZeframLou/call-me), [agent-cli](https://github.com/basnijholt/agent-cli), [voice-chat-ai](https://github.com/bigsk1/voice-chat-ai)
- Browser: [kokoro-extension](https://github.com/Fooftilly/kokoro-extension), [customtts](https://github.com/BassGaming/customtts)
- <sub>Home Assistant: [wyoming_openai](https://github.com/roryeckel/wyoming_openai), [openai_tts](https://github.com/sfortis/openai_tts), [Kokoro-TTS](https://github.com/beecho01/Kokoro-TTS)</sub>
- <sub>App stores and templates: [Umbrel](https://github.com/getumbrel/umbrel-apps/tree/master/kokoro), [Unraid Apps](https://github.com/nwithan8/unraid_templates), [GPUStack](https://github.com/gpustack/gpustack), [jetson-containers](https://github.com/dusty-nv/jetson-containers/tree/master/packages/speech/kokoro-tts)</sub>
- <sub>Readers and audiobooks: [openreader](https://github.com/richardr1126/openreader), [epub_to_audiobook](https://github.com/p0n1/epub_to_audiobook), [audiobook-creator](https://github.com/prakharsr/audiobook-creator), [Zotero-TTS](https://github.com/xujialiu/Zotero-TTS)</sub>
- <sub>Assistants and agents: [xiaozhi-esp32-server](https://github.com/xinnan-tech/xiaozhi-esp32-server), [call-me](https://github.com/ZeframLou/call-me), [agent-cli](https://github.com/basnijholt/agent-cli), [voice-chat-ai](https://github.com/bigsk1/voice-chat-ai)</sub>
- <sub>Browser: [kokoro-extension](https://github.com/Fooftilly/kokoro-extension), [customtts](https://github.com/BassGaming/customtts)</sub>

## Get Started

Expand Down Expand Up @@ -410,6 +408,27 @@ curl -X POST http://localhost:8880/v1/audio/speech \
</details>
<details>
<summary>Voice Tuning (reference clip) 🧪</summary>
`POST /dev/tune` takes a 3 to 30 s clip of one English speaker and speaks with a voice tuned toward it, via [inno-kokoro](https://github.com/remsky/inno-kokoro). A tuner, not a cloner: expect the same neighbourhood, not a match. Only tune voices you have permission to use.
```bash
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F 'request={"input":"Hello there."}' -o out.mp3
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F return_voice_pack=true -o ref.pt
curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F save_voice=am_ref # saves am_ref_tuned, needs ALLOW_LOCAL_VOICE_SAVING=true
```
- `request` is the `/v1/audio/speech` body minus `voice`; the pack only exists for the length of the response unless `save_voice` keeps it in `VOICES_DIR`
- Names follow the existing language prefix pattern (`am_`, `bf_`, `ax_`, etc.) and saved voices get a `_tuned` suffix. Bundled tuned voices end in `_inno`
- Knobs: `prosody_head` (default on), `fmax` pitch ceiling in Hz (default auto)
- Off by default:
- `ENABLE_INNO_TUNER=true` enables with a web player tab
- `ALLOW_LOCAL_VOICE_SAVING=true` allows saving them to the live server.
- Full reference in [docs/inno-tune.md](docs/inno-tune.md)
</details>
<details>
<summary>Multi-Speaker / Dialogue</summary>
Expand Down Expand Up @@ -493,7 +512,7 @@ The city of [Worcester](/wˈʊstər/) is easy. [pause:1s] See?
</details>
<details>
<summary>SSML Input (experimental)</summary>
<summary>SSML Input 🧪</summary>
Send `ssml: true` with `allow_voice_tags: true` on `/v1/audio/speech` or `/dev/captioned_speech` to translate and speak in one call. Both flags are needed, since the translation emits `[voice:]` and `[rate:]` spans that would otherwise be read aloud; `ssml` without them is a 400.
Expand Down
1 change: 1 addition & 0 deletions api/src/core/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,7 @@ class Settings(BaseSettings):
False # Whether to allow saving combined voices locally
)
allow_dev_unload: bool = False # Whether to expose /dev/model, POST /dev/unload, and POST /dev/reload
enable_inno_tuner: bool = False # Whether to expose POST /dev/tune
model_auto_unload_timeout_seconds: float = (
0.0 # Idle seconds before unloading; 0 disables auto-unload
)
Expand Down
35 changes: 14 additions & 21 deletions api/src/core/paths.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,17 @@
}


_API_DIR = os.path.dirname(os.path.dirname(os.path.dirname(__file__)))


def models_dir() -> str:
return os.path.join(_API_DIR, settings.model_dir)


def voices_dir() -> str:
return os.path.join(_API_DIR, settings.voices_dir)


async def _find_file(
filename: str,
search_paths: List[str],
Expand Down Expand Up @@ -112,13 +123,7 @@ async def get_model_path(model_name: str) -> str:
Raises:
FileNotFoundError: If model not found
"""
# Get api directory path (two levels up from core)
api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__)))

# Construct model directory path relative to api directory
model_dir = os.path.join(api_dir, settings.model_dir)

# Ensure model directory exists
model_dir = models_dir()
os.makedirs(model_dir, exist_ok=True)

# Search in model directory
Expand All @@ -140,13 +145,7 @@ async def get_voice_path(voice_name: str) -> str:
Raises:
FileNotFoundError: If voice not found
"""
# Get api directory path
api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__)))

# Construct voice directory path relative to api directory
voice_dir = os.path.join(api_dir, settings.voices_dir)

# Ensure voice directory exists
voice_dir = voices_dir()
os.makedirs(voice_dir, exist_ok=True)

voice_file = f"{voice_name}.pt"
Expand All @@ -164,13 +163,7 @@ async def list_voices() -> List[str]:
Returns:
List of voice names (without .pt extension)
"""
# Get api directory path
api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__)))

# Construct voice directory path relative to api directory
voice_dir = os.path.join(api_dir, settings.voices_dir)

# Ensure voice directory exists
voice_dir = voices_dir()
os.makedirs(voice_dir, exist_ok=True)

# Search in voice directory
Expand Down
45 changes: 45 additions & 0 deletions api/src/inference/decode_clip.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
"""Decodes a reference clip in a child process so a decoder crash cannot take the
server down. Upload on stdin; sample rate as 4 little-endian bytes then float32 mono
samples on stdout. Exits REFUSED with the reason on stderr for a refused or
undecodable clip.
"""

import io
import sys

import numpy as np
import soundfile as sf

MAX_REF_SECONDS = 30
MIN_SAMPLE_RATE = 8000
MAX_SAMPLE_RATE = 96000
MAX_CHANNELS = 2
REFUSED = 64


def decode(data: bytes) -> tuple[int, np.ndarray]:
"""First 30 s, mono, finite, clamped to [-1, 1]. ValueError on a clip over 2
channels or outside 8 to 96 kHz."""
with sf.SoundFile(io.BytesIO(data)) as clip:
sr = clip.samplerate
if clip.channels > MAX_CHANNELS or not MIN_SAMPLE_RATE <= sr <= MAX_SAMPLE_RATE:
raise ValueError(
f"reference must be mono or stereo at {MIN_SAMPLE_RATE // 1000} to "
f"{MAX_SAMPLE_RATE // 1000} kHz"
)
wav = clip.read(MAX_REF_SECONDS * sr, dtype="float32")
if wav.ndim > 1:
wav = wav.mean(-1)
return sr, np.nan_to_num(wav).clip(-1, 1)


if __name__ == "__main__":
try:
sr, wav = decode(sys.stdin.buffer.read())
except ValueError as e:
print(e, file=sys.stderr)
sys.exit(REFUSED)
except sf.LibsndfileError:
print("reference audio could not be decoded", file=sys.stderr)
sys.exit(REFUSED)
sys.stdout.buffer.write(sr.to_bytes(4, "little") + wav.tobytes())
91 changes: 91 additions & 0 deletions api/src/inference/inno_tuner.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
"""Inno clone tuner: reference clip in, stock-shaped Kokoro voice pack out.
Wraps the inno-kokoro package. docker/scripts/download_model.py bakes the pinned
weights next to the Kokoro model; load() runs at startup when ENABLE_INNO_TUNER is
set and any failure leaves available() False, so /dev/tune answers 503.
"""

import os
import subprocess
import sys
import tempfile
import threading
from typing import Optional

import numpy as np
import torch
from loguru import logger

from ..core import paths
from ..core.config import settings
from .decode_clip import REFUSED

MAX_UPLOAD_BYTES = 10 << 20
DECODER = os.path.join(os.path.dirname(__file__), "decode_clip.py")
DECODE_TIMEOUT = 5

_tuner = None
_lock = threading.Lock()


def weights_path() -> str:
return os.path.join(paths.models_dir(), "v1_0", "inno_tuner", "model.safetensors")


def load() -> None:
global _tuner
from inno_kokoro.enroll import Tuner

_tuner = Tuner(weights_path(), device=settings.get_device())
logger.info(f"Inno voice tuner v{_tuner.version} loaded on {_tuner.device}")


def available() -> bool:
return _tuner is not None


def reserve() -> bool:
"""Claim the tuner for one tune() call, False if held. tune() releases it."""
return _lock.acquire(blocking=False)


def decode(data: bytes) -> tuple[int, torch.Tensor]:
"""Run decode_clip.py in a child process. ValueError with the child's reason when
it refuses the clip; a crash or timeout is logged and reported as undecodable."""
try:
proc = subprocess.run(
[sys.executable, DECODER],
input=data,
capture_output=True,
timeout=DECODE_TIMEOUT,
)
except subprocess.TimeoutExpired:
logger.warning(f"Reference decode timed out after {DECODE_TIMEOUT}s")
raise ValueError("reference audio could not be decoded")
reason = proc.stderr.decode(errors="replace").strip()
if proc.returncode == REFUSED:
raise ValueError(reason)
if proc.returncode != 0:
logger.warning(f"Reference decode exited {proc.returncode}: {reason}")
raise ValueError("reference audio could not be decoded")
sr = int.from_bytes(proc.stdout[:4], "little")
return sr, torch.from_numpy(np.frombuffer(proc.stdout[4:], dtype=np.float32).copy())


def tune(data: bytes, head: bool = True, fmax: Optional[float] = None) -> str:
"""Decode, enroll, write the pack to the temp dir, return its path. Blocking,
call off the event loop after reserve(); releases the claim on return.
ValueError on a refused clip or one under 3 s."""
try:
if not available():
raise RuntimeError("inno voice tuner not available")
from inno_kokoro.enroll import enroll

sr, wav = decode(data)
pack, _ = enroll(wav, sr, _tuner, fmax=fmax, head=head)
fd, path = tempfile.mkstemp(prefix="a_tune_", suffix=".pt")
with os.fdopen(fd, "wb") as f:
torch.save(pack, f)
return path
finally:
_lock.release()
13 changes: 13 additions & 0 deletions api/src/inference/kokoro_v1.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
"""Clean Kokoro implementation with controlled resource management."""

import os
import tempfile
from typing import AsyncGenerator, Dict, Optional, Tuple, Union

import numpy as np
Expand Down Expand Up @@ -129,6 +130,18 @@ async def _get_voice_tensor(self, voice_path: str) -> torch.Tensor:
logger.debug(f"Cached voice tensor from {voice_path}")
return self._voice_cache[cache_key]

def forget_voice(self, voice_path: str) -> None:
"""Drop a pack's cached tensor and the temp copy generate() wrote for the pipeline."""
for key in [k for k in self._voice_cache if k.startswith(f"{voice_path}:")]:
del self._voice_cache[key]
temp_copy = os.path.join(
tempfile.gettempdir(), f"temp_voice_{os.path.basename(voice_path)}"
)
for pipeline in self._pipelines.values():
pipeline.voices.pop(temp_copy, None)
if os.path.exists(temp_copy):
os.remove(temp_copy)

async def load_model(self, path: str) -> None:
"""Load pre-baked model.
Expand Down
13 changes: 13 additions & 0 deletions api/src/inference/voice_manager.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,17 @@ def __init__(self):
# Strictly respect settings.use_gpu
self._device = settings.get_device()
self._voices: Dict[str, torch.Tensor] = {}
self._transient: Dict[str, str] = {}

def register_transient(self, voice_name: str, path: str) -> None:
"""Make a pack outside VOICES_DIR resolvable by name for the life of a request."""
self._transient[voice_name] = path

def forget_transient(self, voice_name: str) -> None:
self._transient.pop(voice_name, None)

def is_transient(self, voice_name: str) -> bool:
return voice_name in self._transient

async def get_voice_path(self, voice_name: str) -> str:
"""Get path to voice file.
Expand All @@ -32,6 +43,8 @@ async def get_voice_path(self, voice_name: str) -> str:
Raises:
RuntimeError: If voice not found
"""
if voice_name in self._transient:
return self._transient[voice_name]
return await paths.get_voice_path(voice_name)

async def load_voice(
Expand Down
Loading