diff --git a/.gitignore b/.gitignore index ed7330773..78d1c867d 100644 --- a/.gitignore +++ b/.gitignore @@ -108,3 +108,4 @@ test-results/ ROCM_TRIM.md Kokoro-FastAPI.code-workspace experiments/* +api/src/models/v1_0/inno_tuner/ diff --git a/CHANGELOG.md b/CHANGELOG.md index 262a09730..525d123e0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,10 @@ Per-PR attribution and contributor credits are published automatically on the co ### Added - `normalization_options.remove_emoji` drops emoji before synthesis instead of reading them by name, any language (#353). Off by default. - `normalization_options.caps_normalization` reads all-caps headers and names (`ARNE SAKNUSSEMM`, `TODO_LIST`) as words instead of letter by letter. On by default. Short acronyms (`FBI`, `US GDP`) are still spelled. +- `POST /dev/tune`: tune a voice from a short reference clip and speak with it in one request, via [inno-kokoro](https://github.com/remsky/inno-kokoro). Off by default, `ENABLE_INNO_TUNER=true` turns it on. See [docs/inno-tune.md](docs/inno-tune.md). + - `return_voice_pack=true` returns the tuned `.pt` instead of audio; `save_voice=` keeps it in `VOICES_DIR` as `_tuned`, behind `ALLOW_LOCAL_VOICE_SAVING`. +- Four tuned voices bundled with the server, `_inno` suffix: `af_amelia_inno`, `af_goodall_inno`, `am_price_inno`, `bm_atten_inno`. +- Web player: Tune tab, record or upload a clip and generate with it, download the pack, or save it to the server. ### Changed - Text normalization refactored towards multi-language support: `Normalizer` base class w/ neutral passes, `EnglishNormalizer` implements the rest, registry keyed by lang code. Adding a language is a subclass + table test, see CONTRIBUTING.md. diff --git a/README.md b/README.md index 5eb6a7f4a..395e51159 100644 --- a/README.md +++ b/README.md @@ -10,16 +10,14 @@ [![Misaki](https://img.shields.io/badge/misaki-0.9.4-B8860B)](https://github.com/hexgrad/misaki) [![Tested at Model Commit](https://img.shields.io/badge/model-1.0::41e5892-blue)](https://huggingface.co/hexgrad/Kokoro-82M/commit/41e5892b9d8b43e56fc560f892312a328a410973) -[![Try on Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Try%20on-Spaces-blue)](https://huggingface.co/spaces/Remsky/FastKoko) [![Downloads](https://img.shields.io/badge/downloads-2.5M%2B-2496ED?logo=docker&logoColor=white)](https://github.com/remsky?tab=packages&repo_name=Kokoro-FastAPI) +[![Try on Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Try%20on-Spaces-blue)](https://huggingface.co/spaces/Remsky/FastKoko) [![Downloads](https://img.shields.io/badge/downloads-2.6M%2B-2496ED?logo=docker&logoColor=white)](https://github.com/remsky?tab=packages&repo_name=Kokoro-FastAPI) Dockerized FastAPI wrapper for [Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) text-to-speech model. Generate hours of high quality speech in minutes. -> [!NOTE] -> Looking for custom voices? Try the [Inno Clone-Tuner](https://github.com/remsky/inno-kokoro) - - OpenAI-compatible Speech endpoint, multi-language support - English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, Mandarin Chinese +- Custom voicepack generation via [Inno Clone-Tuner](https://github.com/remsky/inno-kokoro) - Optional integrated WebUI; read-along long-generation - Inline multi-speaker generation & voice mixing + aliasing weighted combinations, SSML support - Per-word, or per-chunk timestamped caption generation @@ -38,11 +36,11 @@ Dockerized FastAPI wrapper for [Kokoro-82M](https://huggingface.co/hexgrad/Kokor Community projects that use, recommend, or enable Kokoro-FastAPI as a backend: -- Home Assistant: [wyoming_openai](https://github.com/roryeckel/wyoming_openai), [openai_tts](https://github.com/sfortis/openai_tts), [Kokoro-TTS](https://github.com/beecho01/Kokoro-TTS) -- App stores and templates: [Umbrel](https://github.com/getumbrel/umbrel-apps/tree/master/kokoro), [Unraid Community Apps](https://github.com/nwithan8/unraid_templates), [GPUStack](https://github.com/gpustack/gpustack), [jetson-containers](https://github.com/dusty-nv/jetson-containers/tree/master/packages/speech/kokoro-tts) -- Readers and audiobooks: [openreader](https://github.com/richardr1126/openreader), [epub_to_audiobook](https://github.com/p0n1/epub_to_audiobook), [audiobook-creator](https://github.com/prakharsr/audiobook-creator), [Zotero-TTS](https://github.com/xujialiu/Zotero-TTS) -- Assistants and agents: [xiaozhi-esp32-server](https://github.com/xinnan-tech/xiaozhi-esp32-server), [call-me](https://github.com/ZeframLou/call-me), [agent-cli](https://github.com/basnijholt/agent-cli), [voice-chat-ai](https://github.com/bigsk1/voice-chat-ai) -- Browser: [kokoro-extension](https://github.com/Fooftilly/kokoro-extension), [customtts](https://github.com/BassGaming/customtts) +- Home Assistant: [wyoming_openai](https://github.com/roryeckel/wyoming_openai), [openai_tts](https://github.com/sfortis/openai_tts), [Kokoro-TTS](https://github.com/beecho01/Kokoro-TTS) +- App stores and templates: [Umbrel](https://github.com/getumbrel/umbrel-apps/tree/master/kokoro), [Unraid Apps](https://github.com/nwithan8/unraid_templates), [GPUStack](https://github.com/gpustack/gpustack), [jetson-containers](https://github.com/dusty-nv/jetson-containers/tree/master/packages/speech/kokoro-tts) +- Readers and audiobooks: [openreader](https://github.com/richardr1126/openreader), [epub_to_audiobook](https://github.com/p0n1/epub_to_audiobook), [audiobook-creator](https://github.com/prakharsr/audiobook-creator), [Zotero-TTS](https://github.com/xujialiu/Zotero-TTS) +- Assistants and agents: [xiaozhi-esp32-server](https://github.com/xinnan-tech/xiaozhi-esp32-server), [call-me](https://github.com/ZeframLou/call-me), [agent-cli](https://github.com/basnijholt/agent-cli), [voice-chat-ai](https://github.com/bigsk1/voice-chat-ai) +- Browser: [kokoro-extension](https://github.com/Fooftilly/kokoro-extension), [customtts](https://github.com/BassGaming/customtts) ## Get Started @@ -410,6 +408,27 @@ curl -X POST http://localhost:8880/v1/audio/speech \ +
+Voice Tuning (reference clip) 🧪 + +`POST /dev/tune` takes a 3 to 30 s clip of one English speaker and speaks with a voice tuned toward it, via [inno-kokoro](https://github.com/remsky/inno-kokoro). A tuner, not a cloner: expect the same neighbourhood, not a match. Only tune voices you have permission to use. + +```bash +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F 'request={"input":"Hello there."}' -o out.mp3 +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F return_voice_pack=true -o ref.pt +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F save_voice=am_ref # saves am_ref_tuned, needs ALLOW_LOCAL_VOICE_SAVING=true +``` + +- `request` is the `/v1/audio/speech` body minus `voice`; the pack only exists for the length of the response unless `save_voice` keeps it in `VOICES_DIR` +- Names follow the existing language prefix pattern (`am_`, `bf_`, `ax_`, etc.) and saved voices get a `_tuned` suffix. Bundled tuned voices end in `_inno` +- Knobs: `prosody_head` (default on), `fmax` pitch ceiling in Hz (default auto) +- Off by default: + - `ENABLE_INNO_TUNER=true` enables with a web player tab + - `ALLOW_LOCAL_VOICE_SAVING=true` allows saving them to the live server. +- Full reference in [docs/inno-tune.md](docs/inno-tune.md) + +
+
Multi-Speaker / Dialogue @@ -493,7 +512,7 @@ The city of [Worcester](/wˈʊstər/) is easy. [pause:1s] See?
-SSML Input (experimental) +SSML Input 🧪 Send `ssml: true` with `allow_voice_tags: true` on `/v1/audio/speech` or `/dev/captioned_speech` to translate and speak in one call. Both flags are needed, since the translation emits `[voice:]` and `[rate:]` spans that would otherwise be read aloud; `ssml` without them is a 400. diff --git a/api/src/core/config.py b/api/src/core/config.py index 0524093ef..a46bdd171 100644 --- a/api/src/core/config.py +++ b/api/src/core/config.py @@ -40,6 +40,7 @@ class Settings(BaseSettings): False # Whether to allow saving combined voices locally ) allow_dev_unload: bool = False # Whether to expose /dev/model, POST /dev/unload, and POST /dev/reload + enable_inno_tuner: bool = False # Whether to expose POST /dev/tune model_auto_unload_timeout_seconds: float = ( 0.0 # Idle seconds before unloading; 0 disables auto-unload ) diff --git a/api/src/core/paths.py b/api/src/core/paths.py index b2e98d4d1..c9308e104 100644 --- a/api/src/core/paths.py +++ b/api/src/core/paths.py @@ -36,6 +36,17 @@ } +_API_DIR = os.path.dirname(os.path.dirname(os.path.dirname(__file__))) + + +def models_dir() -> str: + return os.path.join(_API_DIR, settings.model_dir) + + +def voices_dir() -> str: + return os.path.join(_API_DIR, settings.voices_dir) + + async def _find_file( filename: str, search_paths: List[str], @@ -112,13 +123,7 @@ async def get_model_path(model_name: str) -> str: Raises: FileNotFoundError: If model not found """ - # Get api directory path (two levels up from core) - api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__))) - - # Construct model directory path relative to api directory - model_dir = os.path.join(api_dir, settings.model_dir) - - # Ensure model directory exists + model_dir = models_dir() os.makedirs(model_dir, exist_ok=True) # Search in model directory @@ -140,13 +145,7 @@ async def get_voice_path(voice_name: str) -> str: Raises: FileNotFoundError: If voice not found """ - # Get api directory path - api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__))) - - # Construct voice directory path relative to api directory - voice_dir = os.path.join(api_dir, settings.voices_dir) - - # Ensure voice directory exists + voice_dir = voices_dir() os.makedirs(voice_dir, exist_ok=True) voice_file = f"{voice_name}.pt" @@ -164,13 +163,7 @@ async def list_voices() -> List[str]: Returns: List of voice names (without .pt extension) """ - # Get api directory path - api_dir = os.path.dirname(os.path.dirname(os.path.dirname(__file__))) - - # Construct voice directory path relative to api directory - voice_dir = os.path.join(api_dir, settings.voices_dir) - - # Ensure voice directory exists + voice_dir = voices_dir() os.makedirs(voice_dir, exist_ok=True) # Search in voice directory diff --git a/api/src/inference/decode_clip.py b/api/src/inference/decode_clip.py new file mode 100644 index 000000000..d09fe791e --- /dev/null +++ b/api/src/inference/decode_clip.py @@ -0,0 +1,45 @@ +"""Decodes a reference clip in a child process so a decoder crash cannot take the +server down. Upload on stdin; sample rate as 4 little-endian bytes then float32 mono +samples on stdout. Exits REFUSED with the reason on stderr for a refused or +undecodable clip. +""" + +import io +import sys + +import numpy as np +import soundfile as sf + +MAX_REF_SECONDS = 30 +MIN_SAMPLE_RATE = 8000 +MAX_SAMPLE_RATE = 96000 +MAX_CHANNELS = 2 +REFUSED = 64 + + +def decode(data: bytes) -> tuple[int, np.ndarray]: + """First 30 s, mono, finite, clamped to [-1, 1]. ValueError on a clip over 2 + channels or outside 8 to 96 kHz.""" + with sf.SoundFile(io.BytesIO(data)) as clip: + sr = clip.samplerate + if clip.channels > MAX_CHANNELS or not MIN_SAMPLE_RATE <= sr <= MAX_SAMPLE_RATE: + raise ValueError( + f"reference must be mono or stereo at {MIN_SAMPLE_RATE // 1000} to " + f"{MAX_SAMPLE_RATE // 1000} kHz" + ) + wav = clip.read(MAX_REF_SECONDS * sr, dtype="float32") + if wav.ndim > 1: + wav = wav.mean(-1) + return sr, np.nan_to_num(wav).clip(-1, 1) + + +if __name__ == "__main__": + try: + sr, wav = decode(sys.stdin.buffer.read()) + except ValueError as e: + print(e, file=sys.stderr) + sys.exit(REFUSED) + except sf.LibsndfileError: + print("reference audio could not be decoded", file=sys.stderr) + sys.exit(REFUSED) + sys.stdout.buffer.write(sr.to_bytes(4, "little") + wav.tobytes()) diff --git a/api/src/inference/inno_tuner.py b/api/src/inference/inno_tuner.py new file mode 100644 index 000000000..3294e9211 --- /dev/null +++ b/api/src/inference/inno_tuner.py @@ -0,0 +1,91 @@ +"""Inno clone tuner: reference clip in, stock-shaped Kokoro voice pack out. + +Wraps the inno-kokoro package. docker/scripts/download_model.py bakes the pinned +weights next to the Kokoro model; load() runs at startup when ENABLE_INNO_TUNER is +set and any failure leaves available() False, so /dev/tune answers 503. +""" + +import os +import subprocess +import sys +import tempfile +import threading +from typing import Optional + +import numpy as np +import torch +from loguru import logger + +from ..core import paths +from ..core.config import settings +from .decode_clip import REFUSED + +MAX_UPLOAD_BYTES = 10 << 20 +DECODER = os.path.join(os.path.dirname(__file__), "decode_clip.py") +DECODE_TIMEOUT = 5 + +_tuner = None +_lock = threading.Lock() + + +def weights_path() -> str: + return os.path.join(paths.models_dir(), "v1_0", "inno_tuner", "model.safetensors") + + +def load() -> None: + global _tuner + from inno_kokoro.enroll import Tuner + + _tuner = Tuner(weights_path(), device=settings.get_device()) + logger.info(f"Inno voice tuner v{_tuner.version} loaded on {_tuner.device}") + + +def available() -> bool: + return _tuner is not None + + +def reserve() -> bool: + """Claim the tuner for one tune() call, False if held. tune() releases it.""" + return _lock.acquire(blocking=False) + + +def decode(data: bytes) -> tuple[int, torch.Tensor]: + """Run decode_clip.py in a child process. ValueError with the child's reason when + it refuses the clip; a crash or timeout is logged and reported as undecodable.""" + try: + proc = subprocess.run( + [sys.executable, DECODER], + input=data, + capture_output=True, + timeout=DECODE_TIMEOUT, + ) + except subprocess.TimeoutExpired: + logger.warning(f"Reference decode timed out after {DECODE_TIMEOUT}s") + raise ValueError("reference audio could not be decoded") + reason = proc.stderr.decode(errors="replace").strip() + if proc.returncode == REFUSED: + raise ValueError(reason) + if proc.returncode != 0: + logger.warning(f"Reference decode exited {proc.returncode}: {reason}") + raise ValueError("reference audio could not be decoded") + sr = int.from_bytes(proc.stdout[:4], "little") + return sr, torch.from_numpy(np.frombuffer(proc.stdout[4:], dtype=np.float32).copy()) + + +def tune(data: bytes, head: bool = True, fmax: Optional[float] = None) -> str: + """Decode, enroll, write the pack to the temp dir, return its path. Blocking, + call off the event loop after reserve(); releases the claim on return. + ValueError on a refused clip or one under 3 s.""" + try: + if not available(): + raise RuntimeError("inno voice tuner not available") + from inno_kokoro.enroll import enroll + + sr, wav = decode(data) + pack, _ = enroll(wav, sr, _tuner, fmax=fmax, head=head) + fd, path = tempfile.mkstemp(prefix="a_tune_", suffix=".pt") + with os.fdopen(fd, "wb") as f: + torch.save(pack, f) + return path + finally: + _lock.release() diff --git a/api/src/inference/kokoro_v1.py b/api/src/inference/kokoro_v1.py index a1a0a483b..c559fc990 100644 --- a/api/src/inference/kokoro_v1.py +++ b/api/src/inference/kokoro_v1.py @@ -1,6 +1,7 @@ """Clean Kokoro implementation with controlled resource management.""" import os +import tempfile from typing import AsyncGenerator, Dict, Optional, Tuple, Union import numpy as np @@ -129,6 +130,18 @@ async def _get_voice_tensor(self, voice_path: str) -> torch.Tensor: logger.debug(f"Cached voice tensor from {voice_path}") return self._voice_cache[cache_key] + def forget_voice(self, voice_path: str) -> None: + """Drop a pack's cached tensor and the temp copy generate() wrote for the pipeline.""" + for key in [k for k in self._voice_cache if k.startswith(f"{voice_path}:")]: + del self._voice_cache[key] + temp_copy = os.path.join( + tempfile.gettempdir(), f"temp_voice_{os.path.basename(voice_path)}" + ) + for pipeline in self._pipelines.values(): + pipeline.voices.pop(temp_copy, None) + if os.path.exists(temp_copy): + os.remove(temp_copy) + async def load_model(self, path: str) -> None: """Load pre-baked model. diff --git a/api/src/inference/voice_manager.py b/api/src/inference/voice_manager.py index 35ee62266..a700073e9 100644 --- a/api/src/inference/voice_manager.py +++ b/api/src/inference/voice_manager.py @@ -19,6 +19,17 @@ def __init__(self): # Strictly respect settings.use_gpu self._device = settings.get_device() self._voices: Dict[str, torch.Tensor] = {} + self._transient: Dict[str, str] = {} + + def register_transient(self, voice_name: str, path: str) -> None: + """Make a pack outside VOICES_DIR resolvable by name for the life of a request.""" + self._transient[voice_name] = path + + def forget_transient(self, voice_name: str) -> None: + self._transient.pop(voice_name, None) + + def is_transient(self, voice_name: str) -> bool: + return voice_name in self._transient async def get_voice_path(self, voice_name: str) -> str: """Get path to voice file. @@ -32,6 +43,8 @@ async def get_voice_path(self, voice_name: str) -> str: Raises: RuntimeError: If voice not found """ + if voice_name in self._transient: + return self._transient[voice_name] return await paths.get_voice_path(voice_name) async def load_voice( diff --git a/api/src/main.py b/api/src/main.py index a15a1a88a..898af5ba4 100644 --- a/api/src/main.py +++ b/api/src/main.py @@ -2,6 +2,7 @@ FastAPI OpenAI Compatible API """ +import asyncio import os import sys from contextlib import asynccontextmanager @@ -17,6 +18,7 @@ from .routers.development import router as dev_router from .routers.openai_compatible import router as openai_router from .routers.ssml import router as ssml_router +from .routers.tune import router as tune_router from .routers.web_player import router as web_router @@ -79,6 +81,14 @@ async def lifespan(app: FastAPI): logger.error(f"Failed to initialize model: {e}") raise + if settings.enable_inno_tuner: + from .inference import inno_tuner + + try: + await asyncio.to_thread(inno_tuner.load) + except Exception as e: + logger.error(f"Inno voice tuner not loaded, /dev/tune will answer 503: {e}") + boundary = "░" * 2 * 12 startup_msg = f""" @@ -101,6 +111,10 @@ async def lifespan(app: FastAPI): else: startup_msg += "\nRunning on CPU" startup_msg += f"\n{voicepack_count} voice packs loaded" + if settings.enable_inno_tuner: + startup_msg += "\nInno voice tuner: " + ( + "ready at /dev/tune" if inno_tuner.available() else "not available" + ) # Add web player info if enabled if settings.enable_web_player: @@ -140,6 +154,7 @@ async def lifespan(app: FastAPI): app.include_router(openai_router, prefix="/v1") app.include_router(dev_router) # Development endpoints app.include_router(ssml_router) # SSML translation and capabilities +app.include_router(tune_router) # /dev/tune, 403 unless enabled app.include_router(debug_router) # Debug endpoints (403 unless enabled) if settings.enable_web_player: app.include_router(web_router, prefix="/web") # Web player static files diff --git a/api/src/routers/openai_compatible.py b/api/src/routers/openai_compatible.py index c0c515739..41216846c 100644 --- a/api/src/routers/openai_compatible.py +++ b/api/src/routers/openai_compatible.py @@ -13,6 +13,7 @@ from ..core.config import settings from ..inference.base import AudioChunk +from ..inference.voice_manager import get_manager as get_voice_manager from ..services.streaming_audio_writer import StreamingAudioWriter from ..services.text_processing.text_processor import ( VOICE_TAG_PATTERN, @@ -160,6 +161,7 @@ async def process_and_validate_voices( if available_voices is None: available_voices = await tts_service.list_voices() + voice_manager = await get_voice_manager() for voice_index in range(0, len(voices), 2): token = voices[voice_index] @@ -179,7 +181,7 @@ async def process_and_validate_voices( weight = weight.strip() name = _openai_mappings["voices"].get(name, name) - if name not in available_voices: + if name not in available_voices and not voice_manager.is_transient(name): raise ValueError( f"Voice '{name}' not found. Available voices: {', '.join(sorted(available_voices))}" ) diff --git a/api/src/routers/tune.py b/api/src/routers/tune.py new file mode 100644 index 000000000..cd4c49df3 --- /dev/null +++ b/api/src/routers/tune.py @@ -0,0 +1,199 @@ +import asyncio +import json +import os +import re +import shutil +from typing import AsyncIterator, Optional + +from fastapi import APIRouter, Depends, File, Form, HTTPException, Request, UploadFile +from fastapi.responses import FileResponse, JSONResponse, StreamingResponse +from loguru import logger +from pydantic import ValidationError +from starlette.background import BackgroundTask + +from ..core import paths +from ..core.config import settings +from ..inference import inno_tuner +from ..inference.voice_manager import get_manager as get_voice_manager +from ..services.tts_service import TTSService +from ..structures import OpenAISpeechRequest +from .openai_compatible import create_speech, get_tts_service + +router = APIRouter(tags=["voice tuning"]) + +SAVE_NAME = re.compile(r"^[ab][a-z]?_[a-z0-9]+(_[a-z0-9]+)*$") +SAVE_SUFFIX = "_tuned" +SAVE_NAME_RULE = ( + "save_voice must start with a (US) or b (UK English), an optional second letter, then _ " + "and lowercase letters, digits, single underscores, e.g. ax_me" +) + + +def _bad_request(message: str) -> HTTPException: + return HTTPException( + status_code=400, + detail={ + "error": "validation_error", + "message": message, + "type": "invalid_request_error", + }, + ) + + +def _copy_new(src: str, dest: str) -> None: + with open(src, "rb") as f, open(dest, "xb") as out: + shutil.copyfileobj(f, out) + + +async def _discard(tts_service: TTSService, voice_name: str, pack_path: str) -> None: + (await get_voice_manager()).forget_transient(voice_name) + try: + tts_service.model_manager.get_backend().forget_voice(pack_path) + except Exception: + pass + try: + os.remove(pack_path) + except OSError: + pass + + +async def _discard_after( + body: AsyncIterator, tts_service: TTSService, voice_name: str, pack_path: str +): + try: + async for chunk in body: + yield chunk + finally: + await _discard(tts_service, voice_name, pack_path) + + +@router.post("/dev/tune") +async def tune_speech( + client_request: Request, + audio: UploadFile = File(..., description="Reference clip, 3 to 30 s, one speaker"), + request: Optional[str] = Form( + None, description="JSON body as for /v1/audio/speech, minus voice" + ), + prosody_head: bool = Form(True), + fmax: Optional[float] = Form(None, ge=60, le=1000), + return_voice_pack: bool = Form(False), + save_voice: str = Form(""), + tts_service: TTSService = Depends(get_tts_service), +): + """Tune a voice from the reference clip, then speak `request` with it. + + return_voice_pack=true returns the [510, 1, 256] .pt instead. save_voice= + keeps it as VOICES_DIR/_tuned.pt and needs ALLOW_LOCAL_VOICE_SAVING. + Otherwise the pack lives in the temp dir for the length of the response only. + """ + if not settings.enable_inno_tuner: + raise HTTPException( + status_code=403, + detail={ + "error": "permission_denied", + "message": "Inno voice tuner is disabled", + "type": "permission_error", + }, + ) + if save_voice and not settings.allow_local_voice_saving: + raise HTTPException( + status_code=403, + detail={ + "error": "permission_denied", + "message": "Local voice saving is disabled", + "type": "permission_error", + }, + ) + if not inno_tuner.available(): + raise HTTPException( + status_code=503, + detail={"error": "unavailable", "message": "Inno voice tuner not loaded"}, + ) + if not (request or return_voice_pack or save_voice): + raise _bad_request("request, return_voice_pack, or save_voice is required") + + save_name = save_voice.strip().lower() + if save_name: + if not SAVE_NAME.match(save_name): + raise _bad_request(SAVE_NAME_RULE) + if not save_name.endswith(SAVE_SUFFIX): + save_name += SAVE_SUFFIX + + speech = None + if request: + try: + speech = OpenAISpeechRequest(**{**json.loads(request), "voice": "a_tune"}) + except (TypeError, ValueError, ValidationError) as e: + raise _bad_request(f"request: {e}") + speech.allow_voice_tags = False + speech.ssml = False + + data = await audio.read(inno_tuner.MAX_UPLOAD_BYTES + 1) + if len(data) > inno_tuner.MAX_UPLOAD_BYTES: + raise HTTPException( + status_code=413, + detail={ + "error": "too_large", + "message": f"reference exceeds {inno_tuner.MAX_UPLOAD_BYTES >> 20} MB", + }, + ) + if not inno_tuner.reserve(): + raise HTTPException( + status_code=503, + detail={ + "error": "busy", + "message": "Inno voice tuner is busy, retry shortly", + }, + ) + try: + pack_path = await asyncio.shield( + asyncio.to_thread(inno_tuner.tune, data, prosody_head, fmax) + ) + except ValueError as e: + raise _bad_request(str(e)) + + voice_name = os.path.splitext(os.path.basename(pack_path))[0] + (await get_voice_manager()).register_transient(voice_name, pack_path) + cleanup = BackgroundTask(_discard, tts_service, voice_name, pack_path) + + if save_name: + dest = os.path.join(paths.voices_dir(), f"{save_name}.pt") + try: + await asyncio.to_thread(_copy_new, pack_path, dest) + except FileExistsError: + await _discard(tts_service, voice_name, pack_path) + raise HTTPException( + status_code=409, + detail={"error": "conflict", "message": f"{save_name} already exists"}, + ) + logger.info(f"Saved tuned voice {save_name}") + + if return_voice_pack: + return FileResponse( + pack_path, + media_type="application/octet-stream", + filename=f"{save_name or 'a_tune'}.pt", + headers={"Cache-Control": "no-cache"}, + background=cleanup, + ) + if speech is None: + await _discard(tts_service, voice_name, pack_path) + return JSONResponse({"voice": save_name}) + + speech.voice = voice_name + try: + response = await create_speech( + request=speech, + client_request=client_request, + x_raw_response=None, + ) + except Exception: + await _discard(tts_service, voice_name, pack_path) + raise + if isinstance(response, StreamingResponse): + response.body_iterator = _discard_after( + response.body_iterator, tts_service, voice_name, pack_path + ) + else: + response.background = cleanup + return response diff --git a/api/src/routers/web_player.py b/api/src/routers/web_player.py index 99265999c..782472606 100644 --- a/api/src/routers/web_player.py +++ b/api/src/routers/web_player.py @@ -8,6 +8,7 @@ from ..core.config import settings from ..core.paths import get_content_type, get_web_file_path, read_bytes +from ..inference import inno_tuner router = APIRouter( tags=["Web Player"], @@ -26,6 +27,8 @@ async def get_web_config(): return { "root_path": root_path, "version": settings.api_version, + "tuner": inno_tuner.available(), + "voice_saving": settings.allow_local_voice_saving, } diff --git a/api/src/voices/v1_0/af_amelia_inno.pt b/api/src/voices/v1_0/af_amelia_inno.pt new file mode 100644 index 000000000..2233a688c Binary files /dev/null and b/api/src/voices/v1_0/af_amelia_inno.pt differ diff --git a/api/src/voices/v1_0/af_goodall_inno.pt b/api/src/voices/v1_0/af_goodall_inno.pt new file mode 100644 index 000000000..f1d480313 Binary files /dev/null and b/api/src/voices/v1_0/af_goodall_inno.pt differ diff --git a/api/src/voices/v1_0/am_price_inno.pt b/api/src/voices/v1_0/am_price_inno.pt new file mode 100644 index 000000000..69e4bc5df Binary files /dev/null and b/api/src/voices/v1_0/am_price_inno.pt differ diff --git a/api/src/voices/v1_0/bm_atten_inno.pt b/api/src/voices/v1_0/bm_atten_inno.pt new file mode 100644 index 000000000..6e31ebf71 Binary files /dev/null and b/api/src/voices/v1_0/bm_atten_inno.pt differ diff --git a/api/tests/test_download_model.py b/api/tests/test_download_model.py new file mode 100644 index 000000000..b33113737 --- /dev/null +++ b/api/tests/test_download_model.py @@ -0,0 +1,51 @@ +import importlib.util +import sys +from pathlib import Path +from unittest.mock import MagicMock + +import huggingface_hub.utils +import pytest +from huggingface_hub import constants + +spec = importlib.util.spec_from_file_location( + "download_model", Path(__file__).parents[2] / "docker/scripts/download_model.py" +) +download_model = importlib.util.module_from_spec(spec) +spec.loader.exec_module(download_model) + + +def test_tuner_fetch_failure_is_not_fatal(monkeypatch, tmp_path): + def boom(_): + raise ConnectionError("hub unreachable") + + monkeypatch.setattr(sys, "argv", ["download_model.py", "--output", str(tmp_path)]) + monkeypatch.setattr(download_model, "download_model", lambda _: None) + monkeypatch.setattr(download_model, "download_tuner", boom) + download_model.main() + + monkeypatch.setattr(download_model, "download_model", boom) + with pytest.raises(ConnectionError): + download_model.main() + + +def test_update_check_honours_offline_and_telemetry_opt_outs(monkeypatch): + urls = [] + + class Session: + def get(self, url, timeout): + urls.append(url) + return MagicMock(json=lambda: {"version": "9.9.9"}) + + monkeypatch.setattr(huggingface_hub.utils, "get_session", Session) + monkeypatch.setattr(constants, "HF_HUB_OFFLINE", False) + monkeypatch.setattr(constants, "HF_HUB_DISABLE_TELEMETRY", False) + download_model.check_tuner_update() + assert urls == [ + "https://huggingface.co/remsky/kokoro-inno-clone-tuner/resolve/main/config.json" + ] + + for flag in ("HF_HUB_OFFLINE", "HF_HUB_DISABLE_TELEMETRY"): + monkeypatch.setattr(constants, flag, True) + download_model.check_tuner_update() + monkeypatch.setattr(constants, flag, False) + assert len(urls) == 1 diff --git a/api/tests/test_inno_tuner.py b/api/tests/test_inno_tuner.py new file mode 100644 index 000000000..4d64dfb5e --- /dev/null +++ b/api/tests/test_inno_tuner.py @@ -0,0 +1,459 @@ +import asyncio +import glob +import io +import json +import os +import tempfile +import threading +import time +from unittest.mock import AsyncMock, MagicMock + +import httpx +import numpy as np +import pytest +import soundfile as sf +import torch +from fastapi.testclient import TestClient + +from api.src.core import paths +from api.src.core.config import settings +from api.src.inference import inno_tuner +from api.src.inference.base import AudioChunk +from api.src.inference.voice_manager import VoiceManager +from api.src.main import app +from api.src.routers import openai_compatible +from api.src.routers.openai_compatible import get_tts_service + +client = TestClient(app) + + +def wav_bytes(seconds=4.0, sr=24000, channels=1, fmt="WAV"): + buf = io.BytesIO() + frames = np.zeros((int(seconds * sr), channels), dtype=np.float32) + sf.write(buf, frames, sr, format=fmt) + return buf.getvalue() + + +def post(data=None, seconds=4.0, raw=None): + return client.post( + "/dev/tune", + files={"audio": ("ref.wav", raw or wav_bytes(seconds), "audio/wav")}, + data=data or {}, + ) + + +def transient(): + return VoiceManager._instance._transient + + +def temp_packs(): + return set(glob.glob(os.path.join(tempfile.gettempdir(), "a_tune_*.pt"))) + + +@pytest.fixture +def fake_tuner(monkeypatch): + monkeypatch.setattr(settings, "enable_inno_tuner", True) + monkeypatch.setattr(settings, "allow_local_voice_saving", True) + monkeypatch.setattr(inno_tuner, "_tuner", object()) + calls = [] + + def enroll(wav, sr, tuner, fmax=None, head=True): + if len(wav) / sr < 3: + raise ValueError("reference is too short; need at least 3 s") + calls.append((round(len(wav) / sr, 3), head, fmax)) + return torch.zeros(510, 1, 256), {"af_heart": 1.0} + + import inno_kokoro.enroll + + monkeypatch.setattr(inno_kokoro.enroll, "enroll", enroll) + return calls + + +@pytest.fixture +def service(fake_tuner, monkeypatch): + seen = [] + + async def gen(**kwargs): + seen.append(kwargs) + yield AudioChunk(np.zeros(1, dtype=np.int16), output=b"abc") + + svc = AsyncMock() + svc.generate_audio_stream = gen + monkeypatch.setattr(VoiceManager, "_instance", VoiceManager()) + svc.model_manager = MagicMock() + svc.model_manager.get_backend.return_value.forget_voice = MagicMock() + app.dependency_overrides[get_tts_service] = lambda: svc + monkeypatch.setattr( + openai_compatible, "get_tts_service", AsyncMock(return_value=svc) + ) + yield svc, seen + app.dependency_overrides.pop(get_tts_service, None) + + +def test_503_when_tuner_missing(monkeypatch): + monkeypatch.setattr(settings, "enable_inno_tuner", True) + monkeypatch.setattr(inno_tuner, "_tuner", None) + assert post({"return_voice_pack": "true"}).status_code == 503 + + +def test_403_when_disabled(monkeypatch): + monkeypatch.setattr(settings, "enable_inno_tuner", False) + assert post({"return_voice_pack": "true"}).status_code == 403 + + +def test_return_voice_pack_and_cleanup(service, fake_tuner): + before = temp_packs() + r = post({"return_voice_pack": "true", "prosody_head": "false"}) + assert r.status_code == 200 + assert fake_tuner[-1][1] is False + assert r.headers["content-disposition"].endswith('filename="a_tune.pt"') + assert torch.load(io.BytesIO(r.content), weights_only=True).shape == (510, 1, 256) + assert temp_packs() == before + assert transient() == {} + + +def test_403_when_voice_saving_disabled(service, monkeypatch): + monkeypatch.setattr(settings, "allow_local_voice_saving", False) + assert post({"save_voice": "am_me"}).status_code == 403 + assert post({"return_voice_pack": "true"}).status_code == 200 + + +def test_save_voice_writes_voices_dir(service, monkeypatch, tmp_path): + monkeypatch.setattr(settings, "voices_dir", str(tmp_path)) + before = temp_packs() + r = post({"save_voice": "AM_Me"}) + assert r.status_code == 200 + assert r.json() == {"voice": "am_me_tuned"} + assert torch.load(tmp_path / "am_me_tuned.pt", weights_only=True).shape == ( + 510, + 1, + 256, + ) + assert post({"save_voice": "am_me"}).status_code == 409 + assert post({"save_voice": "am_me_tuned"}).status_code == 409 + assert post({"save_voice": "B_Me"}).json() == {"voice": "b_me_tuned"} + assert post({"save_voice": "ax_me"}).json() == {"voice": "ax_me_tuned"} + for bad in ( + "me", + "am-me", + "am_me!", + "am__me", + "am_me_", + "xf_me", + "am_", + "a_", + "ax_", + "a1_me", + ): + assert post({"save_voice": bad}).status_code == 400, bad + assert temp_packs() == before + assert asyncio.run(paths.list_voices()) == [ + "am_me_tuned", + "ax_me_tuned", + "b_me_tuned", + ] + assert asyncio.run(VoiceManager().get_voice_path("am_me_tuned")) == str( + tmp_path / "am_me_tuned.pt" + ) + + +def test_knobs_and_size_cap(service, fake_tuner, monkeypatch): + assert post({"return_voice_pack": "true", "fmax": "300"}).status_code == 200 + assert fake_tuner[-1] == (4.0, True, 300.0) + assert post({"return_voice_pack": "true", "fmax": "5"}).status_code == 422 + monkeypatch.setattr(inno_tuner, "MAX_UPLOAD_BYTES", 1000) + assert post({"return_voice_pack": "true"}).status_code == 413 + + +def test_decode_is_bounded(service, fake_tuner): + long_flac = wav_bytes(seconds=120, fmt="FLAC") + assert len(long_flac) < inno_tuner.MAX_UPLOAD_BYTES + assert post({"return_voice_pack": "true"}, raw=long_flac).status_code == 200 + assert fake_tuner[-1][0] == 30.0 + for raw in (wav_bytes(channels=4), wav_bytes(sr=192000), wav_bytes(sr=1)): + assert post({"return_voice_pack": "true"}, raw=raw).status_code == 400 + + +def test_400_on_bad_input(service): + assert post({}).status_code == 400 + assert post({"return_voice_pack": "true"}, raw=b"not audio").status_code == 400 + assert post({"return_voice_pack": "true"}, seconds=2).status_code == 400 + assert post({"request": "{not json"}).status_code == 400 + assert ( + post({"request": json.dumps({"input": "hi", "speed": 99})}).status_code == 400 + ) + + +def test_speech_uses_transient_voice_then_forgets_it(service): + svc, seen = service + before = temp_packs() + r = post( + { + "request": json.dumps( + {"input": "hello", "response_format": "wav", "allow_voice_tags": True} + ) + } + ) + assert r.status_code == 200 + assert r.content == b"abc" + assert r.headers["content-type"] == "audio/wav" + voice = seen[0]["voice"] + assert voice.startswith("a_tune_") + assert seen[0]["allow_voice_tags"] is False + assert temp_packs() == before + assert transient() == {} + svc.model_manager.get_backend.return_value.forget_voice.assert_called_once() + assert svc.model_manager.get_backend.return_value.forget_voice.call_args[0][ + 0 + ].endswith(f"{voice}.pt") + + +def test_whole_response_with_download_link(service, monkeypatch, tmp_path): + svc, _ = service + monkeypatch.setattr(settings, "temp_file_dir", str(tmp_path)) + svc.generate_audio.return_value = AudioChunk( + np.zeros(1, dtype=np.int16), output=b"abc" + ) + before = temp_packs() + r = post( + { + "request": json.dumps( + {"input": "hello", "stream": False, "return_download_link": True} + ) + } + ) + assert r.status_code == 200 + assert r.content == b"abc" + assert svc.generate_audio.call_args.kwargs["voice"].startswith("a_tune_") + download = tmp_path / os.path.basename(r.headers["X-Download-Path"]) + assert download.read_bytes() == b"abc" + assert temp_packs() == before + assert transient() == {} + svc.model_manager.get_backend.return_value.forget_voice.assert_called_once() + + +def test_stream_failure_still_discards_pack(service): + svc, _ = service + + async def boom(**kwargs): + yield AudioChunk(np.zeros(1, dtype=np.int16), output=b"abc") + raise RuntimeError("inference fell over") + + svc.generate_audio_stream = boom + before = temp_packs() + with pytest.raises(RuntimeError): + post({"request": json.dumps({"input": "hello", "stream": True})}) + assert temp_packs() == before + assert transient() == {} + svc.model_manager.get_backend.return_value.forget_voice.assert_called_once() + + +def test_forget_voice_evicts_pipeline_cache(tmp_path): + from types import SimpleNamespace + + from api.src.inference.kokoro_v1 import KokoroV1 + + backend = KokoroV1() + pack = str(tmp_path / "a_tune_x.pt") + temp_copy = os.path.join(tempfile.gettempdir(), "temp_voice_a_tune_x.pt") + backend._voice_cache[f"{pack}:cpu"] = torch.zeros(1) + backend._pipelines["a"] = SimpleNamespace(voices={temp_copy: torch.zeros(1)}) + backend.forget_voice(pack) + assert backend._voice_cache == {} + assert backend._pipelines["a"].voices == {} + + +@pytest.mark.asyncio +async def test_transient_voice_resolves_before_disk(tmp_path, monkeypatch): + monkeypatch.setattr(settings, "voices_dir", str(tmp_path)) + manager = VoiceManager() + manager.register_transient("a_tune_x", str(tmp_path / "x.pt")) + assert await manager.get_voice_path("a_tune_x") == str(tmp_path / "x.pt") + manager.forget_transient("a_tune_x") + with pytest.raises(FileNotFoundError): + await manager.get_voice_path("a_tune_x") + + +def test_load_puts_tuner_on_configured_device(monkeypatch): + import inno_kokoro.enroll + + built = [] + + class Tuner: + version = "0.2.0" + + def __init__(self, path, device): + built.append((path, device)) + self.device = device + + monkeypatch.setattr(inno_kokoro.enroll, "Tuner", Tuner) + monkeypatch.setattr(inno_tuner, "_tuner", None) + monkeypatch.setattr(settings, "use_gpu", False) + inno_tuner.load() + assert inno_tuner.available() + assert built == [(inno_tuner.weights_path(), "cpu")] + assert inno_tuner.weights_path().endswith( + os.path.join("v1_0", "inno_tuner", "model.safetensors") + ) + + +def test_tune_without_tuner_raises(monkeypatch): + monkeypatch.setattr(inno_tuner, "_tuner", None) + with pytest.raises(RuntimeError): + inno_tuner.tune(wav_bytes()) + + +def test_stereo_clip_is_mixed_to_mono(service, monkeypatch): + import inno_kokoro.enroll + + shapes = [] + + def enroll(wav, sr, tuner, fmax=None, head=True): + shapes.append(tuple(wav.shape)) + return torch.zeros(510, 1, 256), {} + + monkeypatch.setattr(inno_kokoro.enroll, "enroll", enroll) + r = post({"return_voice_pack": "true"}, raw=wav_bytes(channels=2)) + assert r.status_code == 200 + assert shapes == [(4 * 24000,)] + + +def test_speech_failure_before_streaming_discards_pack(service): + svc, _ = service + svc.generate_audio.side_effect = RuntimeError("inference fell over") + before = temp_packs() + r = post({"request": json.dumps({"input": "hello", "stream": False})}) + assert r.status_code == 500 + assert temp_packs() == before + assert transient() == {} + svc.model_manager.get_backend.return_value.forget_voice.assert_called_once() + + +@pytest.mark.asyncio +async def test_discard_never_raises(monkeypatch): + from api.src.routers.tune import _discard + + monkeypatch.setattr(VoiceManager, "_instance", VoiceManager()) + svc = MagicMock() + svc.model_manager.get_backend.return_value.forget_voice.side_effect = RuntimeError( + "no pipeline" + ) + await _discard( + svc, "a_tune_x", os.path.join(tempfile.gettempdir(), "a_tune_gone.pt") + ) + + +def test_startup_survives_tuner_load_failure(monkeypatch): + from api.src.inference.model_manager import ModelManager + + monkeypatch.setattr(settings, "enable_inno_tuner", True) + monkeypatch.setattr(inno_tuner, "_tuner", None) + monkeypatch.setattr( + inno_tuner, "load", MagicMock(side_effect=OSError("no weights")) + ) + monkeypatch.setattr(ModelManager, "_instance", None) + monkeypatch.setattr(VoiceManager, "_instance", None) + monkeypatch.setattr( + ModelManager, + "initialize_with_warmup", + AsyncMock(return_value=("cpu", "kokoro_v1", 1)), + ) + with TestClient(app) as booted: + inno_tuner.load.assert_called_once() + assert not inno_tuner.available() + assert ( + booted.post( + "/dev/tune", files={"audio": ("ref.wav", wav_bytes(), "audio/wav")} + ).status_code + == 503 + ) + + +def test_decoder_crash_is_a_400_and_the_server_lives( + service, fake_tuner, monkeypatch, tmp_path +): + crasher = tmp_path / "crash.py" + crasher.write_text("import os\nos.abort()\n") + decoder = inno_tuner.DECODER + monkeypatch.setattr(inno_tuner, "DECODER", str(crasher)) + r = post({"return_voice_pack": "true"}) + assert r.status_code == 400 + assert r.json()["detail"]["message"] == "reference audio could not be decoded" + monkeypatch.setattr(inno_tuner, "DECODER", decoder) + assert post({"return_voice_pack": "true"}).status_code == 200 + + +def test_broken_decoders_stay_generic_and_the_server_lives( + service, fake_tuner, monkeypatch, tmp_path +): + raiser = tmp_path / "raise.py" + raiser.write_text("raise RuntimeError('/secret/path')\n") + sleeper = tmp_path / "sleep.py" + sleeper.write_text("import sys, time; sys.stdin.buffer.read(); time.sleep(30)") + decoder = inno_tuner.DECODER + monkeypatch.setattr(inno_tuner, "DECODE_TIMEOUT", 1) + for broken in (raiser, sleeper, tmp_path / "missing.py"): + monkeypatch.setattr(inno_tuner, "DECODER", str(broken)) + r = post({"return_voice_pack": "true"}) + assert r.status_code == 400, broken + assert r.json()["detail"]["message"] == "reference audio could not be decoded" + monkeypatch.setattr(inno_tuner, "DECODER", decoder) + assert post({"return_voice_pack": "true"}).status_code == 200 + + +def test_decode_makes_samples_finite(): + buf = io.BytesIO() + frames = np.array([0.5, np.nan, np.inf, -np.inf, -2.0], dtype=np.float32) + sf.write(buf, frames, 8000, format="WAV", subtype="FLOAT") + sr, wav = inno_tuner.decode(buf.getvalue()) + assert sr == 8000 + assert wav.tolist() == [0.5, 0.0, 1.0, -1.0, -1.0] + + +def test_one_clip_tunes_at_a_time(service, fake_tuner): + assert inno_tuner.reserve() + r = post({"return_voice_pack": "true"}) + inno_tuner._lock.release() + assert r.status_code == 503 + assert r.json()["detail"]["error"] == "busy" + assert post({"return_voice_pack": "true"}).status_code == 200 + + +@pytest.mark.asyncio +async def test_a_cancelled_request_holds_the_tuner_until_the_worker_exits( + service, fake_tuner, monkeypatch +): + import inno_kokoro.enroll + + quick = inno_kokoro.enroll.enroll + started = threading.Event() + + def enroll(*args, **kwargs): + started.set() + time.sleep(0.5) + return quick(*args, **kwargs) + + monkeypatch.setattr(inno_kokoro.enroll, "enroll", enroll) + before = temp_packs() + transport = httpx.ASGITransport(app=app) + async with httpx.AsyncClient(transport=transport, base_url="http://t") as ac: + task = asyncio.create_task( + ac.post( + "/dev/tune", + files={"audio": ("ref.wav", wav_bytes(), "audio/wav")}, + data={"return_voice_pack": "true"}, + ) + ) + await asyncio.to_thread(started.wait, 5) + task.cancel() + with pytest.raises(asyncio.CancelledError): + await task + assert not inno_tuner.reserve() + for _ in range(50): + if inno_tuner.reserve(): + break + await asyncio.sleep(0.1) + inno_tuner._lock.release() + for leftover in temp_packs() - before: + os.unlink(leftover) + assert post({"return_voice_pack": "true"}).status_code == 200 diff --git a/api/tests/test_web_player.py b/api/tests/test_web_player.py index d2e98eddf..a848a1015 100644 --- a/api/tests/test_web_player.py +++ b/api/tests/test_web_player.py @@ -6,6 +6,7 @@ from fastapi.testclient import TestClient from api.src.core.config import settings +from api.src.inference import inno_tuner from api.src.main import app client = TestClient(app) @@ -16,7 +17,12 @@ def test_web_config_reports_root_path_and_version(): response = client.get("/web/config") assert response.status_code == 200 - assert response.json() == {"root_path": "/tts", "version": settings.api_version} + assert response.json() == { + "root_path": "/tts", + "version": settings.api_version, + "tuner": inno_tuner.available(), + "voice_saving": settings.allow_local_voice_saving, + } def test_web_root_serves_index(): diff --git a/docker/cpu/docker-compose.yml b/docker/cpu/docker-compose.yml index 8e8f2fe95..b8f61839b 100644 --- a/docker/cpu/docker-compose.yml +++ b/docker/cpu/docker-compose.yml @@ -17,5 +17,6 @@ services: required: false environment: # full list in docs/configuration.md - API_LOG_LEVEL=DEBUG + # - ENABLE_INNO_TUNER=true # - ALLOW_DEV_UNLOAD=true # - ENABLE_DEBUG_ENDPOINTS=true diff --git a/docker/gpu/docker-compose.yml b/docker/gpu/docker-compose.yml index b8c42793c..6efb40e4c 100644 --- a/docker/gpu/docker-compose.yml +++ b/docker/gpu/docker-compose.yml @@ -22,6 +22,7 @@ services: required: false environment: # full list in docs/configuration.md - API_LOG_LEVEL=DEBUG + # - ENABLE_INNO_TUNER=true # - ALLOW_DEV_UNLOAD=true # - ENABLE_DEBUG_ENDPOINTS=true deploy: diff --git a/docker/rocm/docker-compose.yml b/docker/rocm/docker-compose.yml index 8bbf0ff7b..8e499a4fc 100644 --- a/docker/rocm/docker-compose.yml +++ b/docker/rocm/docker-compose.yml @@ -32,6 +32,7 @@ services: required: false environment: - USE_GPU=true + # - ENABLE_INNO_TUNER=true # - ALLOW_DEV_UNLOAD=true # - ENABLE_DEBUG_ENDPOINTS=true - TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1 diff --git a/docker/scripts/download_model.py b/docker/scripts/download_model.py index e47bafc1d..c5559742e 100644 --- a/docker/scripts/download_model.py +++ b/docker/scripts/download_model.py @@ -126,6 +126,43 @@ def download_model(output_dir: str) -> None: raise +def download_tuner(output_dir: str) -> None: + """Fetch the pinned Inno clone tuner weights into output_dir/inno_tuner, no-op if present.""" + from inno_kokoro.enroll import fetch_weights + + path = fetch_weights(os.path.join(output_dir, "inno_tuner")) + logger.info(f"✓ Inno tuner weights prepared at {path}") + check_tuner_update() + + +def check_tuner_update() -> None: + """Log if the Hub has newer tuner weights than the pinned revision. Capped at 3 s, + never raises, skipped under HF_HUB_OFFLINE or HF_HUB_DISABLE_TELEMETRY.""" + import threading + + from huggingface_hub import constants, hf_hub_url + from huggingface_hub.utils import get_session + from inno_kokoro.enroll import HUB_REPO, HUB_REVISION + + if constants.HF_HUB_OFFLINE or constants.HF_HUB_DISABLE_TELEMETRY: + return + + def get(): + try: + url = hf_hub_url(HUB_REPO, "config.json") + latest = get_session().get(url, timeout=3).json()["version"] + if latest != HUB_REVISION.lstrip("v"): + logger.info( + f"Inno tuner weights {latest} available, pinned {HUB_REVISION}" + ) + except Exception: + pass + + t = threading.Thread(target=get, daemon=True) + t.start() + t.join(3) + + def main(): """Main entry point.""" import argparse @@ -137,6 +174,10 @@ def main(): args = parser.parse_args() download_model(args.output) + try: + download_tuner(args.output) + except Exception as e: + logger.error(f"Inno tuner weights not fetched, /dev/tune will answer 503: {e}") if __name__ == "__main__": diff --git a/docs/INDEX.md b/docs/INDEX.md index 0aa5f916b..1934e0054 100644 --- a/docs/INDEX.md +++ b/docs/INDEX.md @@ -1,6 +1,6 @@ # Docs -[Configuration](configuration.md) · [Troubleshooting](troubleshooting.md) +[Configuration](configuration.md) · [Troubleshooting](troubleshooting.md) · [Voice tuning](inno-tune.md) ### Deployment - [DigitalOcean](deployment/digitalocean.md) diff --git a/docs/configuration.md b/docs/configuration.md index ac52ca5d2..6d6a28fb7 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -107,7 +107,7 @@ Names are the field names from `api/src/core/config.py`, uppercased. Unrecognize | `DEFAULT_VOICE` | `af_heart` | Voice used when a request omits one, preselected in the web player, warms the model at startup | | `DEFAULT_VOICE_CODE` | unset | Override the language code normally taken from the voice name's first letter. Applies to every speaker, so a `[voice:]` dialogue mixing languages is forced onto this one | | `VOICE_WEIGHT_NORMALIZATION` | `true` | Rescale combined voice weights to sum to 1 | -| `ALLOW_LOCAL_VOICE_SAVING` | `false` | Let combined voices be written to disk | +| `ALLOW_LOCAL_VOICE_SAVING` | `false` | Let combined voices be downloaded (`/v1/audio/voices/combine`) and tuned voices be saved into `VOICES_DIR` (`/dev/tune` `save_voice`) | | `ENABLE_VOICE_TAGS` | `true` | Kill switch for `[voice:]` parsing and `/dev/dialogue` | **Text processing** @@ -163,6 +163,7 @@ Names are the field names from `api/src/core/config.py`, uppercased. Unrecognize |---|---|---| | `ENABLE_DEBUG_ENDPOINTS` | `false` | Expose `/debug/*` host and process introspection | | `ALLOW_DEV_UNLOAD` | `false` | Expose `/dev/model`, `POST /dev/unload`, and `POST /dev/reload` | +| `ENABLE_INNO_TUNER` | `false` | Expose `POST /dev/tune`, see [inno-tune.md](inno-tune.md) | | `MODEL_AUTO_UNLOAD_TIMEOUT_SECONDS` | `0.0` | Idle seconds before auto-unload; `0` disables auto-unload | ## Text normalization diff --git a/docs/inno-tune.md b/docs/inno-tune.md new file mode 100644 index 000000000..7a80a04e5 --- /dev/null +++ b/docs/inno-tune.md @@ -0,0 +1,93 @@ +# Voice tuning (`/dev/tune`) + +*Last updated: 2026-09-09* + +`POST /dev/tune` takes a short reference clip and returns speech in a voice tuned toward it, or the voice pack itself. Backed by [inno-kokoro](https://github.com/remsky/inno-kokoro) and its [weights](https://huggingface.co/remsky/kokoro-inno-clone-tuner). + +It is a tuner, not a cloner. The stock model is steered toward the reference's timbre, pitch, pitch range, and speaking rate. Expect a voice in the same neighbourhood, not a match. English only. Only tune voices you have permission to use. + +Off by default, `ENABLE_INNO_TUNER=true` turns it on. + +The web player has a Tune tab with the same controls. + +## Request + +Multipart form. `audio` is the only required field. + +| Field | Default | | +|---|---|---| +| `audio` | required | Reference clip, any format libsndfile reads (wav, flac, ogg, mp3). 3 to 30 s, one speaker, mono or stereo, 8 to 96 kHz, 10 MB cap. Only the first 30 s are decoded | +| `request` | unset | JSON with the same fields as the `/v1/audio/speech` body, minus `voice` (`input`, `response_format`, `speed`, `stream`, `lang_code`, `normalization_options`, `return_download_link`, etc). When set, the clip's voice speaks it | +| `prosody_head` | `true` | Apply the prosody head. `false` uses the plain stock blend, closer to a stock voice | +| `fmax` | auto | Pitch tracking ceiling in Hz, 60 to 1000. Unset, the tuner picks one from the clip's harmonics. Set it if a band-limited clip reads an octave high | +| `return_voice_pack` | `false` | Return the tuned `.pt` instead of audio | +| `save_voice` | unset | Keep the pack as `VOICES_DIR/_tuned.pt`. `af_`, `am_`, `bf_`, or `bm_` plus lowercase letters, digits, single underscores. Needs `ALLOW_LOCAL_VOICE_SAVING=true` | + +`allow_voice_tags` and `ssml` inside `request` are ignored. The reference clip is the voice. + +## Responses + +| Call | Response | +|---|---| +| `request` set | Audio, same shape and headers as `/v1/audio/speech` (streamed by default, `X-Download-Path` with `return_download_link`) | +| `return_voice_pack=true` | The `[510, 1, 256]` pack as `_tuned.pt` (`a_tune.pt` unnamed), wins over `request` | +| `save_voice` only | `{"voice": "_tuned"}` | + +Packs live in the OS temp dir for the length of the response and are removed after, along with their cached tensor. Nothing is kept unless `save_voice` is set. + +```bash +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F 'request={"input":"Hello there."}' -o out.mp3 +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F return_voice_pack=true -o ref.pt +curl -s http://localhost:8880/dev/tune -F audio=@ref.wav -F save_voice=am_ref # saves am_ref_tuned +``` + +## Naming + +Every voice is a flat `_` file in `VOICES_DIR`, a suffix says where it came from: + +| | Example | | +|---|---|---| +| stock | `af_heart` | from the model card | +| shipped tuned | `am_price_inno` | four tuned voices bundled with the server: `af_amelia_inno`, `af_goodall_inno`, `am_price_inno`, `bm_atten_inno` | +| saved by you | `am_ref_tuned` | `save_voice=am_ref`, the suffix is added | +| in flight | `a_tune_` | the transient pack while a `/dev/tune` request runs, never listed | + +`save_voice` must match `^[ab][a-z]?_[a-z0-9]+(_[a-z0-9]+)*$` (uppercase is lowered first; the first letter picks US or UK English, the optional second letter is naming convention only, the web player uses `x`) and gets `_tuned` appended, else 400. Hyphens are out because `-` is a combine operator. An existing name is a 409, delete the file to replace it. + +A saved voice is a plain voice. It works anywhere a voice name does (`voice`, combine syntax, `[voice:]` tags, SSML) and shows up in `/v1/audio/voices`. The language code comes from the first letter, `a` American, `b` British. + +Off by default. With `ALLOW_LOCAL_VOICE_SAVING=true` any caller can write into `VOICES_DIR`, so keep it off on a shared host. + +## Settings + +| Variable | Default | | +|---|---|---| +| `ENABLE_INNO_TUNER` | `false` | Load the tuner at startup and expose the route, otherwise 403 | +| `ALLOW_LOCAL_VOICE_SAVING` | `false` | Allow `save_voice` | +| `VOICES_DIR` | `/app/api/src/voices/v1_0` | Where saved packs go | +| `HF_HUB_OFFLINE` / `HF_HUB_DISABLE_TELEMETRY` | unset | Either one skips the startup version check below | + +`docker/scripts/download_model.py` fetches the pinned tuner weights into `MODEL_DIR/v1_0/inno_tuner/` next to the Kokoro weights. The Docker images ship with them. + +On load the package does one GET of the weights' `config.json` on Hugging Face and logs a line if newer weights exist. `HF_HUB_OFFLINE=1` skips it. + +If the weights are missing or fail to load, the server still starts, logs `Inno voice tuner not loaded`, and the route answers 503. + +## Errors + +| Status | | +|---|---| +| 400 | Clip under 3 s, over 2 channels, outside 8 to 96 kHz, unreadable audio, bad `request` JSON, `save_voice` off the naming rule, nothing to do (no `request`, `return_voice_pack`, or `save_voice`) | +| 403 | `ENABLE_INNO_TUNER=false`, or `save_voice` without `ALLOW_LOCAL_VOICE_SAVING` | +| 409 | `` already exists | +| 413 | Clip over 10 MB | +| 503 | Tuner not loaded, or busy with another clip | + +## How it works + +A Kokoro voice pack is 510 rows of 256 floats: the first 128 drive the decoder (timbre), the last 128 drive the predictor (prosody). + +- Timbre: a speaker embedding of the clip goes through a learned style head, plus a shift along a learned direction for the clip's spectral tilt. +- Prosody: the stock packs' predictor halves are blended by least squares to the clip's pitch mean, pitch spread, and syllable rate. The prosody head then nudges the blend from the same measurements. + +Enrollment takes well under a second on a GPU, a few seconds on CPU. Clips shorter than 5 s, or unusually dark or bright ones (archival, heavy filtering), tune less reliably. diff --git a/pyproject.toml b/pyproject.toml index 359b3267e..bfc686420 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -19,6 +19,8 @@ dependencies = [ "scipy==1.14.1", # Audio processing "soundfile==0.13.0", + "inno-kokoro==0.2.0", + "python-multipart>=0.0.20", "regex>=2025.10.22", "unicode-segmentation-rs>=0.3.3", # Utilities diff --git a/uv.lock b/uv.lock index 15bd98725..aa922d7d5 100644 --- a/uv.lock +++ b/uv.lock @@ -1126,6 +1126,30 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/cb/b1/3846dd7f199d53cb17f49cba7e651e9ce294d8497c8c150530ed11865bb8/iniconfig-2.3.0-py3-none-any.whl", hash = "sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12", size = 7484, upload-time = "2025-10-18T21:55:41.639Z" }, ] +[[package]] +name = "inno-kokoro" +version = "0.2.0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "huggingface-hub" }, + { name = "kokoro" }, + { name = "praat-parselmouth" }, + { name = "safetensors" }, + { name = "scipy" }, + { name = "soundfile" }, + { name = "torch", version = "2.8.0", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "(sys_platform == 'darwin' and extra == 'extra-14-kokoro-fastapi-cpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0", source = { registry = "https://pypi.org/simple" }, marker = "(platform_machine != 'aarch64' and platform_machine != 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu') or (platform_machine != 'aarch64' and platform_machine != 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (platform_machine == 'aarch64' and extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (platform_machine == 'aarch64' and extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (platform_machine == 'aarch64' and extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm') or (platform_machine == 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (platform_machine == 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (platform_machine == 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra != 'extra-14-kokoro-fastapi-cpu' and extra != 'extra-14-kokoro-fastapi-gpu' and extra != 'extra-14-kokoro-fastapi-gpu-cu128' and extra != 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "(sys_platform != 'darwin' and extra == 'extra-14-kokoro-fastapi-cpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0+cu126", source = { registry = "https://download.pytorch.org/whl/cu126" }, marker = "(platform_machine == 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0+cu128", source = { registry = "https://download.pytorch.org/whl/cu128" }, marker = "(platform_machine == 'x86_64' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0+cu129", source = { registry = "https://download.pytorch.org/whl/cu129" }, marker = "(platform_machine == 'aarch64' and extra == 'extra-14-kokoro-fastapi-gpu') or (platform_machine == 'aarch64' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm')" }, + { name = "torch", version = "2.8.0+rocm6.4", source = { registry = "https://download.pytorch.org/whl/rocm6.4" }, marker = "(extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra != 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra != 'extra-14-kokoro-fastapi-cpu' and extra != 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm')" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/4c/25/fc62c06761866e9e645556b942cf49daf62fb5801b340ae775ded3e45a33/inno_kokoro-0.2.0.tar.gz", hash = "sha256:913cf4244d836fe4be8db9ad15ba1940a1780b65d834a6fe3da88b8dcecb8223", size = 21164, upload-time = "2026-09-05T07:45:20.563Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/d5/7e/6dcee78130103f41c7e8a0143aa040a8ce330e5553bf5719642a174ad1d9/inno_kokoro-0.2.0-py3-none-any.whl", hash = "sha256:faebd4006d939f99cf3a8b30c6bdc280f74309c42b88887362e4d28995601664", size = 18324, upload-time = "2026-09-05T07:45:19.597Z" }, +] + [[package]] name = "isodate" version = "0.7.2" @@ -1339,6 +1363,7 @@ dependencies = [ { name = "espeakng-loader" }, { name = "fastapi" }, { name = "inflect" }, + { name = "inno-kokoro" }, { name = "kokoro" }, { name = "loguru" }, { name = "misaki", extra = ["en", "ja", "ko", "zh"] }, @@ -1352,6 +1377,7 @@ dependencies = [ { name = "pydantic-settings" }, { name = "pyopenjtalk-plus", marker = "sys_platform == 'win32' or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-cpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-gpu-cu128') or (extra == 'extra-14-kokoro-fastapi-gpu' and extra == 'extra-14-kokoro-fastapi-rocm') or (extra == 'extra-14-kokoro-fastapi-gpu-cu128' and extra == 'extra-14-kokoro-fastapi-rocm')" }, { name = "python-dotenv" }, + { name = "python-multipart" }, { name = "regex" }, { name = "requests" }, { name = "scipy" }, @@ -1405,6 +1431,7 @@ requires-dist = [ { name = "httpx", marker = "extra == 'test'", specifier = "==0.26.0" }, { name = "hypothesis", marker = "extra == 'test'", specifier = "==6.167.1" }, { name = "inflect", specifier = ">=7.5.0" }, + { name = "inno-kokoro", specifier = "==0.2.0" }, { name = "jinja2", marker = "extra == 'test'", specifier = ">=3.1.6" }, { name = "kokoro", specifier = "==0.9.4" }, { name = "loguru", specifier = "==0.7.3" }, @@ -1422,6 +1449,7 @@ requires-dist = [ { name = "pytest-asyncio", marker = "extra == 'test'", specifier = "==0.25.3" }, { name = "pytest-cov", marker = "extra == 'test'", specifier = "==6.0.0" }, { name = "python-dotenv", specifier = "==1.2.2" }, + { name = "python-multipart", specifier = ">=0.0.20" }, { name = "pytorch-triton-rocm", marker = "extra == 'rocm'", specifier = ">=3.2.0", index = "https://download.pytorch.org/whl/rocm6.4", conflict = { package = "kokoro-fastapi", extra = "rocm" } }, { name = "regex", specifier = ">=2025.10.22" }, { name = "requests", specifier = "==2.33.0" }, @@ -2392,6 +2420,62 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" }, ] +[[package]] +name = "praat-parselmouth" +version = "0.4.7" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "numpy" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/4f/28/2c1204fe3e7aeb6942051ff6776e31da52c7ab5b7df2ca438f371bd60d8a/praat_parselmouth-0.4.7.tar.gz", hash = "sha256:6dd81d246ce1eef5fd93d8cbdaf1bef61ca40ef1d2fc12aa23996a28071181e6", size = 22526491, upload-time = "2025-11-27T20:08:58.276Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/e2/64/8f5d94ae6b3a11828185a6a5b6b8fce0446db0885fa30ed02f180ba08dce/praat_parselmouth-0.4.7-cp310-cp310-macosx_10_9_universal2.whl", hash = "sha256:16243bd11c671829a740560f42e323cd323469e3eff08352610fe8ee1df913f2", size = 17664197, upload-time = "2025-11-27T20:01:15.91Z" }, + { url = "https://files.pythonhosted.org/packages/d5/ed/8b609582c9859becb54f12fed5689738347161fd348b3d0600904f5c1023/praat_parselmouth-0.4.7-cp310-cp310-macosx_10_9_x86_64.whl", hash = "sha256:c1b3468c962d745d1c3bc09535bc5c0cf5416a7c0847637ebeb987311c860b7a", size = 9044039, upload-time = "2025-11-27T20:01:19.742Z" }, + { url = "https://files.pythonhosted.org/packages/fc/ed/7acfcfe09584de89037a832ab759c68107bc5b6260131d396c5eafeeb3bf/praat_parselmouth-0.4.7-cp310-cp310-macosx_11_0_arm64.whl", hash = "sha256:699b4dec1c4251d0cd9ad004eaeb49bd1b11c671677436e8fc6b191b93fa2bc9", size = 8634313, upload-time = "2025-11-27T20:01:23.297Z" }, + { url = "https://files.pythonhosted.org/packages/5b/9a/38a9b2883b9d617398fda35342b31a4e9f6ca3b343fe7b77466748aa6d60/praat_parselmouth-0.4.7-cp310-cp310-manylinux_2_12_i686.manylinux2010_i686.whl", hash = "sha256:29674d590189fef2835c6cdc04d8230e90bb324d2fbde2094065c7b2f7e24ac9", size = 10299898, upload-time = "2025-11-27T20:01:26.812Z" }, + { url = "https://files.pythonhosted.org/packages/12/28/020eaa62e6aa75024a3ab7b04305854fabb6c014bfc7fee9c6ac40b28c85/praat_parselmouth-0.4.7-cp310-cp310-manylinux_2_12_x86_64.manylinux2010_x86_64.whl", hash = "sha256:319cd87cc648b57b49f443249bb30612c1c0ac7e084f49ce219e384d96c16098", size = 10723051, upload-time = "2025-11-27T20:01:30.719Z" }, + { url = "https://files.pythonhosted.org/packages/50/f2/1a9ae7192c47f754bf18668ce23423f9637caeb5168d7a6cbf641204bfa3/praat_parselmouth-0.4.7-cp310-cp310-win32.whl", hash = "sha256:e36f8eb599929e95ae1824c70be27de0c3eacd8c3550440f87cdacf649454ee5", size = 8128042, upload-time = "2025-11-27T20:01:35.927Z" }, + { url = "https://files.pythonhosted.org/packages/87/5f/b25cb817e4727c759cd92c451d5d8a554929adbf0f4b72c98531f98b698e/praat_parselmouth-0.4.7-cp310-cp310-win_amd64.whl", hash = "sha256:fd6f1946e463dfc1b29a8ac4ad5b9792f5bb2436db74f42f2ee4b04168220583", size = 8954447, upload-time = "2025-11-27T20:01:39.262Z" }, + { url = "https://files.pythonhosted.org/packages/11/74/9f27fc62ac4afb06c75f869bc32ed281b507391ad54aa4953562f1381b89/praat_parselmouth-0.4.7-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:7193ecb78b7dde649800aaadbd14364ce8dbba7608c3760da24ed6c1c47731df", size = 17664259, upload-time = "2025-11-27T20:01:44.921Z" }, + { url = "https://files.pythonhosted.org/packages/16/cc/204a663ea0e26c388ded379e444b7c57232108f6fa85a0f5729d701f74ea/praat_parselmouth-0.4.7-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:508b10da7a958c71c6cda8ad22569917774002da4c9841a571a308a50dff0107", size = 9044039, upload-time = "2025-11-27T20:01:47.933Z" }, + { url = "https://files.pythonhosted.org/packages/7e/b3/cf0de61814c676e6b99a6580a05633da90a3f1a3bc1c8a826edcbe3fc7cd/praat_parselmouth-0.4.7-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:88380dd454a26b613eaef810bcc80e842f5eef802631c20e00cdef92aee91d8a", size = 8634343, upload-time = "2025-11-27T20:01:52.85Z" }, + { url = "https://files.pythonhosted.org/packages/03/17/7cd5fe8f24dd5477cdeea708ebbd4e4fc1e555e6eba147643f1737c977f9/praat_parselmouth-0.4.7-cp311-cp311-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:dc4812928cb05e63de0f291fcbd787a0e831b06bcdf74bd9df34500cdfd6edbc", size = 10717365, upload-time = "2025-11-27T20:01:55.919Z" }, + { url = "https://files.pythonhosted.org/packages/d2/73/4475dcc95fc51b6cd6968c2e20df576ba4bba78ebca024034a333be70e5f/praat_parselmouth-0.4.7-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:618310c98b5bb0de5ea3a6f365332bbffb4dc0a7c3c5442a5041bf93d5409036", size = 10753885, upload-time = "2025-11-27T20:01:59.611Z" }, + { url = "https://files.pythonhosted.org/packages/d9/0a/dbb3be27ef74eb6b111af094864c02cca526f6982700cccbd295fff6db73/praat_parselmouth-0.4.7-cp311-cp311-win32.whl", hash = "sha256:f879a0dea38243e46acc6e5ce385508bf4f6131a85bc1cfc67ae5872a25112e2", size = 8127601, upload-time = "2025-11-27T20:02:04.134Z" }, + { url = "https://files.pythonhosted.org/packages/71/c9/f49ba95b171c91770adb1401d401d9cea2d98f05f4304138d74e737d9146/praat_parselmouth-0.4.7-cp311-cp311-win_amd64.whl", hash = "sha256:87671f42cb52ecdb0e1bc8f12786b37dbb5242c80968ff77c85aa30c8767612e", size = 8954605, upload-time = "2025-11-27T20:02:06.074Z" }, + { url = "https://files.pythonhosted.org/packages/71/8b/e3ac60fd972d457ecf6b0b0bdd9b651f578c0f50ce31b4e176065ad1713d/praat_parselmouth-0.4.7-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:fe123f5004e77a3a39150fc0b64b8960404c87ca8bb2a2e031c73ebd32d8e548", size = 17673383, upload-time = "2025-11-27T20:02:09.287Z" }, + { url = "https://files.pythonhosted.org/packages/d5/e3/60f2caa3df1382bf24e171c6be2bf7c8fc0d5f9b443701a9b3397596c899/praat_parselmouth-0.4.7-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:5892deae3631d691c6339455f1bce065e095c77634299e57bb00df94041f7db3", size = 9051957, upload-time = "2025-11-27T20:02:13.435Z" }, + { url = "https://files.pythonhosted.org/packages/75/ea/10e0f1bb64923c27dec9fcd3c28852751565df8ba3beb5ae1d8288fa65f8/praat_parselmouth-0.4.7-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:29e468febc3de2acad8f7718b10cc800610871e229afc549b6b51f009a87ede7", size = 8635367, upload-time = "2025-11-27T20:02:16.791Z" }, + { url = "https://files.pythonhosted.org/packages/f9/a8/d410e7f10ca97293cf271f94c8ddff70deeac44d3b8a524aa44c796f52cc/praat_parselmouth-0.4.7-cp312-cp312-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:e3fc0be825c5d1a4a05d2f85c5d8ff214d7178579cd2dd036cb902d9737a8d80", size = 10711677, upload-time = "2025-11-27T20:02:19.445Z" }, + { url = "https://files.pythonhosted.org/packages/9f/5f/e4495c8ac6ed26a1f16a34269c12579532e747b4ac5db57123093904f2ec/praat_parselmouth-0.4.7-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:6b5fbfab5b3cd4f6146e6145559d5d7c231384d88f942e0c8c0e791058aa61a5", size = 10749363, upload-time = "2025-11-27T20:02:23.082Z" }, + { url = "https://files.pythonhosted.org/packages/50/f3/6b7a92ab093e89b58c7abbda93ca9a7b1e72c1a5c020d49f42e8856c6740/praat_parselmouth-0.4.7-cp312-cp312-win32.whl", hash = "sha256:fb579964b92e7aa5774badfa760570a6645641592cb5eb9e50b5a2225875b210", size = 8125421, upload-time = "2025-11-27T20:02:26.073Z" }, + { url = "https://files.pythonhosted.org/packages/6c/df/a358062226918326825b7af95850ce561a22c6dcdc2cac0665a8efc6a2d9/praat_parselmouth-0.4.7-cp312-cp312-win_amd64.whl", hash = "sha256:7cfd6941a4e8dd1d8398260f1269a831b84e9f230b03c569993e144a20704f99", size = 8958177, upload-time = "2025-11-27T20:02:29.308Z" }, + { url = "https://files.pythonhosted.org/packages/be/1c/a772c9b064a935d40505699e69c6c7c513fbaa4725784bfef7fe36f47270/praat_parselmouth-0.4.7-cp313-cp313-macosx_10_13_universal2.whl", hash = "sha256:97a2f9225fe63b9d6ea4fd2787983faf4ad46b83b3d772abbc8b066e9ef8559f", size = 17673409, upload-time = "2025-11-27T20:02:33.031Z" }, + { url = "https://files.pythonhosted.org/packages/c2/a9/5e38d5ce459795c73600dedca72912955a299eefc4cbd0759529d30fe488/praat_parselmouth-0.4.7-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:6ec2b7c3d59b41794923bf6ca7dd97f64bc295b1a4c7835e732add2cb0fabce9", size = 9051832, upload-time = "2025-11-27T20:02:35.545Z" }, + { url = "https://files.pythonhosted.org/packages/53/16/ac1b44f6247d3abd86fc66634a32f56d43cc58f7314db81c40fd67fa2afd/praat_parselmouth-0.4.7-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:02b2b351e9dce18fb795ef91ce92107343183cff2cd8e56b40050016a45c7778", size = 8635340, upload-time = "2025-11-27T20:02:38.939Z" }, + { url = "https://files.pythonhosted.org/packages/3e/e4/821007a09ca572ae4d4e0f431b925d47aa5ad41fea21de2f6283f7ef8e5c/praat_parselmouth-0.4.7-cp313-cp313-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:b6430856750467b8dc84a7dec87ca52a43dc6e2c44a4f1864e8f01f32901ceeb", size = 10711649, upload-time = "2025-11-27T20:02:42.222Z" }, + { url = "https://files.pythonhosted.org/packages/48/b8/4eb427cc8d03156990360cf21e2e3f9180a8c0a0a03698c9d2dd36c54e0c/praat_parselmouth-0.4.7-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:c22d9f80984e91162e3be13109637f5dfe10d2ca4d3b15edd7e782550ba627f1", size = 10749436, upload-time = "2025-11-27T20:02:45.746Z" }, + { url = "https://files.pythonhosted.org/packages/8f/02/082031f1d6980aaeff05984225a6feadd948047ed6f4cbc0b00281a33950/praat_parselmouth-0.4.7-cp313-cp313-win32.whl", hash = "sha256:768b603cccd1b98db057aed3af52ba26c7a930ac4d48c2371bf6bb6702d7a068", size = 8125197, upload-time = "2025-11-27T20:02:49.291Z" }, + { url = "https://files.pythonhosted.org/packages/83/35/089a5a8dd45800dcbaa23a414d6f4f5ec29a46fec04d654923efdb3a6746/praat_parselmouth-0.4.7-cp313-cp313-win_amd64.whl", hash = "sha256:e3bfe1fc8d0f0252766978affa2e7f9b68876540688dfb9c56535fcaa2f4776a", size = 8958175, upload-time = "2025-11-27T20:02:53.583Z" }, + { url = "https://files.pythonhosted.org/packages/27/a6/50ab4dfee275a55cfe72383f861631f65753eac83b4dc6674ab173f04324/praat_parselmouth-0.4.7-cp314-cp314-macosx_10_15_universal2.whl", hash = "sha256:f022783e8c7fbc3985194d62db64a9ccde60a3ce744633dab8b55f594c12f625", size = 17668901, upload-time = "2025-11-27T20:03:01.486Z" }, + { url = "https://files.pythonhosted.org/packages/94/87/03cef2fc1365000c0bd20f87bf282b3a913a39a2cd9edfc9fb3239c85f4c/praat_parselmouth-0.4.7-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:d656eb4a18edcf52735e95b0a68c32edf7202c591c35afae74cfc54bead88277", size = 9049021, upload-time = "2025-11-27T20:03:21.658Z" }, + { url = "https://files.pythonhosted.org/packages/f5/8d/bb1121fd50ddcecd9e5d2d150c7a8656968c984e3562b623c64b02d90b10/praat_parselmouth-0.4.7-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:998138bf2acb15ae329caa217d523965b897417a1a2df130a2ecf41b664bfbf1", size = 8634455, upload-time = "2025-11-27T20:03:26.958Z" }, + { url = "https://files.pythonhosted.org/packages/89/06/00003910356fb74284e9eb58c64bea0d0c85311061e7e39c080e370a86ef/praat_parselmouth-0.4.7-cp314-cp314-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:313710a579d47effa50652d1cb3e335d926c1c0d1cea13e62dde8041676498f0", size = 10711434, upload-time = "2025-11-27T20:03:30.22Z" }, + { url = "https://files.pythonhosted.org/packages/bf/fa/a9a2bb835c3c13e8212093f060831323c0cb1172223ba3af93f66212e4e0/praat_parselmouth-0.4.7-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:528cec1f4d2bbe02ed58fc6575071e9ccbf0d7ece815613c05d516e076b58a68", size = 10748994, upload-time = "2025-11-27T20:03:33.313Z" }, + { url = "https://files.pythonhosted.org/packages/f7/ef/399a0f8b0c64d8910dc75826ba8c5fbb9490fa4b2e0777fbda7293db55d5/praat_parselmouth-0.4.7-cp314-cp314-win32.whl", hash = "sha256:71ff8990d454d0fa164aef269d1b2acccb6778d5c26a303d7e668b053a81af56", size = 8232788, upload-time = "2025-11-27T20:03:36.432Z" }, + { url = "https://files.pythonhosted.org/packages/38/d4/bbc8153bde9812b33714c210b41ddb50d5af56a6d155f23ee44a01724146/praat_parselmouth-0.4.7-cp314-cp314-win_amd64.whl", hash = "sha256:aa467b5839fa9dd6e1aff146a80e80e157aedc93825772dfb4258f772c062ebb", size = 9110034, upload-time = "2025-11-27T20:03:38.516Z" }, + { url = "https://files.pythonhosted.org/packages/d9/01/fb567977deff7eb4617393b84e7729577e2dd801d847704e884f46c73208/praat_parselmouth-0.4.7-pp310-pypy310_pp73-macosx_10_15_x86_64.whl", hash = "sha256:71f4f157c58db1d43b9c851b7348508628dd7cf0b1382a5f6517eff6c03e4171", size = 9044833, upload-time = "2025-11-27T20:05:56.371Z" }, + { url = "https://files.pythonhosted.org/packages/ad/62/e2feeb05f0b833c350b9fe88d8e6208b703159f448fca6246924ea213a36/praat_parselmouth-0.4.7-pp310-pypy310_pp73-macosx_11_0_arm64.whl", hash = "sha256:d746609bfe8f098f53feb8685e6b7b12d171622ec8a0d98d5eac7d4c1aac2e5f", size = 8634302, upload-time = "2025-11-27T20:06:01.35Z" }, + { url = "https://files.pythonhosted.org/packages/36/8a/caf3668960743c2a7ffc2dbf69bc6ea7f852f685a47c59fde965f25e2f26/praat_parselmouth-0.4.7-pp310-pypy310_pp73-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:00387f2556c57cfc4bf30605d21ec3bb710096b24719dce6aebfdafe1cd64913", size = 10716386, upload-time = "2025-11-27T20:06:04.272Z" }, + { url = "https://files.pythonhosted.org/packages/c7/b7/7f1573ae35cf08960e0642d7e29a2be3eaa584693cf9f2b619f32f6c86f0/praat_parselmouth-0.4.7-pp310-pypy310_pp73-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:dda954a0699e33e17599780c73819facc024087c5e6c3cd5630d223d3c3eaf07", size = 10757758, upload-time = "2025-11-27T20:06:07.165Z" }, + { url = "https://files.pythonhosted.org/packages/42/5f/e1df6d9ab284ee65ce75caed2cd55e760cf4d68330f7d39644743c569c55/praat_parselmouth-0.4.7-pp310-pypy310_pp73-win_amd64.whl", hash = "sha256:027cb0ec6c8f28214b1fe1ef19261e501e929472f37092f62ed68b30076e852a", size = 8954070, upload-time = "2025-11-27T20:06:12.136Z" }, + { url = "https://files.pythonhosted.org/packages/67/1b/9b8d219a0b9ce7358d428585b51c3d33b8faab87b3fb4552cf2a00907e03/praat_parselmouth-0.4.7-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:907ea3d2a466ae9d636a0cf55e4df898157637345e7695422de41b420b7f7fe9", size = 9044824, upload-time = "2025-11-27T20:06:17.976Z" }, + { url = "https://files.pythonhosted.org/packages/5f/c0/ab24f80e230d0a88662ef2627e9a13d0749a7c778f429459bfc76047c0e9/praat_parselmouth-0.4.7-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:6ac46a16e51b265f81aaa0696fe889a088491903d53c28032e1b7c9044fa3493", size = 8634302, upload-time = "2025-11-27T20:06:23.494Z" }, + { url = "https://files.pythonhosted.org/packages/fc/a4/df400fc823ab6ba29a7a6efe2288f0df5d83c92b526a9d32c4483af84a08/praat_parselmouth-0.4.7-pp311-pypy311_pp73-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:a5b0adb5b60bc7006f76afbcea4c806f23c3e0a571e80fb1028a007c2029f0d2", size = 10716355, upload-time = "2025-11-27T20:06:29.758Z" }, + { url = "https://files.pythonhosted.org/packages/30/a3/ecb1c08018de89c5f4ece53b80fd28b5ddf3446907b74d04aefaa489803e/praat_parselmouth-0.4.7-pp311-pypy311_pp73-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:70dfa96c18a3423eaf805c6326218ec7777d2f9e0e21a1c31fdd176eb9770118", size = 10757499, upload-time = "2025-11-27T20:06:34.538Z" }, + { url = "https://files.pythonhosted.org/packages/58/de/1041b4e5673f29eddf6f8cba0ac2a74e24a468e9d3e1f3a5c49ba83171f2/praat_parselmouth-0.4.7-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:6cc2c06534828ba3ab8e2438cbab8672bb476beb1b6b06616c7945344ef115f6", size = 8954414, upload-time = "2025-11-27T20:06:38.802Z" }, +] + [[package]] name = "preshed" version = "3.0.12" @@ -2747,6 +2831,15 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/0b/d7/1959b9648791274998a9c3526f6d0ec8fd2233e4d4acce81bbae76b44b2a/python_dotenv-1.2.2-py3-none-any.whl", hash = "sha256:1d8214789a24de455a8b8bd8ae6fe3c6b69a5e3d64aa8a8e5d68e694bbcb285a", size = 22101, upload-time = "2026-03-01T16:00:25.09Z" }, ] +[[package]] +name = "python-multipart" +version = "0.0.32" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/5b/42/55c32bb9b12693c092ad250a0e82edb5b31ddeda6eb772de5f308b3804ad/python_multipart-0.0.32.tar.gz", hash = "sha256:be54b7f3fa167bb83e4fcd936b887b708f4e57fe75911c02aebf53efaf8d938e", size = 46881, upload-time = "2026-06-04T16:18:58.647Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/e1/04/e8135ebd1ad02c56ec633277529b2602ff99ff634be76cdba5744cf554fd/python_multipart-0.0.32-py3-none-any.whl", hash = "sha256:ff6d3f776f16878c894e52e107296ffc890e913c611b1a4ec6c44e2821fe2e23", size = 30042, upload-time = "2026-06-04T16:18:57.319Z" }, +] + [[package]] name = "pytorch-triton-rocm" version = "3.4.0" diff --git a/web/index.html b/web/index.html index 5e723f549..32ff18e1d 100644 --- a/web/index.html +++ b/web/index.html @@ -69,7 +69,8 @@

FastKoko

- + +
@@ -96,6 +97,39 @@

FastKoko

+