A benchmarking tool that compares speech segments from Korean video content across two LID models × raw / denoise inputs = a 4-way matrix. Outputs per-segment CSV plus a 4-combination timing CSV.
Parent: provides the LID-model selection rationale for the STT pipeline (poc-stt-bench).
| raw audio | denoise audio | |
|---|---|---|
| Whisper LID | W-raw | W-den |
| VoxLingua107 | V-raw | V-den |
Flow: audio → denoise → VAD → speech segments × 4 LID calls → CSV / summary.
LID (Language Identification) classifies which language a given speech segment is spoken in. It sits on the critical path of the STT pipeline — a wrong identification corrupts the entire downstream transcription.
poc-lid-bench/
├── pyproject.toml
├── conf.py # path constants (MODEL_DIR)
├── log.py # file-based logger
├── main.py # entry point — iterates wav list
├── lib/
│ ├── denoise.py # DeepFilterNet v3 wrapper
│ ├── vad.py # Silero VAD wrapper
│ ├── whisper_lid.py # faster-whisper detect_language wrapper
│ ├── voxlingua_lid.py # SpeechBrain VoxLingua107 wrapper
│ ├── util.py # fmt_time / CSV helpers
│ └── bench.py # 4-way comparison driver
└── output/ # artifacts (recommend gitignore)
├── denoise/<stem>.wav # denoised audio (48kHz int16)
├── <stem>.csv # per-segment LID results, one per input
└── timings.csv # 4-combination timings (long format)
uv sync- Python
>=3.11,<3.12(matches DeepFilterNet's official support range) - Single GPU (cuda:0 fixed), CUDA 12.8 wheels
torch / torchaudiopinned to<2.9— works around the removal oftorchaudio.backend.common.AudioMetaDatain 2.9, which breaks DeepFilterNet's imports- All four models (Whisper LID / VoxLingua107 / DeepFilterNet / Silero VAD) are downloaded automatically on first run (network required)
Edit the test_files list in main.py with your wav paths:
test_files = [
"/path/to/audio1.wav",
"/path/to/audio2.wav",
]16kHz mono WAV is recommended; other sample rates are resampled internally.
.venv/bin/python main.pyFlow:
- Pre-load all four models (so load time is excluded from measurements)
- Iterate
test_files— callbench.run()per file - Once all files are done, write
output/timings.csvviabench.save_timings()
| Path | Content |
|---|---|
output/denoise/<stem>.wav |
Denoised audio (48kHz int16) |
output/<stem>.csv |
Per-segment 4-way LID results (lang, prob) |
output/timings.csv |
Per-file, 4-combination timings (long format) |
/usr/service/logs/scenemaker/lid_bench.log |
Progress log (internal path — edit log.py) |
The console prints only the per-file summary (duration / VAD segments / elapsed / agreement rates). Details are in the CSV files above.
Environment notes (for external git users)
- If
HF_HOMEis unset, models download to the default cache (~/.cache/huggingface/).conf.MODEL_DIRdefaults to an internal path (/stg/models). Outside this environment, change it to your own path, or modify the code to use SpeechBrain's default cache.- The log path (
DEFAULT_LOG_FILEinlog.py) is also internal — change it to a path of your own.