Skip to content

About

LID model comparison — Whisper LID vs VoxLingua107 (raw vs denoise)

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

English | 한국어


poc-lid-bench

A benchmarking tool that compares speech segments from Korean video content across two LID models × raw / denoise inputs = a 4-way matrix. Outputs per-segment CSV plus a 4-combination timing CSV.

Parent: provides the LID-model selection rationale for the STT pipeline (poc-stt-bench).

Overview

raw audio denoise audio
Whisper LID W-raw W-den
VoxLingua107 V-raw V-den

Flow: audio → denoise → VAD → speech segments × 4 LID calls → CSV / summary.

LID (Language Identification) classifies which language a given speech segment is spoken in. It sits on the critical path of the STT pipeline — a wrong identification corrupts the entire downstream transcription.

Directory Layout

poc-lid-bench/
├── pyproject.toml
├── conf.py                # path constants (MODEL_DIR)
├── log.py                 # file-based logger
├── main.py                # entry point — iterates wav list
├── lib/
│   ├── denoise.py         # DeepFilterNet v3 wrapper
│   ├── vad.py             # Silero VAD wrapper
│   ├── whisper_lid.py     # faster-whisper detect_language wrapper
│   ├── voxlingua_lid.py   # SpeechBrain VoxLingua107 wrapper
│   ├── util.py            # fmt_time / CSV helpers
│   └── bench.py           # 4-way comparison driver
└── output/                # artifacts (recommend gitignore)
    ├── denoise/<stem>.wav   # denoised audio (48kHz int16)
    ├── <stem>.csv           # per-segment LID results, one per input
    └── timings.csv          # 4-combination timings (long format)

Installation

uv sync
  • Python >=3.11,<3.12 (matches DeepFilterNet's official support range)
  • Single GPU (cuda:0 fixed), CUDA 12.8 wheels
  • torch / torchaudio pinned to <2.9 — works around the removal of torchaudio.backend.common.AudioMetaData in 2.9, which breaks DeepFilterNet's imports
  • All four models (Whisper LID / VoxLingua107 / DeepFilterNet / Silero VAD) are downloaded automatically on first run (network required)

Usage

1) Specify input files

Edit the test_files list in main.py with your wav paths:

test_files = [
    "/path/to/audio1.wav",
    "/path/to/audio2.wav",
]

16kHz mono WAV is recommended; other sample rates are resampled internally.

2) Run

.venv/bin/python main.py

Flow:

  1. Pre-load all four models (so load time is excluded from measurements)
  2. Iterate test_files — call bench.run() per file
  3. Once all files are done, write output/timings.csv via bench.save_timings()

3) Outputs

Path Content
output/denoise/<stem>.wav Denoised audio (48kHz int16)
output/<stem>.csv Per-segment 4-way LID results (lang, prob)
output/timings.csv Per-file, 4-combination timings (long format)
/usr/service/logs/scenemaker/lid_bench.log Progress log (internal path — edit log.py)

The console prints only the per-file summary (duration / VAD segments / elapsed / agreement rates). Details are in the CSV files above.


Environment notes (for external git users)

  • If HF_HOME is unset, models download to the default cache (~/.cache/huggingface/).
  • conf.MODEL_DIR defaults to an internal path (/stg/models). Outside this environment, change it to your own path, or modify the code to use SpeechBrain's default cache.
  • The log path (DEFAULT_LOG_FILE in log.py) is also internal — change it to a path of your own.

About

LID model comparison — Whisper LID vs VoxLingua107 (raw vs denoise)

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages