Proof-of-Concept for benchmarking multilingual STT (Speech-to-Text) systems on real-world video content.
Compares two open-source ASR systems on Korean video content with mixed-language inserts (interviews, foreign clips):
- Whisper —
Systran/faster-whisper-large-v3(faster-whisper / CTranslate2 backend) - Qwen —
Qwen3-ASR-1.7B+Qwen3-ForcedAligner-0.6B
Both systems are scored by Gemini 3.5 Flash as judge — a single multimodal evaluator compares each segment's text against the source audio on a -3 ~ 3 scale.
raw wav
│
[1] denoise — DeepFilterNet v3 (atten_lim_db=-30)
[2] VAD — Silero VAD (raw audio)
[3] LID — Whisper detect_language (raw audio)
[4] ASR — Whisper / Qwen (denoised audio)
[5] post-filter — 5 gates (whisper) / script-based correction (qwen)
[6] evaluate — Gemini judge
[7] report — system × content aggregation
Strategy 2 — raw / denoised split: VAD + LID run on the raw audio (denoise distorts the LID signal), while ASR runs on the denoised audio (less hallucination). Verified by PoC: Whisper LID accuracy 95.2% on raw vs 93.4% on denoised.
Requirements: Python 3.11+, uv (0.6.17+), NVIDIA GPU (tested on RTX 4090 24GB).
# Whisper venv (also runs evaluate / report)
uv venv .venv --python 3.11
.venv/bin/uv sync
# Qwen venv (dependency-conflict isolation — uses python 3.12)
uv venv .venv-qwen --python 3.12
.venv-qwen/bin/uv pip install -r pyproject-qwen.tomlLocal model paths and API keys are configured in conf.py. The .env file requires:
GOOGLE_API_KEY=... # Gemini judge + HuggingFace download
Models used (paths in conf.py):
| Purpose | Model | Path / Source |
|---|---|---|
| Whisper ASR + LID | Systran/faster-whisper-large-v3 |
HF cache |
| Qwen ASR | Qwen3-ASR-1.7B |
local |
| Qwen Timestamp | Qwen3-ForcedAligner-0.6B |
local |
| Qwen LID | mobiuslabsgmbh/faster-whisper-large-v3-turbo |
HF cache |
| Denoise | DeepFilterNet v3 | pip package built-in |
| VAD | Silero VAD | torch.hub |
# 1. Whisper STT
.venv/bin/python main.py
# 2. Qwen STT (separate venv)
.venv-qwen/bin/python main_qwen.py
# 3. Gemini judge evaluation (whisper + qwen)
.venv/bin/python evaluate.py all
# 4. Aggregate report
.venv/bin/python report.pyInput audio paths are configured in the audio_files list at the bottom of each entry-point script (main.py / main_qwen.py / evaluate.py). Comment out files to run a subset.
Outputs land under output/:
output/
├── 1_denoise/<stem>.wav # DF cache (shared between systems)
├── whisper/
│ ├── 2_transcribe/<stem>.md # STT output
│ ├── evaluate/<stem>.csv # Gemini judge scores
│ └── timings.csv # duration / transcribe time / RTF
├── qwen/
│ └── (same structure)
└── report.csv # final system × content comparison
Benchmarked on 6 Korean video contents (~6h 20m total):
| Metric | Whisper | Qwen |
|---|---|---|
| Avg score (-3~3) | 2.71 | 2.63 |
| Usable subtitle rate (≥0) | 97.8% | 97.0% |
| Hallucination rate (-3) | 1.4% | 1.8% |
| Avg RTF | 0.042 | 0.032 |
| VRAM | ~3.5 GB | ~7 GB |
Whisper — more accurate overall, especially on documentary content (docu: Whisper -3 rate 1.6% vs Qwen 6.0%). Lower VRAM. Recommended as default.
Qwen — ~2× faster (batch processing). Slight edge on fast-paced commentary (baseball).
See output/report.csv for full per-content breakdown.
Five post-processing gates (in order of effect):
- VAD pre-filter — Silero VAD skips silence/BGM regions
- MIN_LOGPROB — drop segments with
avg_logprob < -1.0(hallucination catch-all) - LID_TRUST_PROB — if LID
prob < 0.5and non-Korean, force Korean (LID itself untrustworthy) - dual transcribe + MIN_DUAL_LOGPROB — short (<3s) non-Korean speech: run both ko and detected lang, pick higher logprob; drop if both
< -0.6 - Hangul char ratio gate — drop if chosen lang is
kobut Hangul ratio of the text is< 30%(catches Whisper outputting kana/hanja tokens in ko mode)
POC focuses on transcription quality comparison only. The following are intentionally not integrated:
- Video → WAV extraction (assumed as external preprocessing)
- Speaker diarization (PyAnnote)
- SRT / VTT output formatting
- Gemini correction post-processing
- FastAPI / HTTP API / job queue
- Multi-GPU distribution
MIT — see LICENSE.