Skip to content

About

Whisper-Large-v3 vs Qwen3-Audio model comparison

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

poc-stt-bench

Proof-of-Concept for benchmarking multilingual STT (Speech-to-Text) systems on real-world video content.

한국어 README

Overview

Compares two open-source ASR systems on Korean video content with mixed-language inserts (interviews, foreign clips):

  • Whisper — Systran/faster-whisper-large-v3 (faster-whisper / CTranslate2 backend)
  • Qwen — Qwen3-ASR-1.7B + Qwen3-ForcedAligner-0.6B

Both systems are scored by Gemini 3.5 Flash as judge — a single multimodal evaluator compares each segment's text against the source audio on a -3 ~ 3 scale.

Pipeline

raw wav
   │
[1] denoise — DeepFilterNet v3 (atten_lim_db=-30)
[2] VAD     — Silero VAD (raw audio)
[3] LID     — Whisper detect_language (raw audio)
[4] ASR     — Whisper / Qwen (denoised audio)
[5] post-filter — 5 gates (whisper) / script-based correction (qwen)
[6] evaluate    — Gemini judge
[7] report      — system × content aggregation

Strategy 2 — raw / denoised split: VAD + LID run on the raw audio (denoise distorts the LID signal), while ASR runs on the denoised audio (less hallucination). Verified by PoC: Whisper LID accuracy 95.2% on raw vs 93.4% on denoised.

Setup

Requirements: Python 3.11+, uv (0.6.17+), NVIDIA GPU (tested on RTX 4090 24GB).

# Whisper venv (also runs evaluate / report)
uv venv .venv --python 3.11
.venv/bin/uv sync

# Qwen venv (dependency-conflict isolation — uses python 3.12)
uv venv .venv-qwen --python 3.12
.venv-qwen/bin/uv pip install -r pyproject-qwen.toml

Local model paths and API keys are configured in conf.py. The .env file requires:

GOOGLE_API_KEY=...        # Gemini judge + HuggingFace download

Models used (paths in conf.py):

Purpose Model Path / Source
Whisper ASR + LID Systran/faster-whisper-large-v3 HF cache
Qwen ASR Qwen3-ASR-1.7B local
Qwen Timestamp Qwen3-ForcedAligner-0.6B local
Qwen LID mobiuslabsgmbh/faster-whisper-large-v3-turbo HF cache
Denoise DeepFilterNet v3 pip package built-in
VAD Silero VAD torch.hub

Usage

# 1. Whisper STT
.venv/bin/python main.py

# 2. Qwen STT (separate venv)
.venv-qwen/bin/python main_qwen.py

# 3. Gemini judge evaluation (whisper + qwen)
.venv/bin/python evaluate.py all

# 4. Aggregate report
.venv/bin/python report.py

Input audio paths are configured in the audio_files list at the bottom of each entry-point script (main.py / main_qwen.py / evaluate.py). Comment out files to run a subset.

Outputs land under output/:

output/
├── 1_denoise/<stem>.wav         # DF cache (shared between systems)
├── whisper/
│   ├── 2_transcribe/<stem>.md   # STT output
│   ├── evaluate/<stem>.csv      # Gemini judge scores
│   └── timings.csv              # duration / transcribe time / RTF
├── qwen/
│   └── (same structure)
└── report.csv                   # final system × content comparison

Results

Benchmarked on 6 Korean video contents (~6h 20m total):

Metric Whisper Qwen
Avg score (-3~3) 2.71 2.63
Usable subtitle rate (≥0) 97.8% 97.0%
Hallucination rate (-3) 1.4% 1.8%
Avg RTF 0.042 0.032
VRAM ~3.5 GB ~7 GB

Whisper — more accurate overall, especially on documentary content (docu: Whisper -3 rate 1.6% vs Qwen 6.0%). Lower VRAM. Recommended as default.

Qwen — ~2× faster (batch processing). Slight edge on fast-paced commentary (baseball).

See output/report.csv for full per-content breakdown.

Hallucination Handling (Whisper)

Five post-processing gates (in order of effect):

  1. VAD pre-filter — Silero VAD skips silence/BGM regions
  2. MIN_LOGPROB — drop segments with avg_logprob < -1.0 (hallucination catch-all)
  3. LID_TRUST_PROB — if LID prob < 0.5 and non-Korean, force Korean (LID itself untrustworthy)
  4. dual transcribe + MIN_DUAL_LOGPROB — short (<3s) non-Korean speech: run both ko and detected lang, pick higher logprob; drop if both < -0.6
  5. Hangul char ratio gate — drop if chosen lang is ko but Hangul ratio of the text is < 30% (catches Whisper outputting kana/hanja tokens in ko mode)

Out of Scope

POC focuses on transcription quality comparison only. The following are intentionally not integrated:

  • Video → WAV extraction (assumed as external preprocessing)
  • Speaker diarization (PyAnnote)
  • SRT / VTT output formatting
  • Gemini correction post-processing
  • FastAPI / HTTP API / job queue
  • Multi-GPU distribution

License

MIT — see LICENSE.

About

Whisper-Large-v3 vs Qwen3-Audio model comparison

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages