Turn a video into a readable transcript — fillers, stutters and repetitions removed. Runs entirely on your machine (no upload, no API key, no LLM required), CPU-only, with a CLI and a local web GUI. Optional speaker diarization, ASS subtitles and burn-in.
中文文档 · Download a build · Report a bug
video ─▶ VAD ─▶ Whisper verbatim ASR ─▶ explainable clean-up rules ─▶ (optional) LLM polish
(word timestamps) └─▶ raw/clean .txt + .srt + .md + report.json
└─▶ optional: speaker labels, ASS, burn-in, tightened video
Every deletion is recorded with its timestamp in
report.json, so you can always audit why a word disappeared — unlike end-to-end "cleaned" ASR models that silently drop content.
- Features
- Why this exists
- Install
- Quickstart
- Output files
- Cleaning rules
- Requirements
- Benchmarks
- Project layout
- Roadmap
- Contributing
- License
| 🧹 Disfluency removal | fillers (um, uh, 嗯, 呃), stutters (我我我), adjacent repeats (把,把), phrase-level restarts, false starts (我,我,我觉得 → 我觉得), pause-flanked discourse markers (那个, you know) |
| 🎚️ Three aggression levels | --level 1/2/3 — conservative for meeting minutes, aggressive for scripted voice-over |
| 🧾 Two transcripts, always | raw (verbatim, for traceability) + clean (readable), plus a deletion report |
| 🗣️ Speaker diarization | optional, zero-torch (sherpa-onnx, ~35 MB models), 说话人N: labels + speakers.json |
| 💬 Subtitles | SRT for both tiers, ASS export (--ass), burn-in to video (--burn) |
| ✂️ Video tightening | --cut re-encodes the video with fillers physically removed |
| 🖥️ Two entry points | CLI + local web GUI (stdlib http.server, binds 127.0.0.1 only) |
| 🔒 Local-first | no upload, no telemetry, no API key; one optional network call to download model weights |
| 📦 Standalone builds | PyInstaller binaries on the Releases page (no Python needed) |
| 🧪 Tested | 33 pure-function tests (rules / rendering / subtitles / diarization glue), CI on 3 OS × py3.9 + py3.12 |
| Existing option | What it does not do |
|---|---|
| Cloud APIs (AssemblyAI / Deepgram / Rev) | only strip um/uh-style fillers — repetitions and repairs stay; audio must be uploaded |
| Consumer apps (Descript, 剪映, 飞书妙记, 通义听悟) | same filler-only cleaning; not batchable, scriptable or auditable |
| End-to-end "cleaned" ASR adapters | mostly English-only; aggressive deletion — they remove intentional repetition (numbers, emphasis) too |
| Training your own model | word-timestamp rules already cover most of it; the rest is better handled by an optional LLM pass — cheaper and explainable |
This project sits in between: verbatim ASR + explainable rules + optional LLM, for Chinese and English, with every removed word logged.
git clone https://github.com/Weidows/video2script && cd video2script
uv sync # core: CLI + GUI (creates .venv from uv.lock)
uv sync --extra diar # + speaker diarization (sherpa-onnx)
uv run video2script --helpuv.lock is committed, so everyone resolves the exact same dependency set —
uv sync --locked in CI fails rather than silently re-resolving.
Add --extra build when you want to rebuild the standalone binary, and
uv sync --no-install-project when you only need the dev tools.
In China you can point uv at a mirror without touching the lockfile:
UV_DEFAULT_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple uv sync.
python -m venv .venv && . .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e . # core: CLI + GUI
pip install -e ".[diar]" # + speaker diarization (sherpa-onnx)Or grab a standalone binary from Releases — no Python needed:
| Release asset | Platform | Notes |
|---|---|---|
video2script-windows.exe |
Windows | double-click = GUI, drop a file on it = CLI |
video2script-macos |
macOS | unsigned: right-click → Open the first time |
video2script-linux |
Linux | chmod +x first |
video2script-<version>-py3-none-any.whl |
any | pip install <wheel> |
One binary does both: run it with no arguments to open the GUI (http://127.0.0.1:8756), or pass a file to transcribe on the command line.
Model weights are downloaded on first use into ~/.cache/video2script/models
(override with V2S_MODEL_DIR). small ≈ 0.5 GB, medium ≈ 1.5 GB, diarization ≈ 35 MB.
If huggingface.co is unreachable the tool falls back to hf-mirror.com
(override with HF_ENDPOINT).
One command, two modes — no arguments starts the GUI, a file argument runs the CLI.
| Command | What happens |
|---|---|
video2script |
starts the local web GUI and prints the URL |
video2script meeting.mp4 … |
CLI transcription |
video2script --gui meeting.mp4 |
opens the GUI with that file pre-loaded |
video2script --cli meeting.mp4 |
force CLI (useful in scripts) |
video2script --version |
print version |
video2script meeting.mp4 # Chinese, small model, level 2
video2script meeting.mp4 --model medium # recommended for Chinese
video2script talk.mp4 --lang en --model medium
video2script meeting.mp4 --level 3 --cut # aggressive + filler-free video
video2script meeting.mp4 --diarize --ass --burn # speakers + ASS + burned-in subtitles
video2script meeting.mp4 --rewrite llm # optional LLM pass (needs an API key)
python -m video2script meeting.mp4 # works without installing entry points| Flag | Description |
|---|---|
--lang |
zh (default) / en / auto |
--model |
tiny / base / small (default) / medium / large-v3 |
--level |
1 fillers + in-word repeats · 2 (default) + pause-flanked markers · 3 all markers |
--cut |
also emit *_tight.mp4 with fillers physically cut (needs ffmpeg) |
--diarize |
speaker separation (needs [diar]); --num-speakers N, --diar-threshold |
--ass / --burn |
write clean.ass / burn subtitles into *_subtitled.mp4 (needs ffmpeg) |
--rewrite llm |
extra LLM pass (V2S_LLM_BASE / V2S_LLM_KEY / V2S_LLM_MODEL) |
--device, --compute-type |
e.g. --device cuda --compute-type float16 |
video2script # no arguments → starts the GUI (http://127.0.0.1:8756)
video2script --gui meeting.mp4 # open the GUI with a file already loaded
video2script --gui --open # and open the browser automaticallyDrag a file in (or pre-load it with --gui file.mp4) and you get:
- a preview card with the actual video/audio player, size, duration and source path — the player streams from the local server with HTTP Range, so you can scrub before transcribing;
- live feedback while it runs: a progress bar with stage names (
asr 47%,diarizing…,burning subtitles…), the running verbatim transcript appearing segment by segment, and every artifact (.txt/.srt/.md) becoming downloadable the moment it is written; - a complete log — nothing is truncated, newest lines auto-scroll, plus copy/clear buttons;
- the final verbatim vs. cleaned comparison side by side, removal counts, and Open output folder (with "copy path" as a fallback).
The server binds 127.0.0.1 only; nothing leaves your machine.
video2script-gui is kept as an alias for video2script --gui (existing scripts keep working).
from video2script import Options, run
res = run("meeting.mp4",
Options(lang="zh", model="medium", level=2, diarize=True, ass=True),
on_event=lambda kind, data: print(kind, data))
print(res.clean_text, res.counts, res.speakers)uv sync --extra build # or: pip install -e ".[build]"
uv run python scripts/make_icon.py # regenerate the icon (already committed, optional)
uv run python scripts/build_exe.py # dist/video2script(.exe) — one binary, CLI + GUIModel weights are not bundled — the first run downloads them as usual. PNG/ICO icons live in
src/video2script/assets/ and are embedded in the binary (window icon + GUI favicon).
| File | Contents |
|---|---|
raw.txt / raw.srt |
verbatim transcript (fillers included) — the audit trail |
clean.txt / clean.srt / clean.md |
cleaned transcript; Markdown is split by pauses and speakers |
report.json |
every removal: timestamp, word, reason (filler/repeat/discourse/phrase_repeat/partial) |
speakers.json |
speaker turns and per-segment assignment (with --diarize) |
clean.ass |
subtitle file (with --ass/--burn) |
clean.llm.txt |
LLM-polished text (with --rewrite llm) |
*_tight.mp4 |
video with fillers cut out (with --cut) |
*_subtitled.mp4 |
video with burned-in subtitles (with --burn) |
Order matters — see src/video2script/clean.py.
- Punctuation re-attachment — Whisper sometimes emits punctuation as standalone tokens
(
对,对); they are merged back so repeat detection still fires. - Pure fillers —
嗯 呃 额 唔 诶 唉 哦 噢 啊 …/um uh erm hmm … - Partial words —
我-,wou- - In-word repeats —
我我我→我, protected by a reduplication whitelist (谢谢,看看,刚刚,妈妈, …) - Adjacent repeats —
把,把语音识别→把语音识别(dangling comma cleaned up) - Phrase-level restarts — up to 6 words repeated as a block, second copy removed
- Cross-token stutters —
我,我,我觉得→我觉得,那,那我说一下→那我说一下 - Discourse markers (level ≥ 2) —
那个 / 就是 / 然后 / 你知道 / you know / I mean, removed only when flanked by > 0.18 s of silence; level 3 removes them unconditionally - Tidy-up — dangling punctuation removed, rare CJK punctuation (
﹔﹑﹕) normalized
Word matching uses Unicode character categories, not a punctuation whitelist — Whisper occasionally
emits rare punctuation such as ﹔, which a whitelist would miss (making 那﹔那 undetectable).
Customize terminology by editing the word sets at the top of clean.py
(ZH_INTERJ / ZH_DISCOURSE / ZH_KEEP, EN_*) — no code changes needed.
| Minimum | Comfortable | |
|---|---|---|
| CPU | 4 cores (int8 quantized, CPU-only is supported) | 8+ cores |
| RAM | ~605 MB peak with small |
~1.5 GB peak with medium, 4 GB+ advised |
| Disk | ~0.5 GB (deps + small weights) |
3 GB+ (incl. large-v3) |
| GPU | not required | any CUDA GPU, --device cuda is 5–20× faster |
| Python | 3.9+ | 3.11 / 3.12 |
| Stage | Implementation | Required? |
|---|---|---|
| Demux/decode | PyAV (av, bundles FFmpeg libs) |
yes — no system ffmpeg needed |
| Speech recognition | faster-whisper (CTranslate2 Whisper, word timestamps + VAD) |
yes — local inference |
| Cleaning | this project, pure Python | yes — milliseconds |
| Diarization | sherpa-onnx (pyannote-seg 3.0 + 3D-Speaker CAM++, fast clustering) |
optional, [diar] |
| Subtitle burn-in / cutting | system ffmpeg | optional, --burn / --cut |
| LLM polish | any OpenAI-compatible endpoint | optional, --rewrite llm |
| GUI | Python stdlib http.server |
optional, zero extra deps |
No LLM is required. The default pipeline only uses the dedicated Whisper ASR models
(small ≈ 244 M params, large-v3 ≈ 1.55 B), calls no LLM, needs no API key and no network access
(other than the one-time weight download). Diarization runs on onnxruntime — no torch, no uploads.
Measured locally: 16-core CPU, int8, no GPU, 21-second Chinese clip (samples/say_zh.mp4).
| Model | Transcribe + clean | Peak RSS (process tree) |
|---|---|---|
small |
8.3 s | 605 MB |
medium |
24.9 s | 1431 MB |
medium + --diarize |
25 s + 3.7 s | ≈ medium |
Quality (same clip, reference text in samples/say_zh.txt):
| Text | |
|---|---|
raw |
呃、那个,我今天想讲一下这个,嗯,这个项目的一个,一个,就是进度问题。 |
--level 2 |
我今天想讲一下这个项目的一个就是进度问题。 |
--level 3 |
我今天想讲一下项目的一个进度问题。 |
The intended sentence was "我今天想讲一下这个项目的一个进度问题": level 2 keeps one extra 就是
(not flanked by enough silence), level 3 also drops the meaningful 这个. That is the
conservative/aggressive trade-off — use 2 for minutes, 3 for voice-over, or add the LLM pass.
Two-speaker sample (samples/say_two_speakers.mp4):
说话人1:那个,我是产品经理,我今天想讲一下这个项目的进度问题。
说话人2:好的,那我说一下技术这边的情况,我们上周把语音识别的模块做完了。
说话人1:那下周是不是可以开始测试了?
说话人2:对下周我们,我觉得可以开始测试。
uv run pytest -q # 85+ tests, pure functions, no model download
uv run python scripts/gui_smoke.py # real server end-to-end (preview Range / progress / artifacts / reveal)src/video2script/
├── asr.py # faster-whisper wrapper (word timestamps, VAD, progress callbacks)
├── clean.py # the cleaning rules + word lists (the interesting part)
├── render.py # subtitle blocking, SRT/Markdown, cut intervals
├── subtitles.py # ASS generation + burn-in
├── diarize.py # sherpa-onnx diarization, model download, speaker assignment
├── pipeline.py # orchestration: transcribe → clean → diarize → render
├── main.py # unified entry: no args → GUI, file args → CLI
├── cli.py # CLI parser (`video2script 文件.mp4`)
├── gui.py # `video2script --gui` (stdlib http.server + Range preview streaming)
├── config.py # paths, HF mirror fallback, ffmpeg discovery, UTF-8 stdio
└── assets/ # gui.html (the interface) / icon.png / icon.ico
- Speaker diarization (sherpa-onnx, torch-free)
- ASS export + subtitle burn-in
- Standalone binaries (PyInstaller + release CI)
- Domain word lists (
--wordlist) - CapCut / Final Cut project export
- Batch queue mode with
large-v3+ GPU
Issues and PRs are welcome. Please run pip install -e ".[dev]" && pytest -q before opening a PR,
and keep new cleaning rules accompanied by a pure-function test in tests/test_clean.py.
See CONTRIBUTING.md for details.
Source-available, non-commercial. Licensed under the PolyForm Noncommercial License 1.0.0:
- ✅ personal use, study, research, hobby projects, experiments
- ✅ use by non-profits, schools, public research and government institutions
- ✅ modify and redistribute for those purposes (keep the license and the
Required Noticeline) - ❌ commercial use — including internal business use, SaaS, and embedding in a paid product
Commercial licensing is available separately — open an issue at
https://github.com/Weidows/video2script/issues or contact the maintainer
(@Weidows). The Required Notice line lives in NOTICE.
GitHub's license sidebar may show Other / NOASSERTION: PolyForm Noncommercial is deliberately not an OSI-approved license, so GitHub does not auto-detect it. The authoritative text is LICENSE.
The license covers this repository's own code. Third-party components keep their own licenses and are not relicensed here:
| Component | License |
|---|---|
| faster-whisper, CTranslate2 | MIT |
| Whisper model weights (OpenAI, distributed by Systran) | MIT |
| PyAV | BSD-3-Clause |
| onnxruntime | MIT |
| sherpa-onnx | Apache-2.0 |
| Diarization models (pyannote segmentation, 3D-Speaker CAM++) | see upstream repositories |
| NumPy, tokenizers, huggingface_hub | BSD-3-Clause / Apache-2.0 |
If you use this project commercially, make sure you are also compliant with the licenses above.