Skip to content

Repository files navigation

video2script

video2script icon

CI Release License: PolyForm Noncommercial Python Platform

Turn a video into a readable transcript — fillers, stutters and repetitions removed. Runs entirely on your machine (no upload, no API key, no LLM required), CPU-only, with a CLI and a local web GUI. Optional speaker diarization, ASS subtitles and burn-in.

中文文档 · Download a build · Report a bug

video ─▶ VAD ─▶ Whisper verbatim ASR ─▶ explainable clean-up rules ─▶ (optional) LLM polish
  (word timestamps)                     └─▶ raw/clean .txt + .srt + .md + report.json
                                        └─▶ optional: speaker labels, ASS, burn-in, tightened video

Every deletion is recorded with its timestamp in report.json, so you can always audit why a word disappeared — unlike end-to-end "cleaned" ASR models that silently drop content.

Contents

Features

🧹 Disfluency removal fillers (um, uh, 嗯, 呃), stutters (我我我), adjacent repeats (把,把), phrase-level restarts, false starts (我,我,我觉得 → 我觉得), pause-flanked discourse markers (那个, you know)
🎚️ Three aggression levels --level 1/2/3 — conservative for meeting minutes, aggressive for scripted voice-over
🧾 Two transcripts, always raw (verbatim, for traceability) + clean (readable), plus a deletion report
🗣️ Speaker diarization optional, zero-torch (sherpa-onnx, ~35 MB models), 说话人N: labels + speakers.json
💬 Subtitles SRT for both tiers, ASS export (--ass), burn-in to video (--burn)
✂️ Video tightening --cut re-encodes the video with fillers physically removed
🖥️ Two entry points CLI + local web GUI (stdlib http.server, binds 127.0.0.1 only)
🔒 Local-first no upload, no telemetry, no API key; one optional network call to download model weights
📦 Standalone builds PyInstaller binaries on the Releases page (no Python needed)
🧪 Tested 33 pure-function tests (rules / rendering / subtitles / diarization glue), CI on 3 OS × py3.9 + py3.12

Why this exists

Existing option What it does not do
Cloud APIs (AssemblyAI / Deepgram / Rev) only strip um/uh-style fillers — repetitions and repairs stay; audio must be uploaded
Consumer apps (Descript, 剪映, 飞书妙记, 通义听悟) same filler-only cleaning; not batchable, scriptable or auditable
End-to-end "cleaned" ASR adapters mostly English-only; aggressive deletion — they remove intentional repetition (numbers, emphasis) too
Training your own model word-timestamp rules already cover most of it; the rest is better handled by an optional LLM pass — cheaper and explainable

This project sits in between: verbatim ASR + explainable rules + optional LLM, for Chinese and English, with every removed word logged.

Install

uv (recommended — lockfile included)

git clone https://github.com/Weidows/video2script && cd video2script
uv sync                      # core: CLI + GUI (creates .venv from uv.lock)
uv sync --extra diar         # + speaker diarization (sherpa-onnx)
uv run video2script --help

uv.lock is committed, so everyone resolves the exact same dependency set — uv sync --locked in CI fails rather than silently re-resolving. Add --extra build when you want to rebuild the standalone binary, and uv sync --no-install-project when you only need the dev tools. In China you can point uv at a mirror without touching the lockfile: UV_DEFAULT_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple uv sync.

pip

python -m venv .venv && . .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -e .                                    # core: CLI + GUI
pip install -e ".[diar]"                            # + speaker diarization (sherpa-onnx)

Or grab a standalone binary from Releases — no Python needed:

Release asset Platform Notes
video2script-windows.exe Windows double-click = GUI, drop a file on it = CLI
video2script-macos macOS unsigned: right-click → Open the first time
video2script-linux Linux chmod +x first
video2script-<version>-py3-none-any.whl any pip install <wheel>

One binary does both: run it with no arguments to open the GUI (http://127.0.0.1:8756), or pass a file to transcribe on the command line.

Model weights are downloaded on first use into ~/.cache/video2script/models (override with V2S_MODEL_DIR). small ≈ 0.5 GB, medium ≈ 1.5 GB, diarization ≈ 35 MB. If huggingface.co is unreachable the tool falls back to hf-mirror.com (override with HF_ENDPOINT).

Quickstart

One command, two modes — no arguments starts the GUI, a file argument runs the CLI.

Command What happens
video2script starts the local web GUI and prints the URL
video2script meeting.mp4 … CLI transcription
video2script --gui meeting.mp4 opens the GUI with that file pre-loaded
video2script --cli meeting.mp4 force CLI (useful in scripts)
video2script --version print version

CLI

video2script meeting.mp4                          # Chinese, small model, level 2
video2script meeting.mp4 --model medium           # recommended for Chinese
video2script talk.mp4 --lang en --model medium
video2script meeting.mp4 --level 3 --cut          # aggressive + filler-free video
video2script meeting.mp4 --diarize --ass --burn   # speakers + ASS + burned-in subtitles
video2script meeting.mp4 --rewrite llm            # optional LLM pass (needs an API key)
python -m video2script meeting.mp4                # works without installing entry points
Flag Description
--lang zh (default) / en / auto
--model tiny / base / small (default) / medium / large-v3
--level 1 fillers + in-word repeats · 2 (default) + pause-flanked markers · 3 all markers
--cut also emit *_tight.mp4 with fillers physically cut (needs ffmpeg)
--diarize speaker separation (needs [diar]); --num-speakers N, --diar-threshold
--ass / --burn write clean.ass / burn subtitles into *_subtitled.mp4 (needs ffmpeg)
--rewrite llm extra LLM pass (V2S_LLM_BASE / V2S_LLM_KEY / V2S_LLM_MODEL)
--device, --compute-type e.g. --device cuda --compute-type float16

GUI

video2script                      # no arguments → starts the GUI (http://127.0.0.1:8756)
video2script --gui meeting.mp4    # open the GUI with a file already loaded
video2script --gui --open         # and open the browser automatically

Drag a file in (or pre-load it with --gui file.mp4) and you get:

  • a preview card with the actual video/audio player, size, duration and source path — the player streams from the local server with HTTP Range, so you can scrub before transcribing;
  • live feedback while it runs: a progress bar with stage names (asr 47%, diarizing…, burning subtitles…), the running verbatim transcript appearing segment by segment, and every artifact (.txt/.srt/.md) becoming downloadable the moment it is written;
  • a complete log — nothing is truncated, newest lines auto-scroll, plus copy/clear buttons;
  • the final verbatim vs. cleaned comparison side by side, removal counts, and Open output folder (with "copy path" as a fallback).

The server binds 127.0.0.1 only; nothing leaves your machine.

video2script-gui is kept as an alias for video2script --gui (existing scripts keep working).

Library

from video2script import Options, run

res = run("meeting.mp4",
          Options(lang="zh", model="medium", level=2, diarize=True, ass=True),
          on_event=lambda kind, data: print(kind, data))
print(res.clean_text, res.counts, res.speakers)

Standalone build

uv sync --extra build              # or: pip install -e ".[build]"
uv run python scripts/make_icon.py  # regenerate the icon (already committed, optional)
uv run python scripts/build_exe.py  # dist/video2script(.exe) — one binary, CLI + GUI

Model weights are not bundled — the first run downloads them as usual. PNG/ICO icons live in src/video2script/assets/ and are embedded in the binary (window icon + GUI favicon).

Output files

File Contents
raw.txt / raw.srt verbatim transcript (fillers included) — the audit trail
clean.txt / clean.srt / clean.md cleaned transcript; Markdown is split by pauses and speakers
report.json every removal: timestamp, word, reason (filler/repeat/discourse/phrase_repeat/partial)
speakers.json speaker turns and per-segment assignment (with --diarize)
clean.ass subtitle file (with --ass/--burn)
clean.llm.txt LLM-polished text (with --rewrite llm)
*_tight.mp4 video with fillers cut out (with --cut)
*_subtitled.mp4 video with burned-in subtitles (with --burn)

Cleaning rules

Order matters — see src/video2script/clean.py.

  1. Punctuation re-attachment — Whisper sometimes emits punctuation as standalone tokens (对 , 对); they are merged back so repeat detection still fires.
  2. Pure fillers — 嗯 呃 额 唔 诶 唉 哦 噢 啊 … / um uh erm hmm …
  3. Partial words — 我-, wou-
  4. In-word repeats — 我我我 → 我, protected by a reduplication whitelist (谢谢, 看看, 刚刚, 妈妈, …)
  5. Adjacent repeats — 把,把语音识别 → 把语音识别 (dangling comma cleaned up)
  6. Phrase-level restarts — up to 6 words repeated as a block, second copy removed
  7. Cross-token stutters — 我,我,我觉得 → 我觉得, 那,那我说一下 → 那我说一下
  8. Discourse markers (level ≥ 2) — 那个 / 就是 / 然后 / 你知道 / you know / I mean, removed only when flanked by > 0.18 s of silence; level 3 removes them unconditionally
  9. Tidy-up — dangling punctuation removed, rare CJK punctuation (﹔﹑﹕) normalized

Word matching uses Unicode character categories, not a punctuation whitelist — Whisper occasionally emits rare punctuation such as ﹔, which a whitelist would miss (making 那﹔那 undetectable).

Customize terminology by editing the word sets at the top of clean.py (ZH_INTERJ / ZH_DISCOURSE / ZH_KEEP, EN_*) — no code changes needed.

Requirements

Minimum Comfortable
CPU 4 cores (int8 quantized, CPU-only is supported) 8+ cores
RAM ~605 MB peak with small ~1.5 GB peak with medium, 4 GB+ advised
Disk ~0.5 GB (deps + small weights) 3 GB+ (incl. large-v3)
GPU not required any CUDA GPU, --device cuda is 5–20× faster
Python 3.9+ 3.11 / 3.12

What runs at runtime

Stage Implementation Required?
Demux/decode PyAV (av, bundles FFmpeg libs) yes — no system ffmpeg needed
Speech recognition faster-whisper (CTranslate2 Whisper, word timestamps + VAD) yes — local inference
Cleaning this project, pure Python yes — milliseconds
Diarization sherpa-onnx (pyannote-seg 3.0 + 3D-Speaker CAM++, fast clustering) optional, [diar]
Subtitle burn-in / cutting system ffmpeg optional, --burn / --cut
LLM polish any OpenAI-compatible endpoint optional, --rewrite llm
GUI Python stdlib http.server optional, zero extra deps

No LLM is required. The default pipeline only uses the dedicated Whisper ASR models (small ≈ 244 M params, large-v3 ≈ 1.55 B), calls no LLM, needs no API key and no network access (other than the one-time weight download). Diarization runs on onnxruntime — no torch, no uploads.

Benchmarks

Measured locally: 16-core CPU, int8, no GPU, 21-second Chinese clip (samples/say_zh.mp4).

Model Transcribe + clean Peak RSS (process tree)
small 8.3 s 605 MB
medium 24.9 s 1431 MB
medium + --diarize 25 s + 3.7 s ≈ medium

Quality (same clip, reference text in samples/say_zh.txt):

Text
raw 呃、那个,我今天想讲一下这个,嗯,这个项目的一个,一个,就是进度问题。
--level 2 我今天想讲一下这个项目的一个就是进度问题。
--level 3 我今天想讲一下项目的一个进度问题。

The intended sentence was "我今天想讲一下这个项目的一个进度问题": level 2 keeps one extra 就是 (not flanked by enough silence), level 3 also drops the meaningful 这个. That is the conservative/aggressive trade-off — use 2 for minutes, 3 for voice-over, or add the LLM pass.

Two-speaker sample (samples/say_two_speakers.mp4):

说话人1:那个,我是产品经理,我今天想讲一下这个项目的进度问题。
说话人2:好的,那我说一下技术这边的情况,我们上周把语音识别的模块做完了。
说话人1:那下周是不是可以开始测试了?
说话人2:对下周我们,我觉得可以开始测试。
uv run pytest -q                    # 85+ tests, pure functions, no model download
uv run python scripts/gui_smoke.py  # real server end-to-end (preview Range / progress / artifacts / reveal)

Project layout

src/video2script/
├── asr.py         # faster-whisper wrapper (word timestamps, VAD, progress callbacks)
├── clean.py       # the cleaning rules + word lists (the interesting part)
├── render.py      # subtitle blocking, SRT/Markdown, cut intervals
├── subtitles.py   # ASS generation + burn-in
├── diarize.py     # sherpa-onnx diarization, model download, speaker assignment
├── pipeline.py    # orchestration: transcribe → clean → diarize → render
├── main.py        # unified entry: no args → GUI, file args → CLI
├── cli.py         # CLI parser (`video2script 文件.mp4`)
├── gui.py         # `video2script --gui` (stdlib http.server + Range preview streaming)
├── config.py      # paths, HF mirror fallback, ffmpeg discovery, UTF-8 stdio
└── assets/        # gui.html (the interface) / icon.png / icon.ico

Roadmap

  • Speaker diarization (sherpa-onnx, torch-free)
  • ASS export + subtitle burn-in
  • Standalone binaries (PyInstaller + release CI)
  • Domain word lists (--wordlist)
  • CapCut / Final Cut project export
  • Batch queue mode with large-v3 + GPU

Contributing

Issues and PRs are welcome. Please run pip install -e ".[dev]" && pytest -q before opening a PR, and keep new cleaning rules accompanied by a pure-function test in tests/test_clean.py. See CONTRIBUTING.md for details.

License

Source-available, non-commercial. Licensed under the PolyForm Noncommercial License 1.0.0:

  • ✅ personal use, study, research, hobby projects, experiments
  • ✅ use by non-profits, schools, public research and government institutions
  • ✅ modify and redistribute for those purposes (keep the license and the Required Notice line)
  • ❌ commercial use — including internal business use, SaaS, and embedding in a paid product

Commercial licensing is available separately — open an issue at https://github.com/Weidows/video2script/issues or contact the maintainer (@Weidows). The Required Notice line lives in NOTICE.

GitHub's license sidebar may show Other / NOASSERTION: PolyForm Noncommercial is deliberately not an OSI-approved license, so GitHub does not auto-detect it. The authoritative text is LICENSE.

The license covers this repository's own code. Third-party components keep their own licenses and are not relicensed here:

Component License
faster-whisper, CTranslate2 MIT
Whisper model weights (OpenAI, distributed by Systran) MIT
PyAV BSD-3-Clause
onnxruntime MIT
sherpa-onnx Apache-2.0
Diarization models (pyannote segmentation, 3D-Speaker CAM++) see upstream repositories
NumPy, tokenizers, huggingface_hub BSD-3-Clause / Apache-2.0

If you use this project commercially, make sure you are also compliant with the licenses above.

About

视频→文稿:自动剔除语气词/结巴/重复词(CLI + 本地网页 GUI,CPU 可跑,可选说话人分离/字幕/剪片)· Turn videos into clean transcripts — fillers, stutters and repetitions removed

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages