Skip to content

fix(frontend): stop cmvn readers from crashing on blank lines or returning empty statistics - #3760

Merged
LauraGPT merged 1 commit into
modelscope:mainfrom
Lesereingrape:contrib/cmvn-blank-line
Oct 8, 2026
Merged

LauraGPT merged 1 commit into
modelscope:mainfrom
Lesereingrape:contrib/cmvn-blank-line

Conversation

@Lesereingrape

Copy link
Copy Markdown
Contributor

Summary

Both cmvn readers in the installed package walk the file line by line and index the result without checking it: line_item[0] on a blank line raises IndexError: list index out of range, and the unbounded one-line lookahead lines[i + 1] raises the same error when a <AddShift> / <Rescale> tag is the last line. Because the parse is a pure line-matching loop, a file whose tags do not start a line on its own (which Kaldi does emit — the whole <Nnet> on one line) matched nothing and load_cmvn returned empty statistics with no error at all; the failure then only surfaced on the first inference as RuntimeError: The size of tensor a (N) must match the size of tensor b (0).

Fixes #3759

Three-line change per reader (funasr/frontends/wav_frontend.py:15, funasr/frontends/default.py:390):

lines = [line for line in f.readlines() if line.split()]          # skip blank/whitespace-only lines
line_item = lines[i + 1].split() if i + 1 < len(lines) else []     # bounded lookahead
if line_item and line_item[0] == "<LearnRateCoef>":
...
if not means_list or not vars_list:
    raise ValueError(f"No <AddShift>/<Rescale> statistics found in cmvn file: {cmvn_file}")

Skipping blank lines first (rather than guarding inside the loop) also makes <AddShift> + blank line + <LearnRateCoef> parse, which a per-line guard would silently drop. The ValueError follows the convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148 names the offending cmvn file), and blank-line skipping matches how this repository reads other plain-text input (funasr/bin/realtime_ws.py:1363, funasr/utils/compute_det_ctc.py:59).

Type of change

  • Bug fix
  • Documentation
  • Example or demo
  • Runtime or deployment
  • Benchmark or evaluation
  • Model/training change

Validation

Measured on main @ 66d7a4c2, Python 3.13.7 / numpy 2.5.3 / torch 2.14.1+cpu, Windows, with PYTHONPATH pointed at the worktree (verified by printing funasr.__file__):

  • RED with only the two source files restored from origin/main and the new tests present: python -m pytest -q tests/test_numpy_compatibility.py -k cmvn → 2 failed, 3 passed.

  • GREEN with the fix: same command → 5 passed; whole file python -m pytest -q tests/test_numpy_compatibility.py → 9 passed, including the two pre-existing cmvn tests and test_frontend_applies_cmvn_and_preserves_padding.

  • Regression guard as a test of its own: test_cmvn_reader_loads_the_am_mvn_shipped_by_the_runtime asserts the am.mvn this repository ships under runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/ still loads as (2, 560) finite values — it passes on main too, so a file that parses correctly today provably does not change.

  • python -m compileall -q funasr/frontends tests/test_numpy_compatibility.py → exit 0.

  • Formatting: the repository is not black-clean at main (e.g. funasr/frontends/default.py docstring block at 392-395 is already reformatted by black --line-length=100), so black --check fails on both the pristine and the patched file. Checked instead that black --line-length=100 leaves every added line untouched — the only hunks it proposes in these files are pre-existing ones outside my diff.

  • python -m compileall funasr examples tests

  • Docs or links checked

  • Runtime/deployment command tested

User impact

Paraformer / SenseVoice users who point cmvn_file at a hand-edited or concatenated .mvn file — or at a Kaldi artifact written on a single line — currently get either a bare IndexError during frontend construction, or a frontend that builds fine and fails at the first model.generate() call. Both cases now either parse or raise an error that names the file.

Notes for reviewers

  • The same loop is duplicated in five more places under runtime/ (runtime/python/libtorch/, runtime/python/onnxruntime/, runtime/triton_gpu/, export_lfr_cmvn_pe_onnx.py). Those are standalone export/deployment helpers rather than the installed package, so this PR deliberately leaves them alone; the issue lists them.
  • Two of the three possible failure modes were silent or crash-only, so the ValueError is a behaviour change in one narrow case: a cmvn file that yielded empty statistics used to return them. Such a frontend could never produce features (it hit the RuntimeError above on the first call), so this only converts a downstream failure into an actionable one.
  • The practical limitation to note: a fork PR in this repository does not trigger workflow runs, so all evidence above is local, and the reproduction needs no model weights or downloads.
  • Prepared with an AI coding agent; a human maintainer at the PR author's side has not re-reviewed the diff beyond the checks listed above.

Both package cmvn readers index line_item[0] and lines[i + 1] without
checking, so a blank or whitespace-only line anywhere in the file
(trailing newline at EOF, a separator between sections) aborts frontend
construction with IndexError, and a file whose tags the walker does not
recognize returns empty statistics that only fail much later at inference.
Skip blank lines, bound the lookahead, and raise a ValueError naming the
file when neither <AddShift> nor <Rescale> parsed.

@LauraGPT LauraGPT left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head d0b9cc9 against current-main integration. New CMVN cases first fail on old runtime (2 failures / 3 passes), then all 5 selected cases pass. Independent checks cover 30 parser boundaries including empty, partial, truncated, blank/CRLF input and preserved values/dtypes. Replaying the shipped am.mvn with the reporter's blank-line edits and actual WavFrontend construction gives 6 failures / 3 passes before the fix, then 9 passes with exactly unchanged statistics. The current NumPy matrix runs pass 181 tests + 10 subtests in each lane; all 433 exported package Python files match the tested Git tree after restoring omitted setup.py in the test snapshot. Both approved head workflows now succeed; both NumPy artifacts show 123 entries, zero failures/errors/skips, exact head source and all three added regressions. The only main advance since runtime validation is #3764's independently validated SDK image links. No blocking findings. Compact single-line CMVN is rejected early, not newly supported; five standalone runtime/export helper copies remain unchanged. Evidence is Linux/Python3.12/CPU, not Windows hardware or checkpoint inference.

@LauraGPT
LauraGPT merged commit e7e6129 into modelscope:main Oct 8, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: cmvn readers raise IndexError on blank lines and return empty statistics without error

2 participants