Skip to content

Bug: cmvn readers raise IndexError on blank lines and return empty statistics without errorΒ #3759

Description

@Lesereingrape

πŸ› Bug

FunASR has two cmvn readers in the installed package, and both walk the file line by line while indexing the result without any check:

  • funasr/frontends/wav_frontend.py:15 load_cmvn β€” used by WavFrontend / WavFrontendOnline, i.e. the frontend of the Paraformer and SenseVoice models most people load first;
  • funasr/frontends/default.py:390 MultiChannelFrontend._load_cmvn β€” used by MultiChannelFrontend (emotion2vec / campplus).
    for i in range(len(lines)):
        line_item = lines[i].split()
        if line_item[0] == "<AddShift>":          # blank line -> IndexError
            line_item = lines[i + 1].split()      # tag on the last line -> IndexError
            if line_item[0] == "<LearnRateCoef>":

Two distinct symptoms come out of this, both measured on main @ 66d7a4c2:

cmvn_file content wav_frontend.load_cmvn MultiChannelFrontend._load_cmvn
runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn (a real artifact this repository ships) OK, shape (2, 560) OK, shape (560,)
the same file plus one trailing blank line IndexError: list index out of range IndexError: list index out of range
a blank (or whitespace-only) line between <AddShift> and its <LearnRateCoef> line IndexError: list index out of range IndexError: list index out of range
<AddShift> 2 2 as the last line (truncated file) IndexError: list index out of range IndexError: list index out of range
statistics present but the tags do not start a line on their own, e.g. the whole <Nnet> on one line returns a tensor of shape (2, 0) β€” no error returns two empty arrays β€” no error

The first four crash while the frontend is being constructed, so AutoModel(...) or a training run dies with a bare IndexError that never mentions the cmvn file. The fifth is worse because it is silent: load_cmvn returns empty statistics, WavFrontend builds without complaint, and the failure only appears at the first inference.

To Reproduce

  1. Install with: pip install -e . (source checkout of main @ 66d7a4c2)
  2. Copy the cmvn artifact the repository already ships and append a blank line:
    cp runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn /tmp/am.mvn && printf '\n' >> /tmp/am.mvn
  3. Run the code sample below.
  4. See error: IndexError: list index out of range

Code sample

from funasr.frontends.wav_frontend import WavFrontend

WavFrontend(cmvn_file="am.mvn", fs=16000, win_length=400, hop_length=160, n_mels=80)
# IndexError: list index out of range   -- only when the file has an extra blank line;
# without the blank line the same call succeeds.

# silent variant: tags do not start a line on their own
from funasr.frontends.wav_frontend import load_cmvn
open("compact.mvn", "w").write(
    "<Nnet> <AddShift> <LearnRateCoef> 1 [ -1.0 -2.0 ] </LearnRateCoef> "
    "<Rescale> <LearnRateCoef> 1 [ 0.5 0.5 ] </LearnRateCoef> </Nnet>\n"
)
cmvn = load_cmvn("compact.mvn")   # shape (2, 0), no error

Measured with n_mels=3 on the compact file, the frontend then builds fine and the first input_feats += self.mean / apply_cmvn call fails at inference time:

RuntimeError: The size of tensor a (3) must match the size of tensor b (0) at non-singleton dimension 1

MultiChannelFrontend behaves the same way (input_feats += self.mean against an empty registered buffer, funasr/frontends/default.py:355).

Expected behavior

A cmvn file that contains valid statistics should load regardless of blank or whitespace-only separator lines, and a file that yields no statistics should raise an error naming the file instead of returning empty arrays. Files that parse correctly today must keep parsing identically β€” the am.mvn above is the regression case.

Error logs

  File ".../funasr/frontends/wav_frontend.py", line 27, in load_cmvn
    if line_item[0] == "<AddShift>":
IndexError: list index out of range

Environment

  • OS: Windows 11 (also reproduces on Linux β€” the loop is platform-independent)
  • Python version: 3.13.7
  • FunASR version: source, main @ 66d7a4c2
  • PyTorch / torchaudio version: 2.14.1+cpu / matching
  • numpy: 2.5.3
  • Install method: source (pip install -e .)
  • Device: cpu
  • No GPU, no model download needed β€” the reproduction uses the cmvn text file the repository already ships.

Audio details

Not applicable: no audio is involved, the defect is in the plain-text cmvn parser.

Suggested fix

For the two package readers:

  1. drop blank/whitespace-only lines before pairing a section tag with the line below it β€” that also makes <AddShift> + blank line + <LearnRateCoef> parse, and matches how this repository already reads other plain-text inputs (funasr/bin/realtime_ws.py:1363 skips blank hotword lines, funasr/utils/compute_det_ctc.py:59 skips blank stat lines);
  2. bound the one-line lookahead so a truncated file cannot raise IndexError;
  3. raise a ValueError naming the file when neither <AddShift> nor <Rescale> produced statistics, instead of handing back empty arrays β€” the same "say which cmvn file is wrong" convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148 raises FileNotFoundError("cmvn file not exits")).

Five more copies of the same loop exist under runtime/ (runtime/python/libtorch/, runtime/python/onnxruntime/, runtime/triton_gpu/, export_lfr_cmvn_pe_onnx.py); they are standalone export/deployment helpers rather than the installed package, so I would scope any PR to the two package readers.

Filed with an AI coding agent, based on the measurements above that I ran locally on this machine.

Activity

  1. LauraGPT commented on Oct 8, 2026

    @LauraGPT
    Collaborator

    Thanks for the exact file-based reproduction and focused PR. #3760 is now merged in e7e6129.

    I replayed the repository's shipped am.mvn using your trailing-blank and section-separator edits, including actual WavFrontend construction with the reported arguments: old runtime gives 6 failures / 3 passes; the candidate passes all 9 checks. Both readers preserve the original statistics exactly (float32 tensor for load_cmvn, float64 arrays for MultiChannelFrontend). Another 30 parser boundaries pass, including empty, partial and truncated files. The current expanded NumPy 1.26.4 / 2.4.0 matrix passes 181 tests + 10 subtests in each lane; both hosted head artifacts pass 123 entries and include all three new regressions.

    Scope remains the two installed-package readers. Compact single-line CMVN now raises an early file-naming ValueError rather than returning empty statistics; this change does not add compact-format parsing. The five standalone runtime/export helper copies are unchanged.

    The Fixes clause auto-closed this issue on merge. It is now open with needs feedback under our lifecycle policy while you confirm the source fix on your original Windows / Python 3.13.7 / NumPy 2.5.3 setup. Our validation used Linux / Python 3.12 / CPU and the shipped text fixture, not a real Windows environment or checkpoint inference. This is a source-main fix, not a new PyPI release. No audio or model download is needed for the parser/constructor retest.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs feedbackWaiting for reporter feedback or retest results

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions