Skip to content

fix(fun-asr-nano): restore repetitive-garbage trimming in streaming vLLM - #3766

Open
MincongZhou wants to merge 2 commits into
modelscope:mainfrom
MincongZhou:fix/fun-asr-nano-streaming-repetition-trim
Open

MincongZhou wants to merge 2 commits into
modelscope:mainfrom
MincongZhou:fix/fun-asr-nano-streaming-repetition-trim

Conversation

@MincongZhou

Copy link
Copy Markdown

Summary

funasr/models/fun_asr_nano/inference_vllm_streaming.py carried a raw 0x01 control byte where its sibling inference_vllm_pipeline.py has the two characters \1:

# streaming  (line 37)   -- the 0x01 is a literal U+0001 inside a raw string
text = re.sub(r'(>.{2,8}?).{3,}', '', text)

# pipeline   (line 43)
text = re.sub(r"(>.{2,8}?)\1{3,}", "", text)

xxd on the two source lines:

streaming: 62 28 72 27 28 3e 2e 7b 32 2c 38 7d 3f 29 01 7b   -> b"r'(>.{2,8}?).{"
pipeline : 62 28 72 22 28 3e 2e 7b 32 2c 38 7d 3f 29 5c 31 7b -> b"r\"(>.{2,8}?)\1{"

So the streaming pattern is (>.{2,8}?)\x01{3,} — it requires three or more literal U+0001 characters, which no model output contains, and therefore never fires. This is the only control character anywhere under funasr/ (grep -rlP '\x01' funasr --include=*.py → this one file).

The intent is unambiguous from the history:

  • 3817ca31b3 (feature introduction) had r"(>.{2,8}?)\1{3,}".
  • 4bb5b543fe ("fix: improve streaming vLLM with two-stage approach for long audio") rewrote this function, and its own diff shows the corruption — git renders 0x01 as ^A:
    +        text = re.sub(r'(>.{2,8}?)^A{3,}', '', text)
    
  • The pipeline copy, added separately, kept \1.

streaming_generate calls _clean_text for the returned text, for the fixed_text = text[:-rollback_chars] fixed/unfixed window and for prev_text of later chunks. The streaming engine therefore keeps repetition runs of exactly the shape this rule exists to remove — the same model on the same audio gets cleaned by the pipeline path and not by the streaming path. This rule is the "后处理截断" remedy already relied on by the pipeline path (cf. #2930, #2950).

Type of change

  • Bug fix
  • Documentation
  • Example or demo
  • Runtime or deployment
  • Benchmark or evaluation
  • Model/training change

Validation

  • python -m compileall funasr examples tests → exit 0 on the patched tree.
  • Docs or links checked — not applicable, no docs touched.
  • Runtime/deployment command tested — see the limitation below.

The two _clean_text implementations were loaded verbatim from the real files (the function AST extracted and executed with only re in its namespace; the modules' torch/vLLM imports are unused by this function and were not needed) and compared on strings of exactly the shape the rule targets:

strings the pipeline modifies: 22
streaming BEFORE cleans :  3/22
streaming AFTER  cleans : 22/22
streaming AFTER == pipeline on every case: True

Concretely, unchanged at main:

in  '你好。>这个表面>这个表面>这个表面>这个表面'
  before='你好。>这个表面>这个表面>这个表面>这个表面'
  after ='你好。'
  pipe  ='你好。'

in  '<|zh|>你好<|nospeech|>>啊呀呀>啊呀呀>啊呀呀>啊呀呀[breath]'
  before='你好>啊呀呀>啊呀呀>啊呀呀>啊呀呀'
  after ='你好'
  pipe  ='你好'

I could not run the real streaming engine end to end (no GPU, no vLLM, no checkpoint here), so I cannot show a real transcript that the pipeline cleans and streaming does not. The claim is limited to the function's behaviour, which is where the divergence is, and both copies are now identical on every case tried.

black --line-length=100 --check output is byte-identical before and after — this file is already not black-formatted at HEAD, so it is deliberately not reformatted here (a one-byte fix, 1 insertion / 1 deletion).

User impact

Anyone serving Fun-ASR-nano through the streaming vLLM engine gets cleaner transcripts, matching what the VAD+vLLM pipeline engine already produces for the same audio — repetition runs like 你好。>这个表面>这个表面>… no longer leak into the returned text, the fixed/unfixed window or the next chunk's prev_text.

Notes for reviewers

  • One byte of source; no logic, signature or config change, and no other file touched.
  • The offline vLLM engine (inference_vllm.py:665-670) has no repetition clause at all. That is a separate question and deliberately out of scope for this minimal fix.
  • Not verified locally: a real end-to-end streaming run (no GPU/checkpoint in this environment).

Competition

No competing item. gh api search/issues for _clean_text → 0, for inference_vllm_streaming → 0, and for a control character in this context → 1 unrelated issue. The repo currently has only two open PRs (#3726, #3686), neither of which touches this file. The neighbouring issues that discuss this remedy (#2930, #2950, #3595, #3597) are all closed and address it by other means.

inference_vllm_streaming.py carried a raw 0x01 control byte where its sibling
inference_vllm_pipeline.py has the two characters \1:

  streaming: re.sub(r'(>.{2,8}?).{3,}', '', text)   <- literal U+0001 in the raw string
  pipeline : re.sub(r"(>.{2,8}?)\1{3,}", "", text)  <- backreference

The streaming pattern therefore requires three or more U+0001 characters, which
no model output contains, so the rule never fires. It is the only control
character in funasr/. The intent is not in doubt: the commit that introduced
the feature (3817ca3) had \1, and the streaming rewrite (4bb5b54) shows
the 0x01 in its own diff, where git renders it as ^A:

  +        text = re.sub(r'(>.{2,8}?)^A{3,}', '', text)

Effect: streaming_generate calls _clean_text for the returned text, for the
fixed_text = text[:-rollback_chars] window and for prev_text, so the streaming
engine kept repetition runs ("你好。>这个表面>这个表面>...") that the pipeline
engine strips from the same audio -- the post-processing remedy the pipeline
path already relies on.

Restore the backreference so both runtimes trim the same garbage.
Import both production _clean_text implementations and require them to produce
the same output: on repetitive garbage (which the pipeline engine trims, the
streaming engine did not) and on text without repetition (which neither should
touch).

The cases enumerate the rule's real boundaries rather than just the happy path:
four repeats trim, three do not; the repeated group needs at least two
characters; a genuine 0x01 byte in the text must survive.

No weights, GPU or model artifact are needed - the modules only import torch,
numpy and re at import time, which is what the neighbouring fun_asr_nano tests
rely on too.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant