Skip to content

Proposal: Indic code-switched data (Hindi-English / Tamil-English) + code-switch metric #177

Description

@balaguhanesh

#175 adds Hindi data, but monolingual. In production Indian-market voice agents, callers code-switch intra-sentence ("mera appointment reschedule karna hai next Tuesday ko") — exactly where ASR degrades hardest. The ServiceNow code-switching benchmark covers ES/FR/DE-English; no Indic pairs exist anywhere.

I'd like to contribute, in stages:

  1. Hindi-English + Tamil-English code-switched utterance sets across existing domains, following Adding Korean and Hindi data #175's conventions (I'm a native speaker of both; I build STT/TTS eval harnesses for code-switching at 2care.ai)
  2. Code-Switch F1 scoring alongside WER — accuracy at switch-point words, so boundary degradation isn't averaged away
  3. (Follow-up PR) a code-switch perturbation for the perturbation pipeline

Before building: is this something you'd take, and is anyone working on it internally? If green-lit, first set + scoring ready for review in ~a week.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions