Viyan

Viyan AI

Why Global WER Misleads on Code-Switched Speech

Aggregate Word Error Rate hides how speech models fail at language transitions, necessitating switch-specific metrics for reliable evaluation.

Aggregate Word Error Rate (WER) is an insufficient metric for evaluating modern speech models on code-switched data. Research testing eleven ASR and audio language models on English-Yoruba speech shows that the best-performing ASR model and a leading audio language model exhibit statistically indistinguishable aggregate WER scores. Yet, they show vastly different performance when transitioning between languages. WER provides a single global error percentage, but it obscures that errors cluster at switch points—the moments where a speaker shifts from English to Yoruba.

The Mechanism of Switch Point Failure

The failure occurs because the autoregressive decoder architecture maintains a strong bias toward the previous language. In a standard Transformer decoder, causal self-attention allows each predicted token to attend to all preceding tokens. When the acoustic encoder processes a switch from English to Yoruba, the decoder’s hidden state is dominated by the linguistic patterns of the English tokens just generated. The model’s probability distribution for the next token is conditioned on these preceding English embeddings, causing the attention heads to place high probability mass on English vocabulary tokens. This suppresses the activation of Yoruba-specific tokens in the model’s output layer, creating a persistent inertia. Even when the acoustic input signals a shift, the weight assigned to the recent English context overwhelms the incoming acoustic signal.

Evaluation on English-Yoruba speech shows that this systemic failure in diacritic-rich, low-resource language processing occurs across nearly all tested systems. Yoruba token recognition consistently collapses, with error rates reaching 0.97 across almost all evaluated models. Aggregate WER effectively flattens this disparity. The leading audio language model outperformed the leading ASR model on switch-localized metrics. However, audio LMs often fail as transcribers due to their training objectives. Unlike standard ASR models trained on a strictly supervised sequence-to-sequence loss, audio LMs are typically trained as generative language models. This leads them to prioritize fluent output over verbatim accuracy. Consequently, when the model faces ambiguous acoustic input, it defaults to its language modeling prior—resulting in translation of the content, unintended verbosity, or the generation of text matching a likely prompt rather than the actual audio.

Metric Category Standard ASR Audio Language Model Insight
Aggregate WER Low Low Hides local failures
Switch Point Performance Poor Better Audio LMs adapt to transitions
Yoruba Token Accuracy Collapsed Collapsed Universal weakness
Output Fidelity Verbatim Erratic/Verbose LMs struggle with constraints

Measuring the Transition

To identify where these systems break, you must isolate the transition moment in the audio signal. Standard WER averages performance across an entire file, washing out the high error rates occurring in the tiny temporal window of the language shift. Metrics like Switch Entry Token Error Rate (SETER) operate differently by anchoring evaluation to the token where the language change occurs. By calculating error rates within a fixed window surrounding this switch, SETER captures whether the model recognized the shift or continued hallucinating in the previous language. Windowed error rates allow you to see if the model's accuracy recovers ten or twenty tokens after the switch. This helps you distinguish between a temporary transition failure and a complete collapse of language identification.

For example, if a user says "I am going to the market, mo n lo si oja," a model might transcribe the English correctly but produce a string of gibberish or English words for the Yoruba clause. If you rely only on a 15% aggregate WER, you might assume the system is functioning well. The system has failed to transcribe the second language segment. The difficulty in resolving this lies in the technical tradeoff: reducing this inertia requires lowering the context window's influence on the current token prediction. This is problematic because global sentence structure relies on long-distance dependencies where a subject at the start of a sentence determines a verb conjugation much later. If you weaken the influence of past tokens to allow a rapid language shift, you decrease the model's ability to maintain coherent grammar and syntactic structure across that same long-range span.

What remains unknown is the degree to which multi-task training or specific fine-tuning on diverse code-switched datasets can decouple language identification from sequence generation. We can see that audio LMs manage the transition better, but we do not yet know if this is due to superior acoustic encoder representations or simply the model's inherent ability to translate between languages when it recognizes the shift. Moving forward, developers should look for models that treat language identification as a distinct gated task within the decoder, rather than relying on the general attention mechanism to infer language shifts from acoustic cues alone.

Sources