Viyan

Viyan AI

Multimodal ICL beats fine-tuning in MLC-SLM challenge

In the 2nd MLC-SLM challenge, a multimodal in-context learning strategy on a frozen 24B model achieved higher accuracy than fine-tuned smaller models.

The Eloquence team reached 0.81 macro-accuracy in Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026 by using multimodal in-context learning on a frozen Voxtral-24B model. The task requires models to map acoustic inputs to semantic labels across 21 languages. By applying in-context learning, the team outperformed a LoRA-tuned baseline on the same task, which achieved 0.72 macro-accuracy. The shift to a larger, frozen model was motivated by the need to handle complex cross-lingual classification without the label biases that often emerge during fine-tuning on limited datasets.

Comparison of System Approaches

Method Model Base Strategy Macro-Accuracy
Fine-Tuning Voxtral-Mini-3B LoRA + Augmentation 0.72
Retrieval Custom Training-free memory 0.68
Multimodal ICL Voxtral-24B Frozen weights 0.81

Fine-tuning the Voxtral-Mini-3B model involved cross-lingual data augmentation, ASR transcript augmentation, and timestamp-aware audio cropping. While these steps improve performance on specific language pairs, the model still struggled with consistent classification across the full set of 21 languages. The 24B model avoids these errors because its larger parameter space creates more robust internal representations of multilingual phonemes, allowing it to correctly map diverse acoustic inputs to the same semantic label where the 3B model fails due to feature overlap or insufficient linguistic coverage.

The retrieval system, which scored 0.68, utilizes a three-layer architecture. First, it extracts acoustic identity, then maps that identity to semantic content, and finally verifies the result against an external knowledge graph. The graph acts as a secondary gate, checking the proposed label against known multilingual constraints. The system failed to reach the accuracy of the 24B model because the retrieval process introduces a step of abstraction; the model has to resolve the retrieved information before it can classify the input. This extra step appears to introduce noise into the classification boundary, whereas the large frozen model binds features to labels directly through its internal weights.

It remains uncertain whether this performance gap would persist if the 3B model were trained on a dataset more representative of the 21-language distribution rather than using augmentation. Similarly, the exact mechanism that causes the retrieval system to lose accuracy during the binding process is still an area of active investigation. Future work in this challenge will likely need to focus on how to bridge this gap without needing to scale to 24 billion parameters for real-time multilingual classification.

Sources