Viyan

Viyan AI

Nuha-Speech Adds Arabic Audio Tokens to Qwen-Omni

The Nuha-Speech project fine-tunes Qwen-Omni on 1.5 million Arabic samples to enable native speech-to-speech interaction without intermediate transcription.

The Nuha-Speech initiative leverages 1.5 million audio-text samples to fine-tune the Qwen-Omni architecture, specifically creating an Arabic-native Speech Question-Answering (SQA) model. By bypassing the traditional two-stage transcription pipeline—where audio must first be converted to text by an ASR engine before reaching the language model—this approach allows the model to process Arabic audio for instructional tasks directly. It transitions the processing architecture from a cascaded pipeline to a unified instruction-tuning framework for SQA tasks.

Translating Audio into Model Inputs

To bridge the gap between raw audio waveforms and the Transformer’s architecture, the model must first convert audio into a representation it can read. Nuha-Speech processes audio by converting waveforms into a sequence of continuous feature representations, such as mel-spectrograms, which are then passed through a projection layer. This layer acts as a bridge, mapping these features into the same latent embedding space that the model uses for text tokens. Once the audio is projected into this shared space, the model perceives it as a sequence of input tokens. Because the input is continuous, the model can preserve information that is typically stripped away during the transcription process, such as prosody and emotional inflection. Unlike text, where each character is a discrete token, these continuous latent tokens carry the spectral signature of the original speech, which allows the model to respond to the tone and phrasing of the input directly.

Feature Traditional Pipeline Nuha-Speech SQA Approach
Data Flow Audio -> ASR -> Text -> LLM Audio -> Embedding -> LLM
Pipeline Stages Two-stage, decoupled Single, unified pass
Information Text-only (lossy) Audio-aware (prosody preserved)
Primary Use General ASR/Transcription Arabic Speech Question-Answering

Operational Considerations

In a standard ASR-based pipeline, the language model is only as accurate as the text provided by the transcription engine. When ASR fails to transcribe a word correctly, the downstream reasoning model receives distorted input, which often leads to errors in logic or understanding. By inputting latent audio tokens directly, you bypass the transcription bottleneck, effectively ignoring the concept of word-error rates at the input stage. If your system requires nuanced responses to spoken Arabic questions, this model is built to retain the original acoustic properties of the speaker's intent.

However, this model's performance on non-standard hardware or in noisy environments is unverified, as current documentation only covers standard speech patterns. The model is trained on specific spectral characteristics; when background noise or a different microphone changes the spectral profile of the input, the distribution of the latent tokens shifts. Because the model relies on a specific mapping learned during training, these environmental shifts can push the input into a space the model does not recognize, resulting in unpredictable output. If you are building with this, you must test how your deployment environment's audio capture affects the latent mapping. The main constraint remains the lack of clear data on how well the model handles dialectal diversity, as the current training set emphasizes standard Arabic patterns over localized variants.

Sources