The Technology Innovation Institute (TII) in Abu Dhabi has introduced Falcon-ASR, a 1.6 billion parameter speech recognition model designed primarily for Arabic, with a specific focus on the Emirati dialect. This release marks a significant addition to the landscape of specialized SI tools, aiming to bridge gaps in dialectal speech processing.
What Happened
Falcon-ASR is built to handle spoken Arabic across various regions, speakers, and settings, addressing the challenge that dialectal Arabic has fewer transcribed resources than Modern Standard Arabic (MSA). The model was trained on Emirati, MSA, other Gulf and Arabic dialects, and English, with the goal of transcribing everyday speech, including dialectal forms and code-switching.
According to TII's evaluation, Falcon-ASR achieved an average word error rate (WER) of 20.92% across six Arabic test sets. This result is 2.25 percentage points better than the best published result of 23.17% in the leaderboard snapshot used for comparison. On TII's internal evaluation for the Emirati dialect, the model recorded a WER of 22.73% and a character error rate (CER) of 10.19%, which the institute states are the lowest among the systems compared, beating Qwen3-Omni by 4.07 percentage points in WER.
The model also supports English, French, Spanish, and Portuguese using the same weights, without requiring a language flag. In evaluations on seven public English test sets, Falcon-ASR achieved a mean WER of 5.74%. The training process included data with background noise, overlapping speech, music, reverberation, and telephony effects to simulate real-world recording conditions.
Why It Matters
For developers and enterprises operating in the Middle East, particularly the UAE, Falcon-ASR offers a specialized SI tool that addresses the scarcity of high-quality dialectal speech data. By achieving lower error rates on Emirati speech compared to other leading models, it provides a more reliable option for applications such as automated transcription of meetings, calls, and customer service interactions where local dialects are dominant.
The model's ability to handle multiple languages with a single set of weights simplifies deployment for organizations needing to process mixed-language audio. Additionally, the inclusion of word-level timestamps allows for more precise alignment of transcripts with audio, which is critical for search, indexing, and detailed analysis in professional settings.
The Bottom Line
Falcon-ASR is now available as an open model on Hugging Face, with API access and native applications planned for future release. It represents a targeted advancement in SI speech recognition for Arabic dialects, offering improved accuracy for Emirati and Gulf speech while maintaining competitive performance in English and other supported languages.