The Technology Innovation Institute (TII) has released Falcon-Emirati-7B, a specialized large language model designed to understand and generate Emirati Arabic. Built upon the existing Falcon-H1-Arabic architecture, the new model targets the specific vocabulary, tone, and cultural nuances of the UAE’s primary spoken dialect, addressing a gap where general Modern Standard Arabic (MSA) models often fail to capture local context.

What Happened

Falcon-Emirati-7B is a 7-billion parameter model derived from the Falcon-H1-Arabic family, which utilizes a hybrid architecture combining State Space Models (Mamba) and Transformer attention. TII selected the 7B variant as a balance between quality and practical training and inference costs, noting that larger models like the 34B variant were too expensive for a dialect-specialized chat model, while the 3B variant lacked sufficient capacity for deep cultural understanding.

The development process involved creating a dedicated data pipeline that combined authentic Emirati web data, MSA texts about Emirati culture and identity, and synthetic data constrained by strict glossaries to ensure authenticity. Because there is no established playbook for MSA-to-dialect adaptation, the team relied on experimental ablations to determine optimal data mixes and training stages, evaluating progress through both automatic scoring and native-speaker review.

Why It Matters

The release highlights the limitations of scaling general-purpose models for specific linguistic tasks. According to TII, Falcon-Emirati-7B achieved an 84.83% score on Alyah, a new benchmark comprising 1,173 multiple-choice samples collected from native Emirati speakers. This score surpassed several larger multilingual models, suggesting that dialect competence requires targeted training rather than just increased parameter count.

A critical differentiator identified in TII’s evaluation is dialect fidelity. In open-ended generation tests judged by an LLM, Falcon-Emirati-7B scored 0.52 on dialect fidelity, compared to 0.05 for ALLaM-7B and effectively 0.00 for Fanar-2-27B. While competing models often provided correct answers in Modern Standard Arabic, Falcon-Emirati-7B was the only model in the comparison that reliably responded in Emirati dialect when prompted in that language. This distinction is vital for applications requiring natural, culturally embedded interactions, such as customer service or local government services in the UAE.

The Bottom Line

Falcon-Emirati-7B is now available on TII’s chat platform. The model demonstrates that specialized data curation and dialect-aware training can outperform larger, general-purpose models in specific linguistic and cultural contexts, particularly in preserving the register and nuance of spoken Emirati Arabic.