The UAE-based AI research team at TIIUAE, in partnership with Hugging Face, has released Falcon-Emirati-7B, a specialized large language model designed to understand and generate Emirati Arabic with native-speaker proficiency. The 7-billion-parameter model addresses a critical gap in AI language capabilities: most Arabic language models are trained primarily on Modern Standard Arabic (MSA), the formal written dialect, and fail to capture the vocabulary, grammar, and cultural nuance of Emirati Arabic as spoken in everyday conversation.
Built on the foundation of Falcon-H1-Arabic, which already demonstrated strong multilingual Arabic capabilities, Falcon-Emirati-7B represents a focused effort to master a single dialect at depth. The model uses a hybrid Mamba-Transformer architecture that combines the computational efficiency of State Space Models with the precision of attention mechanisms—particularly important for Arabic's morphologically rich structure. The team selected the 7B scale specifically as a practical middle ground: large enough to capture dialect nuance, but small enough to keep training and inference costs manageable.
The challenge of dialect adaptation proved far more complex than simply fine-tuning a general Arabic model. Emirati Arabic exists primarily in spoken form, with limited written online representation compared to MSA or other regional dialects. Beyond vocabulary and grammar, the dialect relies heavily on cultural references, idiomatic expressions, and poetic allusions that require deeper contextual understanding. To overcome these obstacles, TIIUAE built a three-part data pipeline: authentic Emirati web content scraped from native sources, MSA-language materials about Emirati culture and heritage, and synthetically generated dialectal data constrained by strict glossaries and style rules to ensure quality.
Key Points
Falcon-Emirati-7B specializes in Emirati Arabic dialect, addressing limitations of general models trained only on Modern Standard Arabic
The hybrid Mamba-Transformer architecture balances computational efficiency with long-range contextual understanding for morphologically rich language
Dialect adaptation proved harder than expected due to minimal written Emirati Arabic data and reliance on cultural context
Three-part data strategy combined authentic web content, culturally-focused MSA materials, and synthetically-generated data with strict linguistic constraints
7B parameters selected as optimal balance between capturing dialect nuance and maintaining practical training and inference costs