Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Falcon-Emirati-7B: A Model Built for Emirati Arabic Dialect - OpenSmartRoute
Falcon-Emirati-7B is a new model designed for Emirati Arabic dialect. Tiiuae released it to handle local language nuances better than general models. It scores 84.83% on the Alyah benchmark, beating many larger competitors. This release marks a shift toward specialized regional language support in large language models.
Falcon-Emirati-7B: A Dialect-Specialized Model for Emirati Arabic
Arabic is a family of languages that often share one name but differ in daily use. Modern Standard Arabic appears in news and textbooks, yet people rarely speak it this way. In the UAE, conversation, humor, and storytelling happen in Emirati Arabic. This Gulf dialect has its own vocabulary, rhythm, and deep cultural roots.
Poetry, proverbs, and short anecdotes carry meaning that disappears with literal translation. A model knowing only Modern Standard Arabic can translate every word correctly. It still misses the actual intent behind the sentence. The gap between surface words and cultural meaning is where Falcon-Emirati-7B aims to help.
This model focuses on understanding and generating Emirati Arabic like a native speaker would. It captures the specific vocabulary, tone, and context of the region. General models struggle with this depth because they lack local exposure. Falcon-Emirati-7B fills that gap by training specifically on dialectal data.
The Falcon-H1-Arabic Family and Its Hybrid Architecture
Falcon-Emirati-7B is not a standalone creation but part of a larger family. It builds directly on top of Falcon-H1-Arabic, which set new benchmarks earlier this year. This base model uses a unique hybrid architecture combining two different technologies. State Space Models, or Mamba, run in parallel with Transformer attention blocks.
The outputs from both systems fuse together before each block's projection layer. This design gives linear-time efficiency for long sequences while keeping precision for long-range dependencies. Arabic is morphologically rich, so handling long-range context is crucial for understanding grammar and meaning.
The Falcon-H1 family spans three scales: 3 billion, 7 billion, and 34 billion parameters. Context windows reach up to 128K or 256K tokens in some configurations. The base model trained on a broad mix of Modern Standard Arabic and various dialects. It included Gulf, Levantine, Egyptian, and Maghrebi dialects alongside English and multilingual data.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Mistral released a preview of Mistral Large 4, a one-trillion-parameter model trained in Europe.
This foundation gave the team a strong starting point for specialization. The base model already understood Arabic broadly and handled long context well. It had some initial exposure to dialectal patterns baked into its training. Falcon-Emirati-7B takes this general capability and pushes it specifically toward the Emirati dialect.
Building a Specialist: Data Sources and Training Challenges
The team chose the 7 billion parameter variant for Falcon-Emirati-7B specifically. This size offers the best balance of quality against training and inference cost. The 34 billion model would likely push quality further but is too expensive for a chat model. The 3 billion model lacks the headroom for deep cultural understanding.
Turning a general Arabic model into a dialect specialist sounds easy, yet it is genuinely difficult. Emirati Arabic is mostly spoken, so there is far less written text online than Modern Standard Arabic. There is simply not enough raw data to learn from without synthetic help.
Meaning in the dialect is often non-literal. Idioms and proverbs rely on shared cultural context rather than surface vocabulary. There is no established playbook for how much dialectal data is enough. Experts do not know which training stage matters most for picking up a dialect.
The team relied heavily on trial and error to build Falcon-Emirati-7B. They tested different data mixes, training stages, and supervision strategies. Human judgment and benchmark scores helped figure out what actually moved the needle. No single recipe worked perfectly across all scenarios.
Evaluation on Alyah: Benchmark Scores and Competitor Comparison
To measure progress, the team created a dedicated Emirati data pipeline. They drew on three complementary sources to build their training dataset. First, they crawled content from Emirati websites and forums written natively in the dialect. This provided ground truth on how people actually write and speak online.
Second, they included Modern Standard Arabic material about Emirati culture and identity. Articles covered local customs, values, history, and social norms. This taught the model what it is talking about when Emirati topics come up. It understood heritage, etiquette, and the context a native speaker knows instinctively.
Third, they generated synthetic data guided by glossaries and style rules. Authentic text alone did not cover the range of topics a chat model needs. They constrained generators with strict rules built for Emirati vocabulary and grammar. These guardrails ensured synthetic output sounded authentically Emirati rather than just grammatically fine.
They tracked progress using native speaker reviews and quantitative benchmarks. Native speakers judged outputs on naturalness, tone, and cultural appropriateness. Automatic metrics alone cannot capture these nuances well enough to trust them fully. For quantitative tracking, they used Alyah, a benchmark released by the community.
Alyah is a fully native multiple-choice benchmark of 1,173 samples. It spans categories from greetings to poetry where dialect matters most. Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other model compared. This includes several models many times its size in terms of parameters.
Size alone does not buy dialect competence in this comparison. Some large multilingual models score well below smaller, more dialect-aware ones. Emirati proficiency must be trained for on purpose, not picked up as a side effect of scale. The best performing models are Arabic-native or Arabic-focused to begin with.
Dialect Fidelity vs. Modern Standard Arabic in Open-Ended Tasks
Multiple-choice accuracy tells you if a model can recognize the right answer among four options. It does not tell you if the model will produce Emirati Arabic on its own during open conversation. So the team ran a second evaluation: open-ended generation on the same 1,173 Alyah questions.
They scored these answers using an LLM judge named Gemini 3.7 Flash. The judge evaluated five models including Falcon-Emirati-7B and others like ALLaM-7B-Instruct-preview. Scoring covered two dimensions: content correctness and dialect fidelity. Dialect fidelity checks if the answer comes back in Emirati rather than Modern Standard Arabic.
Falcon-Emirati-7B leads on correctness scores across the board. However, the real gap appears in the second chart regarding dialect fidelity. Falcon-Emirati-7B scores 0.52 partial credit while ALLaM scores only 0.05. The difference is close to two orders of magnitude at the low end.
This means other models often know the right answer but say it in Modern Standard Arabic by default. Falcon-Emirati-7B is the only one that reliably answers back in the dialect it was asked in. Fanar-2-27B-Instruct stands out for another reason: it abstains far more than any other model. It declines to answer 26.2% of the time versus under 5% for others.
Dialect fidelity holds across every single category in Alyah, from greetings to poetry. This suggests a general shift in register rather than a narrow trick learned for specific question types. Competing models stay in Modern Standard Arabic across nearly every category. The one place they do relatively better is Greetings & Daily Expressions. Here, Emirati and Modern Standard Arabic overlap the most.
Why Dialect Adaptation Requires Targeted Data and Evaluation
Beyond multiple choice, LLM-as-Judge evaluation revealed deeper issues with competing models. Pairwise comparison showed Falcon-Emirati-7B consistently outperforming others in head-to-head judging. The judge picked Falcon-Emirati-7B's answers as better across the board.
Breaking dialect fidelity down by category makes the pattern even clearer. It holds true for everyday greetings, heritage knowledge, and figurative language. This indicates a fundamental change in how the model defaults to register when Emirati is expected. Generic Arabic models struggle to adapt their output style dynamically.
The team learned that data and evaluation must be built specifically for the dialect. General Arabic coverage is a necessary starting point but not sufficient. Targeted, dialect-specific work closes the rest of the gap on hard parts like poetry. Cultural context requires more than just linguistic patterns.
Native speaker review remains essential alongside automatic scoring. Human judgment catches naturalness and tone that metrics miss. Automatic scores alone cannot capture cultural fit well enough to trust them fully. The combination of both approaches provided the most reliable feedback loop.
What to Do: Choosing Models for Regional Language Support
Engineers and managers need clear steps when choosing models for regional language support. First, check benchmark results on dialect-specific evaluations like Alyah. Look for scores that reflect actual performance rather than just parameter count. Size alone does not guarantee competence in specialized domains.
Second, consider the trade-off between quality and cost for your specific use case. The 7 billion variant offers a practical balance for many chat applications. Larger models may be needed for extreme depth but come with higher training and serving costs. Smaller models lack the headroom for complex cultural understanding.
Third, verify dialect fidelity through open-ended generation tests. Multiple-choice accuracy is not enough to ensure the model speaks your target language. Open-ended generation reveals if the model defaults to Modern Standard Arabic when asked in a dialect. This is critical for user experience and perceived authenticity.
Fourth, plan for evaluation strategies that include native speaker review. Automatic metrics cannot fully capture cultural appropriateness or tone. Human judgment complements quantitative scores to ensure the output feels right to the target audience.
Finally, consider data availability and synthetic generation needs. Spoken dialects often lack sufficient written training data. Teams may need to generate synthetic data guided by glossaries and style rules. This ensures coverage across diverse topics without overfitting to specific patterns.