- The Falcon team reports that its Falcon-Emirati-7B model scores 84.83% on Alyah, a 1,173-question Emirati-dialect benchmark, and answers in dialect far more often than four rivals.
- The results are the developer's own, judged partly by another AI model. They still show that a correct answer in the wrong register can fail a customer.
- Before buying a language model for a regional market, test it in the dialect your customers speak, with native speakers reviewing the output.
A correct answer in the wrong voice
An AI model can know the right answer and still fail the person who asked. The team behind Falcon found this when it tested models in Emirati Arabic. That is the Gulf dialect people use for daily talk, humor and negotiation.
The team's report says the rival models often knew the answer. But they replied in Modern Standard Arabic (MSA), the formal style of news and textbooks. The lesson here is that accuracy and fluency are different things. A customer-facing system is judged on how it sounds, not only on what it says.
How the model was built
Falcon-Emirati-7B builds on Falcon-H1-Arabic, the team's existing Arabic model family. Every block of that design runs two methods side by side. One is a State Space Model (Mamba), which handles long text efficiently. The other is Transformer attention, which keeps precision across long passages.
The team chose the 7-billion-parameter size. Its stated reason was a good balance of quality against the cost of training and running the model.
Next came a dedicated data pipeline with three sources. The first was material gathered from Emirati sites and forums, where people write the dialect as they speak it. The second was formal Arabic writing about Emirati culture, heritage and social norms. The third was synthetic dialect data from a generator model.
The team limited that generator with rules, glossaries and dictionaries for Emirati words and grammar. Without those limits, it said, the text can be correct yet still sound off to a native speaker.
The team also says no standard recipe exists for this kind of adaptation. So it tested how much dialect data to add and when in training. It also tested how to mix real and synthetic text. Native speakers reviewed outputs alongside the automated scores.
What the tests show
The team used Alyah, a benchmark of 1,173 multiple-choice questions collected by hand from native Emirati speakers. Topics include greetings, figurative language, heritage and poetry. The team's own model came out on top at 84.83%. That included models many times its size.
Multiple choice only tests whether a model can spot the right answer. So the team ran a second test. Five models, Falcon-Emirati-7B and four rivals, answered the same questions in free text.
An AI judge, Gemini 3.7 Flash, scored two things separately. One was correctness. The other was whether the reply came back in dialect rather than MSA.
On dialect fidelity (partial credit), Falcon-Emirati-7B scored 0.52. ALLaM-7B-Instruct-preview scored 0.05, gemma-3-27b-it 0.03 and Jais-2-8B-Chat 0.02. Fanar-2-27B-Instruct scored effectively zero.
Fanar-2-27B-Instruct also declined to answer 26.2% of the time. Every other model in the comparison declined under 5% of the time.
The team also says size did not decide the result. Some of the largest multilingual models scored well below smaller, dialect-aware ones.
The gap was widest where the dialect is most distinctive, such as poetry. It was narrowest in greetings, where Emirati and MSA overlap most. There, rivals held their own.
Read the evidence with care
This is a developer reporting on its own model. The Falcon team says it and the community released Alyah. An AI model did much of the judging. The report does not say anyone else reproduced the results.
The result is also narrow. It covers Emirati Arabic and four rivals that the team picked from the Alyah leaderboard. It says nothing about other dialects or languages.
Even so, the pattern helps buyers. The report suggests regional fluency has to be trained and tested on purpose. Size alone does not provide it.
Questions to put to your team
Planning to put AI in front of customers in a regional market? Ask four things.
Which dialect do our customers actually write in? Was the model tested in that dialect, or only in the formal language? Do native speakers review the outputs, and how often? What share of answers come back in the register the customer used?
Ask vendors for dialect results separately from accuracy results. A model can give the right answer in the wrong voice. It may pass a benchmark and still sound foreign to the customer.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Hugging Face.





