TII Launches 1.6B Falcon-ASR for Arabic and Emirati Speech

Abu Dhabi’s Technology Innovation Institute (TII) introduced Falcon-ASR on 7 October 2026, a 1.6-billion-parameter speech recognition model for Arabic with a particular focus on the Emirati dialect. It also transcribes English, French, Spanish and Portuguese.

In TII’s own evaluation, the model averaged a 20.92% word error rate across six Arabic test sets, against 23.17% for the best published result in the leaderboard snapshot TII used. On an internal Emirati test, TII says it had the lowest word and character error rates of the systems it compared.

One Model in a Three-Model Launch

Falcon-ASR arrived alongside two other TII models, according to the launch statement carried by Gulf Today on 7 October: Falcon-Emirati, a 7-billion-parameter language model built for Emirati Arabic, and Falcon-OCR-Arabic, a model that extracts Arabic text and structured content from images and documents.

TII is the applied research arm of Abu Dhabi’s Advanced Technology Research Council (ATRC), the statement says. It lists uses for Falcon-ASR including multilingual customer service, meeting and interview transcription, subtitling, accessibility tools and voice-enabled digital services.

TII’s chief executive, Najwa Aaraj, framed the launch as sovereign AI in the statement: “Sovereign AI capability must reflect the language used in daily life.”

TII’s post says all five languages run on the same weights without a language flag, and the output is a transcript in the language spoken. The model also gives word-level timestamps, linking each transcribed word to its position in the audio.

Average word error rate across six Arabic test sets, lower is better. Falcon-ASR’s 20.92% is TII’s
Average word error rate across six Arabic test sets, lower is better. Falcon-ASR’s 20.92% is TII’s own run on the leaderboard’s pinned manifests; the other figures are the leaderboard’s published averages, checked by TII on 30 September 2026. Figures transcribed from TII’s post.

Business Pill · SOVEREIGN AI

A one-minute explainer of sovereign AI: who controls the model and the data. It teaches the general idea only and says nothing about any company in this story.

The key insight: As we read it, the part of this launch that is hardest to copy is the Emirati test, not the model size: TII evaluated on its own held-out recordings with human-validated transcripts, in a dialect with fewer transcribed resources than Modern Standard Arabic, and that is where its widest reported margin sits.

What the Arabic Score Is Measured Against

TII scores the model on the Open Universal Arabic ASR Leaderboard, run by the ELM Research Center, which ranks systems by the equal-weight average word error rate across six test sets. Lower is better.

The comparison mixes two kinds of number, and the post says so. TII ran Falcon-ASR on the six benchmarks itself, using the leaderboard’s pinned manifests; the competitor figures are the leaderboard’s published averages, checked on 30 September 2026.

In TII’s chart, Falcon-ASR’s 20.92% compares with 23.17% for Audar-ASR-V1-Turbo, 25.87% for Cohere Transcribe Arabic (07-2026) and 28.32% for omniASR LLM 7B. Character error rates follow the same order: 8.79% for Falcon-ASR against 9.23% for Audar-ASR-V1-Turbo.

The Emirati Test Is TII’s Own

Public data already has some Emirati speech: the Casablanca dataset has a UAE subset, the post notes. TII added an internal evaluation of further Emirati and Gulf speech, using held-out recordings and human-validated transcripts.

On that internal test, Falcon-ASR scored 22.73% word error rate and 10.19% character error rate. The next best was Qwen3-Omni-30B-A3B-Instruct at 26.80%, which TII puts 4.07 percentage points behind; Audar-ASR-V1-Turbo followed at 27.89%.

The launch statement frames that result by size: “On Emirati speech, it surpassed the performance of a 30-billion-parameter multimodal model.” The test set is TII’s internal one, so the Emirati figures rest on TII’s own evaluation.

Word error rate on TII internal Emirati test: Falcon-ASR 22.73%, Qwen3-Omni-30B-A3B-Instruct 26.80%, Audar-ASR-V1-Turbo 27.89%, Cohere Transcribe Arabic 31.05%, Qwen3-ASR-1.7B-hf 31.52%, Audar-ASR-V1-Flash 32.87%
Word error rate on TII’s internal evaluation of held-out Emirati and Gulf recordings, lower is better. Figures transcribed from TII’s post; the test is TII’s own.

English From the Same Weights

On the seven public English test sets used by the Hugging Face Open ASR Leaderboard, TII reports a mean word error rate of 5.74%.

TII’s chart breaks that down by set. The lowest is LibriSpeech clean at 1.75% and the highest is Earnings-22 at 11.86%, with AMI at 8.33% and GigaSpeech at 8.15%.

Training for Dialect and Noise

TII trained the model on Emirati, Modern Standard Arabic, other Gulf and Arabic dialects, and English. The post gives the reason: dialectal Arabic has fewer transcribed resources than Modern Standard Arabic, which makes training and evaluation harder.

The training data included background noise, overlapping speech, music, room reverberation and telephony effects, plus changes in speed and pitch, the post says, with the same treatment applied to Emirati recordings. The model builds on TII’s Falcon3-Audio work, described in an arXiv paper.

The Structural Read

TII’s two Arabic results are different kinds of evidence. On the public sets, the 20.92% is TII’s own run set against published leaderboard averages; on Emirati speech, every system was scored on TII’s internal test, where the reported gap to the next model is 4.07 points.

The launch statement frames the result by size: a 1.6-billion-parameter model ahead of a 30-billion-parameter multimodal model on Emirati speech. As we read it, the claim is that specialised training data, not scale, did the work on this test.

The release came as one of three dialect- and document-focused models, alongside a 7-billion-parameter Emirati language model and an Arabic OCR model, according to the statement. As we read it, TII is building a set of Arabic tools layer by layer, from text to speech to documents.

Najwa Aaraj, Chief Executive Officer of TII, in the launch statement carried by Gulf Today (7 October 2026)

“Falcon-Emirati and Falcon-ASR will help ensure that the next generation of AI understands not only Arabic but how Emirati communities actually speak it.”

Three Implications

BUILDERS SERVING ARABIC SPEAKERS TII lists customer service, meeting transcription, subtitling and accessibility as uses, and the model returns word-level timestamps; for now access is a demo, with API access and applications described as planned.

SPEECH-MODEL VENDORS TII’s public Arabic comparison names Audar-ASR-V1-Turbo, Cohere Transcribe Arabic and omniASR LLM 7B; its internal Emirati test adds two Qwen models and Audar-ASR-V1-Flash.

READERS OF BENCHMARK CLAIMS The public Arabic comparison pairs TII’s own run with published averages from a 30 September snapshot, and the Emirati result comes from a test only TII holds.

The Business Engineer Lens

This story maps onto the Business Engineer framework The Four Intelligence Moats.

The framework puts it this way: “The transformer is a commodity now; what changes between paradigms is where intelligence accumulates and who can capture it.” It describes the first of the four as “The Corpus Moat — built by pretraining, eroding as the public web is exhausted and synthetic data closes gaps.”

As we read it, dialect speech is a place where that corpus is still thin: TII’s post says dialectal Arabic has fewer transcribed resources than Modern Standard Arabic, and its widest reported margin is on its internal test of held-out recordings with human-validated transcripts.

Access, and What Is Not Established

For now TII offers a demo on Hugging Face. The post says “API access and native applications are planned.” It gives no date for either, and neither the post nor the launch statement as carried by Gulf Today states a licence or a price.

Every score here is TII’s: its own run on the public Arabic sets, its own internal Emirati test and its own English evaluation. We have not run the model ourselves, and the leaderboard snapshot TII compared against was the one checked on 30 September.

Business Engineer Framework

The Four Intelligence Moats

A Business Engineer framework on where intelligence accumulates in the AI stack and who can capture it.

Read the Map of AI →

The Bottom Line

TII’s Falcon-ASR is a 1.6-billion-parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect, that also transcribes four other languages. By TII’s evaluation it averaged 20.92% word error rate on six Arabic test sets, below the best published leaderboard result of 23.17%, and 22.73% on an internal Emirati test. For now it is a demo, with API access and applications described as planned.

94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

A note on sourcing. We read TII’s post on Hugging Face in full on 10 October 2026, including its three result charts, whose figures we transcribed; every score here is TII’s. The three-model launch and the executives’ words come from TII’s launch statement as carried by Gulf Today on 7 October; We did not run the model. Nothing here is a forecast, and nothing here is financial or investment advice.

Sources: TII: Introducing Falcon ASR (Hugging Face, 7 Oct 2026) · Gulf Today: TII launches Falcon-Emirati alongside new Arabic AI models (launch statement, 7 Oct 2026) · ELM Research Center: Open Universal Arabic ASR Leaderboard

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA