Skip to content

Voice AI

How Voice AI Understands Arabic Dialects on the Phone

Why Arabic dialects challenge speech recognition on phone calls, how modern voice AI handles them, and how to test a vendor on your own call recordings.

Ranen teamUpdated 7 October 20267 min read

Saudi callers rarely speak textbook Arabic to a contact center. They speak Najdi, Hijazi or another Gulf variety, switch to English for a product name or an error message, read out an ID number, and often do it over a phone line from a car. A voice agent has to understand all of that before it can complete anything.

Public benchmarks show the size of the problem. In the Open Universal Arabic ASR Leaderboard, the best-ranked open-source model reached a word error rate (WER) of 19.23% on Modern Standard Arabic (MSA) in the Saudi SADA dataset, but 36.34% on Najdi and 48.23% on Khaliji speech from the same dataset. This article explains why, how modern systems narrow the gap, and how to test a vendor on your own calls.

Key takeaways

Dialect, code-switching, numbers and telephone audio each add errors, and they compound. Published scores on broadcast or read speech do not predict accuracy on your calls. The reliable test is a sample of your own recordings, scored by dialect, by critical details such as ID numbers and by task outcome, not by WER alone.

Why Arabic dialects are hard for speech recognition

Training data is skewed toward MSA

Most transcribed Arabic speech comes from news and broadcast in MSA. The leaderboard authors note that in the widely used MGB-2 corpus "over 70% of the samples represent the MSA dialect," and link the drop on dialects to "the imbalanced distribution of dialects in most public datasets." Djanibekov and colleagues reach a similar conclusion: existing systems "mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects."

Variety (SADA test data)WER, best open model
Modern Standard Arabic19.23%
Najdi36.34%
Hijazi36.96%
Egyptian40.97%
Khaliji48.23%

These figures come from television audio, not phone calls, and from open-source models, not commercial systems. They show the direction of the gap, not the accuracy of any specific vendor.

Sounds shift between dialects

The same letter is pronounced differently across regions. A study in the Journal of King Saud University describes how, in Najdi Arabic, the classical qaf is realized as [g], and in the Qassimi dialect it can be further fronted to [dz] in certain environments. Written Arabic also omits most short vowels, and the leaderboard paper lists the "lack of diacritics in written text" among the core difficulties. A model trained mainly on MSA pronunciation hears these regular patterns as different words.

Vocabulary and spelling vary

Everyday words for "want", "now" or "how much" differ by region, and dialects have no standard spelling. The Casablanca dataset, covering eight dialects, documents this lexical variation. Because the same dialect word can be written several ways, Hamed and colleagues note that word and character error rates alone are not adequate for unstandardized orthography.

Code-switching with English

Callers mix Arabic and English inside one sentence, especially for products, technical terms and job titles. A survey of code-switched Arabic NLP lists Saudi Arabia among the countries where Arabic-English switching is observed, and reports WER "in the range of 24.8-53.8%" across code-switched Arabic speech corpora. A monolingual model sees two sound systems and two scripts in one utterance.

Numbers, names and dates

The details a contact center needs most are the hardest: national ID and Iqama numbers, phone numbers, IBANs, order numbers, amounts, Hijri and Gregorian dates, and personal names. Callers group digits differently, mix Arabic and English numbers, and pronounce the same name in several ways. One wrong digit breaks a lookup even when the overall WER looks acceptable.

Telephone audio

Phone calls are narrowband, sampled at about 8 kHz, while many speech models are trained on 16 kHz audio. Researchers at IIT Madras note that "wide band (sampled at ∼16 kHz) ASR models do not perform well for narrow band speech data," and that labelled narrowband data is scarce. Speakerphones, road noise, compression and packet loss add further errors on top of dialect.

How modern voice AI handles dialects

  • More dialect data. SADA, built by SDAIA's National Center for AI with the Saudi Broadcasting Authority, holds about 667 hours from more than 57 TV shows. ADI17 covers about 3,000 hours from 17 countries, and Casablanca adds eight dialects with code-switching labels.
  • One model across dialects and languages. Chowdhury and colleagues showed that a single model trained on Arabic, English and French outperformed state-of-the-art monolingual dialectal and code-switching Arabic ASR. Djanibekov and colleagues released models covering 17 Arabic-speaking countries, including code-switched speech.
  • Dialect identification. Classifiers trained on data such as ADI17 detect the caller's variety, so the system can adapt recognition and reply in the same dialect.
  • Telephone adaptation. Training on narrowband audio, as in channel-aware pretraining, reduces the mismatch between 16 kHz models and 8 kHz calls.
  • Context and confirmation. Domain vocabulary (product, branch and medicine names) guides recognition, a language model interprets intent despite small transcription errors, and critical details are read back to the caller before any action.

What to measure beyond word error rate

MeasureWhat it showsWatch for
WER per dialect and channelTranscription accuracy for each caller groupSpelling variants inflate it; normalize first
Entity accuracyWhether IDs, numbers, dates and names are exactOne wrong digit is a failure
Intent accuracyWhether the agent understood the requestCan be high even when WER is moderate
Task completionWhether the request was completed end to endNeeds live or simulated calls
Reply dialectWhether answers match the caller's varietyRated by native listeners

How to test a vendor on your own recordings

  1. Build a stratified sample. Select real calls across your main dialects, age groups, mobile and landline, quiet and noisy conditions, and calls with English terms. A few hundred well-chosen calls tell you more than thousands of clean ones.
  2. Handle the data lawfully. Share only recordings you are permitted to share, under a data processing agreement. See our guide to PDPL and call recordings.
  3. Create reference transcripts. Native speakers transcribe with a written style guide covering dialect spelling, numbers and English words. Without it, scores measure spelling disagreements.
  4. Score by segment. Report WER per dialect and channel, entity accuracy and intent accuracy. Averages hide weak segments.
  5. Run live calls. Staff from different regions call a test line with scripted and free requests, including code-switching and asking for a person.
  6. Rate the replies. Native listeners judge whether answers are correct, natural and in the caller's dialect, and whether handovers to staff happen at the right moment.
  7. Agree thresholds in advance, then pilot on one line before expanding.

For how this differs from menu-based systems, see AI voice agent vs IVR, and for the wider picture, our guide to AI voice agents in Saudi contact centers.

How Ranen handles Arabic dialects

Ranen understands Modern Standard Arabic and Gulf, Egyptian, Levantine and Maghrebi dialects, plus English, including Arabic and English mixed in one sentence, and replies in the caller's dialect. Card and ID numbers are masked automatically in transcripts and recordings, which are hosted in data centers inside Saudi Arabia. Callers can ask for a person at any time. We run demos on your own call recordings, so you can apply the test above before committing. Request a demo.

Frequently asked questions

Can AI understand Saudi dialects on phone calls?

Yes, systems trained on dialect and telephone data can, but accuracy varies by dialect, audio quality and domain. Test on your own calls before deciding.

What is a good word error rate for Arabic speech recognition?

There is no universal threshold. For a contact center, accuracy on numbers, names and intent matters more than WER, so set targets per use case.

Does mixing Arabic and English confuse voice AI?

It is harder for monolingual models. Models trained on code-switched speech handle it better, so include mixed-language calls in any test.

Sources

  1. arXiv: Wang, Alhmoud and Alqurishi, Open Universal Arabic ASR Leaderboard
  2. IEEE DataPort: SADA, Saudi Audio Dataset for Arabic
  3. arXiv: Djanibekov et al., Dialectal Coverage and Generalization in Arabic Speech Recognition
  4. Journal of King Saud University: Al-Rojaie, The Effect of Social Factors on Sound Change in Najdi Arabic
  5. arXiv: Talafha et al., Casablanca: Data and Models for Multidialectal Arabic Speech Recognition
  6. arXiv: Hamed et al., A Survey of Code-switched Arabic NLP
  7. arXiv: Sukhadia and Umesh, Channel-Aware Pretraining for Telephonic-Speech ASR
  8. MIT CSAIL: The Arabic Dialect Identification for 17 Countries (ADI17) Dataset
  9. ISCA Interspeech 2021: Chowdhury et al., Multilingual Strategy for Dialectal Code-switching Arabic ASR

Hear Ranen on your own calls

Book a demo and we run Ranen on a sample of your call recordings, then size the plan with you.

Related articles

All articles