Voice AI8 min read
AI Voice Agents for Saudi Contact Centers: The Complete Guide
How AI voice agents work in Saudi contact centers: speech recognition, actions in your systems, handover, Arabic dialects, PDPL data residency and launch.
Voice AI
Why Arabic dialects challenge speech recognition on phone calls, how modern voice AI handles them, and how to test a vendor on your own call recordings.
Saudi callers rarely speak textbook Arabic to a contact center. They speak Najdi, Hijazi or another Gulf variety, switch to English for a product name or an error message, read out an ID number, and often do it over a phone line from a car. A voice agent has to understand all of that before it can complete anything.
Public benchmarks show the size of the problem. In the Open Universal Arabic ASR Leaderboard, the best-ranked open-source model reached a word error rate (WER) of 19.23% on Modern Standard Arabic (MSA) in the Saudi SADA dataset, but 36.34% on Najdi and 48.23% on Khaliji speech from the same dataset. This article explains why, how modern systems narrow the gap, and how to test a vendor on your own calls.
Dialect, code-switching, numbers and telephone audio each add errors, and they compound. Published scores on broadcast or read speech do not predict accuracy on your calls. The reliable test is a sample of your own recordings, scored by dialect, by critical details such as ID numbers and by task outcome, not by WER alone.
Most transcribed Arabic speech comes from news and broadcast in MSA. The leaderboard authors note that in the widely used MGB-2 corpus "over 70% of the samples represent the MSA dialect," and link the drop on dialects to "the imbalanced distribution of dialects in most public datasets." Djanibekov and colleagues reach a similar conclusion: existing systems "mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects."
| Variety (SADA test data) | WER, best open model |
|---|---|
| Modern Standard Arabic | 19.23% |
| Najdi | 36.34% |
| Hijazi | 36.96% |
| Egyptian | 40.97% |
| Khaliji | 48.23% |
These figures come from television audio, not phone calls, and from open-source models, not commercial systems. They show the direction of the gap, not the accuracy of any specific vendor.
The same letter is pronounced differently across regions. A study in the Journal of King Saud University describes how, in Najdi Arabic, the classical qaf is realized as [g], and in the Qassimi dialect it can be further fronted to [dz] in certain environments. Written Arabic also omits most short vowels, and the leaderboard paper lists the "lack of diacritics in written text" among the core difficulties. A model trained mainly on MSA pronunciation hears these regular patterns as different words.
Everyday words for "want", "now" or "how much" differ by region, and dialects have no standard spelling. The Casablanca dataset, covering eight dialects, documents this lexical variation. Because the same dialect word can be written several ways, Hamed and colleagues note that word and character error rates alone are not adequate for unstandardized orthography.
Callers mix Arabic and English inside one sentence, especially for products, technical terms and job titles. A survey of code-switched Arabic NLP lists Saudi Arabia among the countries where Arabic-English switching is observed, and reports WER "in the range of 24.8-53.8%" across code-switched Arabic speech corpora. A monolingual model sees two sound systems and two scripts in one utterance.
The details a contact center needs most are the hardest: national ID and Iqama numbers, phone numbers, IBANs, order numbers, amounts, Hijri and Gregorian dates, and personal names. Callers group digits differently, mix Arabic and English numbers, and pronounce the same name in several ways. One wrong digit breaks a lookup even when the overall WER looks acceptable.
Phone calls are narrowband, sampled at about 8 kHz, while many speech models are trained on 16 kHz audio. Researchers at IIT Madras note that "wide band (sampled at ∼16 kHz) ASR models do not perform well for narrow band speech data," and that labelled narrowband data is scarce. Speakerphones, road noise, compression and packet loss add further errors on top of dialect.
| Measure | What it shows | Watch for |
|---|---|---|
| WER per dialect and channel | Transcription accuracy for each caller group | Spelling variants inflate it; normalize first |
| Entity accuracy | Whether IDs, numbers, dates and names are exact | One wrong digit is a failure |
| Intent accuracy | Whether the agent understood the request | Can be high even when WER is moderate |
| Task completion | Whether the request was completed end to end | Needs live or simulated calls |
| Reply dialect | Whether answers match the caller's variety | Rated by native listeners |
For how this differs from menu-based systems, see AI voice agent vs IVR, and for the wider picture, our guide to AI voice agents in Saudi contact centers.
Ranen understands Modern Standard Arabic and Gulf, Egyptian, Levantine and Maghrebi dialects, plus English, including Arabic and English mixed in one sentence, and replies in the caller's dialect. Card and ID numbers are masked automatically in transcripts and recordings, which are hosted in data centers inside Saudi Arabia. Callers can ask for a person at any time. We run demos on your own call recordings, so you can apply the test above before committing. Request a demo.
Yes, systems trained on dialect and telephone data can, but accuracy varies by dialect, audio quality and domain. Test on your own calls before deciding.
There is no universal threshold. For a contact center, accuracy on numbers, names and intent matters more than WER, so set targets per use case.
It is harder for monolingual models. Models trained on code-switched speech handle it better, so include mixed-language calls in any test.
Book a demo and we run Ranen on a sample of your call recordings, then size the plan with you.
Voice AI8 min read
How AI voice agents work in Saudi contact centers: speech recognition, actions in your systems, handover, Arabic dialects, PDPL data residency and launch.
Voice AI6 min read
AI voice agent vs IVR phone menus: how each works, a side-by-side comparison, where IVR still fits, and how Saudi contact centers can move from menu to agent.
Voice AI7 min read
When a voice agent should hand over to a human: the triggers, what context to pass, warm vs cold transfer, common mistakes, and how to measure handover quality.