NVIDIA fine-tunes speech recognition for Najdi and Hijazi Arabic
NVIDIA says it adapted its Nemotron 3.5 ASR speech-recognition model using Saudi Arabic audio, reducing word and character error rates on a test split focused on Najdi and Hijazi dialects.

Key takeaways · 4
- 01
The fine-tuning used 133.7 hours of Najdi and Hijazi audio from the Saudi Audio Dataset for Arabic.
- 02
NVIDIA reports a 25.09 percentage-point drop in word error rate on the target test split.
- 03
NVIDIA says latency for the enhanced model starts at 80 milliseconds.
- 04
NVIDIA shared workflows and tools intended to help developers replicate the method across other languages.
A focused dialect adaptation
NVIDIA used the Saudi Audio Dataset for Arabic to train Nemotron 3.5 ASR, with the initial experiment focused on Najdi and Hijazi.[1][2] The work used 133.7 hours of audio from those dialects.[1] NVIDIA’s developer page describes Nemotron 3.5 ASR as supporting multilingual streaming transcription across 40 language-locales, including Arabic.[2] The developer page gives September 30, 2026, as its date.[2] The dataset is a resource launched by the Saudi Data and AI Authority in partnership with the Saudi Broadcasting Authority.[1]
Reported error-rate changes
On NVIDIA’s Najdi and Hijazi target test split, word error rate fell from 55.05% to 29.96%.[2] The character-level error rate fell from 31.63% to 12.18%.[3] NVIDIA says the enhanced model’s latency starts at 80 milliseconds and describes it as suited to voice assistants, conversational agents, live broadcast subtitling, and transcription.[1] The reported improvements are tied to the target test split; the evidence does not establish the same error rates for other dialects or recording conditions.[2]
Training choices and trade-offs
The training took four and a half hours using two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.[1][2] NVIDIA says its adaptation workflow uses the NeMo framework and a fine-tuning recipe.[2] Its guidance says replay mixing can help limit catastrophic forgetting of previously learned data.[2] NVIDIA also says partial encoder unfreezing changes fewer parameters and is cheaper and faster than a full fine-tune, with some accuracy trade-off.[2] In the reported comparison, updating all 24 encoder layers achieved lower error rates than updating only the top eight.[4]
A larger dataset and wider use
The Saudi Audio Dataset for Arabic contains about 667 hours of transcribed audio and more than 125,000 categorized clips.[1] More than 600 hours came from 57 television programs and series covering more than 10 Saudi dialects.[1] The Saudi Data and AI Authority released the dataset on Kaggle to support researchers and developers.[1] NVIDIA says regional dialects and local recording conditions are often underrepresented in speech-recognition training data.[2] It has shared workflows and tools for developers to replicate the method across other languages.[1]
For teams evaluating speech recognition, the results offer a concrete dialect-focused benchmark and a reported training setup to investigate. They do not establish performance across all dialects or real-world recording conditions, so validation on a team’s own audio remains important.
Why it matters
Test yourself on this story — 2 questions.
Create a free account to take the quiz, earn XP, and get a daily session built for your industry.
Take the quizHow this developed
4 October 2026
NVIDIA fine-tunes speech recognition for Najdi and Hijazi Arabic
Sources
- Arab News | Saudi dataset cuts AI errors in Arabic dialectsarabnews.com
- Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages | NVIDIA Technical Blogdeveloper.nvidia.com
- NVIDIA's AI Models Master Najdi and Hijazi Dialects: Reducing Speech Recognition Errors by More Than Half | Sabqsabq.org
- NVIDIA halves Saudi dialect errors with SDAIA audio datasetmiddleeastainews.com