Skip to main content

NVIDIA fine-tunes speech recognition for Najdi and Hijazi Arabic

4 OCTOBER 2026·2 MIN READ·4 SOURCES

NVIDIA says it adapted its Nemotron 3.5 ASR speech-recognition model using Saudi Arabic audio, reducing word and character error rates on a test split focused on Najdi and Hijazi dialects.

NVIDIA fine-tunes speech recognition for Najdi and Hijazi Arabic

Key takeaways · 4

  • 01

    The fine-tuning used 133.7 hours of Najdi and Hijazi audio from the Saudi Audio Dataset for Arabic.

  • 02

    NVIDIA reports a 25.09 percentage-point drop in word error rate on the target test split.

  • 03

    NVIDIA says latency for the enhanced model starts at 80 milliseconds.

  • 04

    NVIDIA shared workflows and tools intended to help developers replicate the method across other languages.

A focused dialect adaptation

NVIDIA used the Saudi Audio Dataset for Arabic to train Nemotron 3.5 ASR, with the initial experiment focused on Najdi and Hijazi.[1][2] The work used 133.7 hours of audio from those dialects.[1] NVIDIA’s developer page describes Nemotron 3.5 ASR as supporting multilingual streaming transcription across 40 language-locales, including Arabic.[2] The developer page gives September 30, 2026, as its date.[2] The dataset is a resource launched by the Saudi Data and AI Authority in partnership with the Saudi Broadcasting Authority.[1]

Reported error-rate changes

On NVIDIA’s Najdi and Hijazi target test split, word error rate fell from 55.05% to 29.96%.[2] The character-level error rate fell from 31.63% to 12.18%.[3] NVIDIA says the enhanced model’s latency starts at 80 milliseconds and describes it as suited to voice assistants, conversational agents, live broadcast subtitling, and transcription.[1] The reported improvements are tied to the target test split; the evidence does not establish the same error rates for other dialects or recording conditions.[2]

Training choices and trade-offs

The training took four and a half hours using two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.[1][2] NVIDIA says its adaptation workflow uses the NeMo framework and a fine-tuning recipe.[2] Its guidance says replay mixing can help limit catastrophic forgetting of previously learned data.[2] NVIDIA also says partial encoder unfreezing changes fewer parameters and is cheaper and faster than a full fine-tune, with some accuracy trade-off.[2] In the reported comparison, updating all 24 encoder layers achieved lower error rates than updating only the top eight.[4]

A larger dataset and wider use

The Saudi Audio Dataset for Arabic contains about 667 hours of transcribed audio and more than 125,000 categorized clips.[1] More than 600 hours came from 57 television programs and series covering more than 10 Saudi dialects.[1] The Saudi Data and AI Authority released the dataset on Kaggle to support researchers and developers.[1] NVIDIA says regional dialects and local recording conditions are often underrepresented in speech-recognition training data.[2] It has shared workflows and tools for developers to replicate the method across other languages.[1]

For teams evaluating speech recognition, the results offer a concrete dialect-focused benchmark and a reported training setup to investigate. They do not establish performance across all dialects or real-world recording conditions, so validation on a team’s own audio remains important.

Why it matters
Story quiz

Test yourself on this story — 2 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 4 October 2026

    NVIDIA fine-tunes speech recognition for Najdi and Hijazi Arabic

Sources

AI fluency, one session a day, built for your work.