
On September 30, 2026, NVIDIA shared an experiment using NVIDIA NeMo to fine-tune Nemotron 3.5 ASR for Saudi Arabic, with a focus on the Najdi and Hijazi dialects. Nemotron 3.5 ASR supports streaming speech recognition across 40 languages and locales. The technical blog focuses on how a limited amount of dialect data can reduce errors in a particular deployment context.
The workflow described in the article includes light data curation, weighted replay that mixes target-dialect data with FLEURS, duration bucketing, and fine-tuning with only part of the encoder unfrozen. NVIDIA reports that 133.7 hours of Najdi and Hijazi data reduced WER on the target split from 55.05% to 29.96%. English WER changed from 11.04% to 10.42%, and the other Arabic dialects did not regress in the reported experiment. These are tutorial-experiment results, not a guarantee of the same improvement on every dataset or production environment.
The fine-tuning configurations show a clear trade-off. The article says that unfreezing all 24 encoder layers produced the lowest error, but required training about 230.4 million parameters. Freezing more layers can reduce compute while potentially limiting adaptation. Weighted replay is intended to keep the model exposed to existing language data while learning the new dialects, but it only protects the distributions represented in the replay mix. It does not replace evaluation on real speakers, devices, noise conditions, or domain vocabulary.
For streaming use, NVIDIA also evaluates lookahead and beam search. The post reports that a configuration with 13 lookahead frames and beam-8 MALSD reduced WER by another 2.71 percentage points without retraining, with roughly 800 milliseconds of latency for batch transcription. Speech systems therefore cannot optimize offline accuracy alone. Live support, meeting captions, and voice control have to balance error rate, latency, GPU cost, and the user experience.
Nemotron 3 diarization is part of the same deployment path and supports up to eight speakers. Separating who said what and when can improve summaries, ticket creation, and compliance review for calls or meetings. But diarization errors can propagate into the transcript and attribution, so it still needs independent testing in the target languages, accents, and recording environments.
The practical value of the article is that dialect adaptation needs a complete evaluation chain. Teams must first check whether the data represents the real use case, then compare replay, layer-freezing, and inference settings, and measure the target dialect alongside other languages, latency, and resource cost. NVIDIA's result is a reproducible starting point, but it is a vendor tutorial and experiment report, not a universal guarantee for Saudi Arabic or other language deployments.



