NVIDIA’s ASR Model Achieves Major Improvement in Saudi Dialect Recognition

Accra:NVIDIA announced a significant improvement in its "Nemotron 3.5 ASR" multilingual automatic speech recognition model through the use of the Saudi Audio Dataset for Arabic (SADA), developed by the Saudi Data and AI Authority (SDAIA) and the Saudi Broadcasting Authority.

According to Saudi Press Agency, the integration of the SADA dataset, which includes 133.7 hours of Najdi and Hijazi audio, reduced comprehension errors in Saudi dialects by nearly half. Prior to the dataset training, the model incorrectly identified about 55 out of every 100 words in Saudi dialects, which decreased to approximately 30 words post-training.

Overall, the model's error rate across multiple dialects dropped from 58.8% to 35.6%, with character-level errors reducing from 31.6% to 12.2%. The training process took just four and a half hours using two GPUs.

The improved model, designed for real-time interactive systems, boasts a latency starting at 80 milliseconds, supporting various applications such as intelligent voice assistants, conversational agents, and live broadcast subtitling. NVIDIA has also provided workflows and tools for developers to apply the methodology to other languages.

SDAIA's SADA dataset, released on Kaggle, comprises 667 hours of transcribed audio, including 600 hours from the Saudi Broadcasting Authority. It spans over 10 Saudi dialects, with 20 hours reserved for validation, and supports a range of audio model developments and the enrichment of digital Arabic content.

NVIDIA's use of the SADA dataset underscores the importance of high-quality national data in improving AI systems' alignment with local dialects and enhancing technology for Arabic speakers.