Sarvam AI launches dataset to test speech recognition in 22 Indian languages

Sarvam AI, in collaboration with AI4Bharat, has introduced Indic DiarBench, an open benchmark dataset to evaluate automatic speech recognition (ASR) and speaker diarisation across all 22 scheduled Indian languages. The dataset contains 108 hours of natural, multi-speaker conversations featuring 485 unique speakers from 189 districts across urban and rural India, and is available on Hugging Face.

Load More