Hugging Face and Voice Arena have announced a significant expansion of the Open ASR Leaderboard, introducing Hindi and Indian English evaluation datasets in a partnership that marks the first inclusion of a Global South language on the multilingual benchmark. The Monsoon datasets comprise over 11 hours of audio from 4,888 speakers across hundreds of districts in India, making this the most demographically diverse automatic speech recognition benchmark to date and the first to include an Indic language.
The initiative directly addresses a well-documented but often invisible problem in speech recognition technology: ASR systems exhibit substantial performance disparities across demographic groups. Prior research has shown commercial systems perform roughly twice as poorly for Black speakers compared to white speakers, with additional gaps by gender, age, and accent. Traditional benchmarks lack demographic metadata about speakers, rendering these disparities invisible to developers. The Monsoon datasets were deliberately designed to vary along nine axes—geography, age, gender, vocabulary, device type, acoustic environment, speech rate, speech content, and transcription ambiguity—enabling researchers to measure how performance diverges across populations rather than reporting only aggregate accuracy.
The datasets follow the leaderboard's established methodology with both public and private splits to prevent benchmark overfitting, and include extensive speaker attributes including occupation, education, income, and device information. Hindi speakers comprise more than half a billion people globally, and the inclusion signals a broader industry push to develop speech recognition technology for underrepresented languages and populations. The public splits are immediately available for the open-source community to begin evaluation and development.
Key Points
Hindi becomes the first Global South and Indic language on the Open ASR Leaderboard
Monsoon datasets comprise 4,888 speakers with 12+ attributes recorded per contributor to expose performance disparities
Datasets deliberately vary across nine demographic and acoustic dimensions to reveal how ASR systems fail for specific populations
Addresses documented racial, gender, age, and accent-based disparities in commercial speech recognition systems
Public and private splits released simultaneously to enable benchmarking while preventing benchmark-specific model optimization