AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicVoices

AI4Bharat
gated / Overall score: 3.7

IndicVoices is a speech corpus of read, extempore and conversational audio in all 22 scheduled Indian languages, recorded from about 51,000 speakers in more than 400 districts. More than 23,000 hours have been recorded, and about half of them transcribed by people under published guidelines. AI4Bharat collected it with its open Kathbath and Shoonya tools and trained its 22-language IndicASR model on it.

Hour and speaker totals grow with each release; the card figures are used here, and they are well above the paper's.

Openness

3 high confidence
3.0
license
cc-by-4.0(card "License" section and metadata)
access
auto(Hugging Face gated: auto
dataset_card
present(card describes collection, composition and file layout

The recordings and transcripts carry CC BY 4.0. The Hugging Face copy asks for an access request that is approved automatically, and the card also links the full-length audio as open archive downloads.

Adoption

3 high confidence
3.0

Hugging Face downloads of the IndicVoices repository. Full-length audio is also served from AI4Bharat's own object store, which publishes no count, so downloads from there are not included.

Capability

4 medium confidence
4.0

IndicVoices has more human-transcribed spontaneous speech in all twenty-two scheduled Indian languages than any other open corpus, a paper documents its collection and transcription, and its builders made IndicASR, the first recognizer to cover all of them, on it. About eleven thousand of its hours are transcribed, a large set for these languages but a tenth or less of the hundreds of thousands of hours in the biggest English speech corpora.

Verified 2026-09-24