AI Potluck
Back to Gap Map Model components / Language-specific datasets

IndicVoices-R

AI4Bharat
gated / Overall score: 2.7

IndicVoices-R is a text-to-speech corpus of 1,704 hours from 10,496 speakers across the 22 scheduled Indian languages. AI4Bharat built it by restoring recordings from the IndicVoices speech recognition collection with demixing, dereverberation and enhancement models, then filtering for clarity. Each 48 kHz clip carries verbatim and normalized transcripts and speaker metadata, and a benchmark tests speaker generalization.

Openness

3 high confidence
3.0
license
cc-by-4.0(card License section and metadata)
access
auto(Hugging Face gate asks for contact information and approves automatically)
dataset_card
present(card describes the pipeline, format and benchmark)

The audio and transcripts are under an attribution-only license, and the Hugging Face form that asks for contact details approves requests automatically. The processing pipeline is described on the card and in the paper.

Adoption

2 high confidence
2.0

Hugging Face downloads of the corpus repository through the automatic gate. Language-specific re-uploads by other organizations and use through trained voices are not counted.

Capability

3 medium confidence
3.0

IndicVoices-R turns recognition recordings into a speech synthesis corpus for all twenty-two scheduled Indian languages with more than ten thousand voices, and a voice model from its own team is trained on it. At under two thousand hours of enhanced field audio it is small beside English speech corpora of hundreds of thousands of hours, and it is derived from IndicVoices, the larger corpus it restores.

Verified 2026-09-24