IndicVoices-R
AI4BharatIndicVoices-R is a text-to-speech corpus of 1,704 hours from 10,496 speakers across the 22 scheduled Indian languages. AI4Bharat built it by restoring recordings from the IndicVoices speech recognition collection with demixing, dereverberation and enhancement models, then filtering for clarity. Each 48 kHz clip carries verbatim and normalized transcripts and speaker metadata, and a benchmark tests speaker generalization.
Openness
3 high confidence- license
- cc-by-4.0(card License section and metadata)
- access
- auto(Hugging Face gate asks for contact information and approves automatically)
- dataset_card
- present(card describes the pipeline, format and benchmark)
The audio and transcripts are under an attribution-only license, and the Hugging Face form that asks for contact details approves requests automatically. The processing pipeline is described on the card and in the paper.
- https://huggingface.co/api/datasets/ai4bharat/indicvoices_r recorded 2026-09-24
gated: auto; cardData license: cc-by-4.0
- https://huggingface.co/datasets/ai4bharat/indicvoices_r recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; License: "CC-BY-4.0"; data processing pipeline listed (HTDemucs, VoiceFixer, DeepFilterNet3)
Adoption
2 high confidenceHugging Face downloads of the corpus repository through the automatic gate. Language-specific re-uploads by other organizations and use through trained voices are not counted.
- https://huggingface.co/api/datasets/ai4bharat/indicvoices_r recorded 2026-09-24
7415 downloads in the trailing 30 days for ai4bharat/indicvoices_r
Capability
3 medium confidenceIndicVoices-R turns recognition recordings into a speech synthesis corpus for all twenty-two scheduled Indian languages with more than ten thousand voices, and a voice model from its own team is trained on it. At under two thousand hours of enhanced field audio it is small beside English speech corpora of hundreds of thousands of hours, and it is derived from IndicVoices, the larger corpus it restores.
- https://arxiv.org/abs/2409.05356 recorded 2026-09-24
"the largest multilingual Indian TTS dataset derived from an ASR dataset, with 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages"
- https://huggingface.co/datasets/ai4bharat/indicvoices_r recorded 2026-09-24
Models trained on the dataset include ai4bharat/IndicF5, ai4bharat/Cadence-Fast and ai4bharat/Cadence
Verified 2026-09-24