Rasa
AI4BharatRasa is an expressive text-to-speech corpus covering the 22 officially recognized languages of India, with a female and a male voice per language and about 1,145 hours in total. Its 48 kHz recordings mix neutral Wikipedia readings, speech in the six Ekman emotions, voice-assistant style commands, conversation, news reading and book narration. AI4Bharat built it, and the method appeared at Interspeech 2024.
The card cites an Interspeech 2024 paper but links no arXiv entry, so none is listed.
Openness
3 high confidence- license
- cc-by-4.0(card License section and metadata)
- access
- auto(Hugging Face gate asks for contact information and approves automatically)
- dataset_card
- present(card describes speaking styles and per-speaker hours)
Every recording is under an attribution-only license, released through a Hugging Face form that approves automatically once contact details are given.
- https://huggingface.co/api/datasets/ai4bharat/Rasa recorded 2026-09-24
gated: auto; cardData license: cc-by-4.0
- https://huggingface.co/datasets/ai4bharat/Rasa recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; License "CC-BY-4.0"; statistics table totals 1145.44 hours and 640,950 utterances
Adoption
3 high confidenceHugging Face downloads of the single corpus repository through the automatic gate. Use through voices trained on it, such as IndicF5, is not counted.
- https://huggingface.co/api/datasets/ai4bharat/Rasa recorded 2026-09-24
10167 downloads in the trailing 30 days for ai4bharat/Rasa
Capability
3 medium confidenceRasa records a female and a male studio voice in each of India's twenty-two official languages across emotions and speaking styles, and speech synthesis models from AI Bharat and SPRING Lab are trained on it. At about eleven hundred hours from a few dozen speakers it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices from the same team records tens of thousands of speakers.
- https://huggingface.co/datasets/ai4bharat/Rasa recorded 2026-09-24
"the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language"; "44 speaker-language pairs across all 22 Indian languages"; models trained include ai4bharat/IndicF5, SPRINGLab/Indic-Mio
Verified 2026-09-24