Shrutilipi
AI4BharatShrutilipi is a labeled speech recognition corpus of more than 6,400 hours in 12 Indian languages, from Bengali and Hindi to Sanskrit and Urdu. Its audio and sentence transcripts were mined from All India Radio news bulletins by aligning broadcast audio with the published bulletin text. AI4Bharat built it, and the paper reports 4.95 million sentences.
Openness
3 high confidence- license
- cc-by-4.0(card License section and metadata)
- access
- auto(Hugging Face gate asks for contact information and approves automatically)
- dataset_card
- present(card describes the source and languages
The whole corpus is under an attribution-only license, and anyone who shares contact details on Hugging Face gets the files at once. The gate is a form, not a review.
- https://huggingface.co/api/datasets/ai4bharat/Shrutilipi recorded 2026-09-24
gated: auto; cardData license: cc-by-4.0
- https://huggingface.co/datasets/ai4bharat/Shrutilipi recorded 2026-09-24
"You need to agree to share your contact information to access this dataset"; "This dataset is released under the CC BY 4.0."
Adoption
2 high confidenceHugging Face downloads of the single corpus repository. Downloads through the gate count, but copies mirrored elsewhere and use through derived checkpoints do not.
- https://huggingface.co/api/datasets/ai4bharat/Shrutilipi recorded 2026-09-24
2692 downloads in the trailing 30 days for ai4bharat/Shrutilipi
Capability
3 medium confidenceShrutilipi adds thousands of hours of transcribed news speech in 12 Indian languages, its paper shows it improves recognition on IndicSUPERB, and community speech models are trained on it. It covers one domain, read broadcast news, and the same team's IndicVoices supersedes it with spontaneous speech in all 22 scheduled languages.
- https://arxiv.org/abs/2208.12666 recorded 2026-09-24
"Shrutilipi, a dataset which contains over 6,400 hours of labelled audio across 12 Indian languages totalling to 4.95M sentences"; mined from All India Radio archives
- https://huggingface.co/datasets/ai4bharat/Shrutilipi recorded 2026-09-24
"Models trained or fine-tuned on ai4bharat/Shrutilipi": shunyalabs/zero-stt-hinglish, lemuralabs/tamil-asr-qwen3 and others
Verified 2026-09-24