AI Potluck
Back to Gap Map Model components / Language-specific datasets

Shrutilipi

AI4Bharat
gated / Overall score: 2.7

Shrutilipi is a labeled speech recognition corpus of more than 6,400 hours in 12 Indian languages, from Bengali and Hindi to Sanskrit and Urdu. Its audio and sentence transcripts were mined from All India Radio news bulletins by aligning broadcast audio with the published bulletin text. AI4Bharat built it, and the paper reports 4.95 million sentences.

Openness

3 high confidence
3.0
license
cc-by-4.0(card License section and metadata)
access
auto(Hugging Face gate asks for contact information and approves automatically)
dataset_card
present(card describes the source and languages

The whole corpus is under an attribution-only license, and anyone who shares contact details on Hugging Face gets the files at once. The gate is a form, not a review.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single corpus repository. Downloads through the gate count, but copies mirrored elsewhere and use through derived checkpoints do not.

Capability

3 medium confidence
3.0

Shrutilipi adds thousands of hours of transcribed news speech in 12 Indian languages, its paper shows it improves recognition on IndicSUPERB, and community speech models are trained on it. It covers one domain, read broadcast news, and the same team's IndicVoices supersedes it with spontaneous speech in all 22 scheduled languages.

Verified 2026-09-24