IndicVoices
AI4BharatIndicVoices is a speech corpus of read, extempore and conversational audio in all 22 scheduled Indian languages, recorded from about 51,000 speakers in more than 400 districts. More than 23,000 hours have been recorded, and about half of them transcribed by people under published guidelines. AI4Bharat collected it with its open Kathbath and Shoonya tools and trained its 22-language IndicASR model on it.
Hour and speaker totals grow with each release; the card figures are used here, and they are well above the paper's.
Openness
3 high confidence- license
- cc-by-4.0(card "License" section and metadata)
- access
- auto(Hugging Face gated: auto
- dataset_card
- present(card describes collection, composition and file layout
The recordings and transcripts carry CC BY 4.0. The Hugging Face copy asks for an access request that is approved automatically, and the card also links the full-length audio as open archive downloads.
- https://huggingface.co/api/datasets/ai4bharat/IndicVoices recorded 2026-09-24
"gated": "auto"; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/ai4bharat/IndicVoices recorded 2026-09-24
"This dataset is released under the CC BY 4.0." Card links full-length audio at iv-release.objectstore.e2enetworks.net archives v1 to v5.
- https://raw.githubusercontent.com/AI4Bharat/IndicVoices/master/README.md recorded 2026-09-24
README describes the Kathbath collection platform, the Shoonya transcription platform and IndicASR training steps.
Adoption
3 high confidenceHugging Face downloads of the IndicVoices repository. Full-length audio is also served from AI4Bharat's own object store, which publishes no count, so downloads from there are not included.
- https://huggingface.co/api/datasets/ai4bharat/IndicVoices recorded 2026-09-24
25056 downloads in the trailing 30 days for ai4bharat/IndicVoices
Capability
4 medium confidenceIndicVoices has more human-transcribed spontaneous speech in all twenty-two scheduled Indian languages than any other open corpus, a paper documents its collection and transcription, and its builders made IndicASR, the first recognizer to cover all of them, on it. About eleven thousand of its hours are transcribed, a large set for these languages but a tenth or less of the hundreds of thousands of hours in the biggest English speech corpora.
- https://arxiv.org/abs/2403.01926 recorded 2026-09-24
Abstract: "Using INDICVOICES, we build IndicASR, the first ASR model to support all the 22 languages listed in the 8th schedule".
- https://huggingface.co/datasets/ai4bharat/IndicVoices recorded 2026-09-24
"a total of 23.7K hours ... from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have already been transcribed."
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
"~31278 hours ... 105 languages. From this audio data, 2,122 hours of transcribed data".
Verified 2026-09-24