Project Vaani
ARTPARK, Indian Institute of ScienceProject Vaani is a speech corpus of spontaneous, image-prompted speech from about 156,000 speakers in 165 Indian districts, spanning some 105 languages. It holds over 31,000 hours of audio, about 2,000 hours of it transcribed, together with the roughly 289,000 prompt images. IISc Bangalore and ARTPARK run the project with funding from Google, and also publish a transcribed-only subset and a Hindi speech-recognition benchmark.
Openness
3 high confidence- license
- cc-by-4.0(metadata of all three repositories)
- access
- auto(Hugging Face gated: auto, behind a form asking first and last name, country, company and intended use)
- dataset_card
- present(card gives per-district audio and transcription hours and language breakdowns)
All three Vaani repositories carry CC BY 4.0. Each sits behind a form asking for name, country, company and intended use, which Hugging Face approves automatically, so the data is open but not anonymous to fetch.
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
"gated": "auto"; license "cc-by-4.0"; extra_gated_fields: First Name, Last Name, Country, Company, "I want to use this dataset for" (Research, Academic, Commercial, Education).
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
"gated": "auto"; tag "license:cc-by-4.0".
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
"gated": "auto"; tag "license:cc-by-4.0".
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
Card: "~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts"; per-district audio and transcription hours.
Adoption
3 high confidenceHugging Face downloads summed over the main corpus, the transcribed-only subset and the benchmark, most of them for the main corpus. Copies distributed through Bhashini are not counted.
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
18629 downloads in the trailing 30 days for ARTPARK-IISc/Vaani
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
180 downloads in the trailing 30 days for ARTPARK-IISc/Vaani-Benchmark-V1.0
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
1626 downloads in the trailing 30 days for ARTPARK-IISc/Vaani-transcription-part
Capability
3 medium confidenceVaani records image-prompted speech in about a hundred and five languages from a hundred and sixty-five districts, more hours and far more languages than any other Indian speech corpus here, and ARTPARK's SraVaani and Whisper-Vaani recognizers and a Hindi recognition leaderboard are built on it. Only about two thousand of its hours are transcribed, roughly one in fifteen, well short of IndicVoices and a sliver of the hundreds of thousands in the largest English speech corpora.
- https://arxiv.org/abs/2603.28714 recorded 2026-09-24
Abstract: "We release approximately 289K images, 31,255 hours of speech, and 2,043 hours of transcribed audio spanning 105 languages from 28 states and 3 union territories."
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
"Models trained or fine-tuned on ARTPARK-IISc/Vaani" lists SraVaani-1.0, whisper-large-v3-vaani-hindi, Vaani-LID_v0.
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
"A curated ASR evaluation set drawn from the Vaani project ... 5,050 audio segments from 1,103 speakers ... each with three independent human transcriptions."
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
"This dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages."
Verified 2026-09-24