ARTPARK, Indian Institute of Science
lab · IndiaScores
1 product on the map — 1 open-ish.
Openness
3 high confidence- license
- cc-by-4.0(metadata of all three repositories)
- access
- auto(Hugging Face gated: auto, behind a form asking first and last name, country, company and intended use)
- dataset_card
- present(card gives per-district audio and transcription hours and language breakdowns)
All three Vaani repositories carry CC BY 4.0. Each sits behind a form asking for name, country, company and intended use, which Hugging Face approves automatically, so the data is open but not anonymous to fetch.
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
"gated": "auto"; license "cc-by-4.0"; extra_gated_fields: First Name, Last Name, Country, Company, "I want to use this dataset for" (Research, Academic, Commercial, Education).
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
"gated": "auto"; tag "license:cc-by-4.0".
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
"gated": "auto"; tag "license:cc-by-4.0".
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
Card: "~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts"; per-district audio and transcription hours.
Adoption
3 high confidenceHugging Face downloads summed over the main corpus, the transcribed-only subset and the benchmark, most of them for the main corpus. Copies distributed through Bhashini are not counted.
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
18629 downloads in the trailing 30 days for ARTPARK-IISc/Vaani
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
180 downloads in the trailing 30 days for ARTPARK-IISc/Vaani-Benchmark-V1.0
- https://huggingface.co/api/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
1626 downloads in the trailing 30 days for ARTPARK-IISc/Vaani-transcription-part
Capability
3 medium confidenceVaani records image-prompted speech in about a hundred and five languages from a hundred and sixty-five districts, more hours and far more languages than any other Indian speech corpus here, and ARTPARK's SraVaani and Whisper-Vaani recognizers and a Hindi recognition leaderboard are built on it. Only about two thousand of its hours are transcribed, roughly one in fifteen, well short of IndicVoices and a sliver of the hundreds of thousands in the largest English speech corpora.
- https://arxiv.org/abs/2603.28714 recorded 2026-09-24
Abstract: "We release approximately 289K images, 31,255 hours of speech, and 2,043 hours of transcribed audio spanning 105 languages from 28 states and 3 union territories."
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani recorded 2026-09-24
"Models trained or fine-tuned on ARTPARK-IISc/Vaani" lists SraVaani-1.0, whisper-large-v3-vaani-hindi, Vaani-LID_v0.
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0 recorded 2026-09-24
"A curated ASR evaluation set drawn from the Vaani project ... 5,050 audio segments from 1,103 speakers ... each with three independent human transcriptions."
- https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part recorded 2026-09-24
"This dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages."