AI Potluck
Back to Gap Map Model components / Language-specific datasets

Project Vaani

ARTPARK, Indian Institute of Science
gated / Overall score: 3.0

Project Vaani is a speech corpus of spontaneous, image-prompted speech from about 156,000 speakers in 165 Indian districts, spanning some 105 languages. It holds over 31,000 hours of audio, about 2,000 hours of it transcribed, together with the roughly 289,000 prompt images. IISc Bangalore and ARTPARK run the project with funding from Google, and also publish a transcribed-only subset and a Hindi speech-recognition benchmark.

Openness

3 high confidence
3.0
license
cc-by-4.0(metadata of all three repositories)
access
auto(Hugging Face gated: auto, behind a form asking first and last name, country, company and intended use)
dataset_card
present(card gives per-district audio and transcription hours and language breakdowns)

All three Vaani repositories carry CC BY 4.0. Each sits behind a form asking for name, country, company and intended use, which Hugging Face approves automatically, so the data is open but not anonymous to fetch.

Adoption

3 high confidence
3.0

Hugging Face downloads summed over the main corpus, the transcribed-only subset and the benchmark, most of them for the main corpus. Copies distributed through Bhashini are not counted.

Capability

3 medium confidence
3.0

Vaani records image-prompted speech in about a hundred and five languages from a hundred and sixty-five districts, more hours and far more languages than any other Indian speech corpus here, and ARTPARK's SraVaani and Whisper-Vaani recognizers and a Hindi recognition leaderboard are built on it. Only about two thousand of its hours are transcribed, roughly one in fifteen, well short of IndicVoices and a sliver of the hundreds of thousands in the largest English speech corpora.

Verified 2026-09-24