AI Potluck
Back to Gap Map Organization

ARTPARK, Indian Institute of Science

lab · India

Scores

1 product on the map — 1 open-ish.

Project Vaani

Openness

3 high confidence
3.0
license
cc-by-4.0(metadata of all three repositories)
access
auto(Hugging Face gated: auto, behind a form asking first and last name, country, company and intended use)
dataset_card
present(card gives per-district audio and transcription hours and language breakdowns)

All three Vaani repositories carry CC BY 4.0. Each sits behind a form asking for name, country, company and intended use, which Hugging Face approves automatically, so the data is open but not anonymous to fetch.

Adoption

3 high confidence
3.0

Hugging Face downloads summed over the main corpus, the transcribed-only subset and the benchmark, most of them for the main corpus. Copies distributed through Bhashini are not counted.

Capability

3 medium confidence
3.0

Vaani records image-prompted speech in about a hundred and five languages from a hundred and sixty-five districts, more hours and far more languages than any other Indian speech corpus here, and ARTPARK's SraVaani and Whisper-Vaani recognizers and a Hindi recognition leaderboard are built on it. Only about two thousand of its hours are transcribed, roughly one in fifteen, well short of IndicVoices and a sliver of the hundreds of thousands in the largest English speech corpora.