AI Potluck
Back to Gap Map Model components / Language-specific datasets

SPRING-INX

SPRING Lab, IIT Madras
restricted / Overall score: 2.7

SPRING-INX is a speech recognition corpus of about 2,000 hours of manually transcribed speech in ten Indian languages, from Assamese and Bengali to Punjabi and Tamil. Agencies recorded monologues and mock conversations across at least ten domains under common guidelines, and a KL University team checked quality. SPRING Lab at IIT Madras releases it per language, with ESPnet baseline recipes.

The Hugging Face repositories have empty cards, so the corpus is read from its paper; no Hindi repository was found.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(cards and API metadata of all sixteen repositories carry no license
access
public(all repositories ungated)
dataset_card
no(cards contain only generated schema metadata)

Every language downloads without a gate, and the paper calls the release open source. No repository states a license, so the terms a reuser would rely on are not written down.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the sixteen per-language repositories. Use through the ESPnet recipes or copies held elsewhere is not counted.

Capability

3 medium confidence
3.0

SPRING-INX supplies about two thousand hours of hand-transcribed speech in ten Indian languages, documented only by a short paper. No named model reports training on it, and IndicVoices covers all 22 scheduled languages with far more transcribed speech.

  • https://arxiv.org/abs/2310.14654 recorded 2026-09-24

    "about 2000 hours of legally sourced and manually transcribed speech data for ASR system building in Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi and Tamil"

  • https://arxiv.org/pdf/2310.14654 recorded 2026-09-24

    Table 1: per-language hours from 61 (Assamese) to 420 (Bengali); ESPnet recipes at github.com/Speech-Lab-IITM

Verified 2026-09-24