AI Potluck
Back to Gap Map Model components / Language-specific datasets

Lahaja

AI4Bharat
gated / Overall score: 2.4

Lahaja is a Hindi speech recognition benchmark of 12.5 hours of read and extempore speech from 132 speakers across 83 districts of India, built to test models on regional accents. Each clip carries verbatim and normalized transcripts with the speaker's native language, district and occupation. AI4Bharat built it and released four NeMo baseline models with it.

Openness

3 medium confidence
3.0
license
cc-by-4.0(card License section)+mit(card metadata tag
access
auto(Hugging Face gate asks for contact information and approves automatically)
dataset_card
present(card and GitHub README describe the recordings and metadata fields)

Both licenses the card names allow commercial reuse with attribution, and the Hugging Face form approves requests automatically. The card should say which one governs.

Adoption

1 high confidence
1.0

Hugging Face downloads of the benchmark repository through the automatic gate. An evaluation set is downloaded once per evaluation setup, so downloads say little about how often scores on it are reported.

Capability

3 medium confidence
3.0

Lahaja is a documented test set for one language and one question, how Hindi recognition copes with regional accents. No benchmark or model outside its paper reports on it, and Kathbath covers 12 languages with far more speech.

Verified 2026-09-24