Lahaja
AI4BharatLahaja is a Hindi speech recognition benchmark of 12.5 hours of read and extempore speech from 132 speakers across 83 districts of India, built to test models on regional accents. Each clip carries verbatim and normalized transcripts with the speaker's native language, district and occupation. AI4Bharat built it and released four NeMo baseline models with it.
Openness
3 medium confidence- license
- cc-by-4.0(card License section)+mit(card metadata tag
- access
- auto(Hugging Face gate asks for contact information and approves automatically)
- dataset_card
- present(card and GitHub README describe the recordings and metadata fields)
Both licenses the card names allow commercial reuse with attribution, and the Hugging Face form approves requests automatically. The card should say which one governs.
- https://huggingface.co/api/datasets/ai4bharat/Lahaja recorded 2026-09-24
gated: auto; cardData license: mit
- https://huggingface.co/datasets/ai4bharat/Lahaja recorded 2026-09-24
License: mit (header); "This dataset is released under the CC BY 4.0."; "You need to agree to share your contact information to access this dataset"
- https://raw.githubusercontent.com/AI4Bharat/Lahaja/master/README.md recorded 2026-09-24
Metadata CSV fields: verbatim, normalized, native_language, native_district, occupation; links to models M1-M4
Adoption
1 high confidenceHugging Face downloads of the benchmark repository through the automatic gate. An evaluation set is downloaded once per evaluation setup, so downloads say little about how often scores on it are reported.
- https://huggingface.co/api/datasets/ai4bharat/Lahaja recorded 2026-09-24
232 downloads in the trailing 30 days for ai4bharat/Lahaja
Capability
3 medium confidenceLahaja is a documented test set for one language and one question, how Hindi recognition copes with regional accents. No benchmark or model outside its paper reports on it, and Kathbath covers 12 languages with far more speech.
- https://arxiv.org/abs/2408.11440 recorded 2026-09-24
"a total of 12.5 hours of Hindi audio, sourced from 132 speakers spanning 83 districts of India"
Verified 2026-09-24