AI Potluck
Back to Gap Map Model components / Language-specific datasets

Kathbath

AI4Bharat
gated / Overall score: 2.7

Kathbath is a speech recognition corpus of 1,684 hours of labeled speech in 12 Indian languages, recorded by 1,218 contributors in 203 districts. Contributors read aloud sentences drawn from IndicCorp text, collected through the Karya crowdsourcing platform, and the corpus includes clean and noisy test sets. AI4Bharat built it as the data behind IndicSUPERB, its benchmark of six speech tasks.

Openness

3 medium confidence
3.0
license
cc0-1.0(card "Licensing Information": CC0 for the packaging
access
auto(Hugging Face gated: auto
dataset_card
present(card names languages, contributors, text source and license

AI4Bharat waives its rights under CC0 while noting that it does not own the IndicCorp source text. The Hugging Face copy needs an access request that is approved automatically, and the IndicSUPERB README links the same audio as open downloads.

Adoption

2 high confidence
2.0

Hugging Face downloads of the Kathbath repository. The same audio is also served as direct archives linked from the IndicSUPERB README, which publish no count, so those downloads are not included.

Capability

3 medium confidence
3.0

Kathbath gives twelve Indian languages crowdsourced read speech that the IndicSUPERB paper documents and turns into six speech benchmarks, and NVIDIA's Magpie multilingual speech synthesis model lists it as training data. At under two thousand hours it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices from the same team adds spontaneous speech in all twenty-two scheduled languages.

  • https://arxiv.org/abs/2208.11761 recorded 2026-09-24

    Abstract: "We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts ... we create benchmarks across 6 speech tasks".

  • https://huggingface.co/datasets/ai4bharat/Kathbath recorded 2026-09-24

    "Models trained or fine-tuned on ai4bharat/Kathbath" lists nvidia/magpie_tts_multilingual_357m, shunyalabs/zero-stt-hinglish.

Verified 2026-09-24