Kathbath
AI4BharatKathbath is a speech recognition corpus of 1,684 hours of labeled speech in 12 Indian languages, recorded by 1,218 contributors in 203 districts. Contributors read aloud sentences drawn from IndicCorp text, collected through the Karya crowdsourcing platform, and the corpus includes clean and noisy test sets. AI4Bharat built it as the data behind IndicSUPERB, its benchmark of six speech tasks.
Openness
3 medium confidence- license
- cc0-1.0(card "Licensing Information": CC0 for the packaging
- access
- auto(Hugging Face gated: auto
- dataset_card
- present(card names languages, contributors, text source and license
AI4Bharat waives its rights under CC0 while noting that it does not own the IndicCorp source text. The Hugging Face copy needs an access request that is approved automatically, and the IndicSUPERB README links the same audio as open downloads.
- https://huggingface.co/api/datasets/ai4bharat/Kathbath recorded 2026-09-24
"gated": "auto"; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/ai4bharat/Kathbath recorded 2026-09-24
"We license the actual packaging of all this data under the Creative Commons CC0 license" and "We do not own any of the raw text used in creating this dataset."
- https://raw.githubusercontent.com/AI4Bharat/IndicSUPERB/master/README.md recorded 2026-09-24
README lists Train 85GB, Valid, Test Known and Test Unknown archives, clean and noisy, at objectstore.e2enetworks.net.
Adoption
2 high confidenceHugging Face downloads of the Kathbath repository. The same audio is also served as direct archives linked from the IndicSUPERB README, which publish no count, so those downloads are not included.
- https://huggingface.co/api/datasets/ai4bharat/Kathbath recorded 2026-09-24
4586 downloads in the trailing 30 days for ai4bharat/Kathbath
Capability
3 medium confidenceKathbath gives twelve Indian languages crowdsourced read speech that the IndicSUPERB paper documents and turns into six speech benchmarks, and NVIDIA's Magpie multilingual speech synthesis model lists it as training data. At under two thousand hours it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices from the same team adds spontaneous speech in all twenty-two scheduled languages.
- https://arxiv.org/abs/2208.11761 recorded 2026-09-24
Abstract: "We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts ... we create benchmarks across 6 speech tasks".
- https://huggingface.co/datasets/ai4bharat/Kathbath recorded 2026-09-24
"Models trained or fine-tuned on ai4bharat/Kathbath" lists nvidia/magpie_tts_multilingual_357m, shunyalabs/zero-stt-hinglish.
Verified 2026-09-24