Kazakh Speech Corpus 2
Institute of Smart Systems and Artificial Intelligence, Nazarbayev UniversityKazakh Speech Corpus 2 (KSC2) provides about 1,200 hours of transcribed Kazakh speech in over 600,000 utterances for speech recognition. It merges the earlier Kazakh Speech Corpus and Kazakh Text-To-Speech 2 with recordings from TV, radio, the Senate, and podcasts, and includes Kazakh-Russian code-switching. The Institute of Smart Systems and Artificial Intelligence (ISSAI) at Nazarbayev University releases it with an ESPnet training recipe.
Openness
4 medium confidence- license
- cc-by-4.0(ISSAI corpus page
- access
- public(ungated on the Hub, also downloadable from the ISSAI site)
- dataset_card
- partial(short card with summary and paper link
The audio and transcripts download freely for research and commercial use, with an attribution line the publisher asks commercial products to carry. The Hub card is a short summary, so the details of collection come from the paper and ISSAI's page.
- https://huggingface.co/api/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
API JSON: "gated": false, "private": false.
- https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2/raw/main/README.md recorded 2026-09-24
Card metadata "license: mit"; summary of sources and "around 1.2k hours of high-quality transcribed data comprising over 600k utterances".
- https://issai.nu.edu.kz/kz-speech-corpus/ recorded 2026-09-24
"the KSC2 dataset is freely available to both academic researchers and industry practitioners ... available under a Creative Commons Attribution 4.0 International License."
- https://raw.githubusercontent.com/IS2AI/Kazakh_ASR/main/LICENSE.md recorded 2026-09-24
Creative Commons "Attribution 4.0 International" license text.
Adoption
2 high confidenceHugging Face downloads of the Hub repository, which holds the corpus as split archive parts, so one full download counts several times. Downloads from the ISSAI site are not counted.
- https://huggingface.co/api/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
1866 downloads in the trailing 30 days for issai/Kazakh_Speech_Corpus_2
Capability
3 medium confidenceKSC gathers earlier Kazakh corpora with broadcast, Senate and podcast audio into what its publisher calls the first large open Kazakh speech corpus, documented in an Interspeech paper, and several community Whisper fine-tunes for Kazakh are trained on it. At about twelve hundred hours it is small beside English speech corpora of hundreds of thousands of hours, and it serves only recognition in one language, where IndicVoices spans twenty-two.
- https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
Hub page "Models trained or fine-tuned on" lists shyngys879/kazakh-whisper-large-v3-turbo and abilmansplus/whisper-turbo-ksc2.
- https://issai.nu.edu.kz/kz-speech-corpus/ recorded 2026-09-24
"Kazakh Speech Corpus 2 (KSC2) is the first industrial-scale open-source Kazakh speech corpus ... contains utterances with the Kazakh-Russian code-switching".
- https://raw.githubusercontent.com/IS2AI/Kazakh_ASR/main/README.md recorded 2026-09-24
"This repository provides the recipe for the paper KSC2"; "Our code builds upon ESPnet".
Verified 2026-09-24