Institute of Smart Systems and Artificial Intelligence, Nazarbayev University
lab · KazakhstanScores
1 product on the map — 1 open.
Openness
4 medium confidence- license
- cc-by-4.0(ISSAI corpus page
- access
- public(ungated on the Hub, also downloadable from the ISSAI site)
- dataset_card
- partial(short card with summary and paper link
The audio and transcripts download freely for research and commercial use, with an attribution line the publisher asks commercial products to carry. The Hub card is a short summary, so the details of collection come from the paper and ISSAI's page.
- https://huggingface.co/api/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
API JSON: "gated": false, "private": false.
- https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2/raw/main/README.md recorded 2026-09-24
Card metadata "license: mit"; summary of sources and "around 1.2k hours of high-quality transcribed data comprising over 600k utterances".
- https://issai.nu.edu.kz/kz-speech-corpus/ recorded 2026-09-24
"the KSC2 dataset is freely available to both academic researchers and industry practitioners ... available under a Creative Commons Attribution 4.0 International License."
- https://raw.githubusercontent.com/IS2AI/Kazakh_ASR/main/LICENSE.md recorded 2026-09-24
Creative Commons "Attribution 4.0 International" license text.
Adoption
2 high confidenceHugging Face downloads of the Hub repository, which holds the corpus as split archive parts, so one full download counts several times. Downloads from the ISSAI site are not counted.
- https://huggingface.co/api/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
1866 downloads in the trailing 30 days for issai/Kazakh_Speech_Corpus_2
Capability
3 medium confidenceKSC gathers earlier Kazakh corpora with broadcast, Senate and podcast audio into what its publisher calls the first large open Kazakh speech corpus, documented in an Interspeech paper, and several community Whisper fine-tunes for Kazakh are trained on it. At about twelve hundred hours it is small beside English speech corpora of hundreds of thousands of hours, and it serves only recognition in one language, where IndicVoices spans twenty-two.
- https://huggingface.co/datasets/issai/Kazakh_Speech_Corpus_2 recorded 2026-09-24
Hub page "Models trained or fine-tuned on" lists shyngys879/kazakh-whisper-large-v3-turbo and abilmansplus/whisper-turbo-ksc2.
- https://issai.nu.edu.kz/kz-speech-corpus/ recorded 2026-09-24
"Kazakh Speech Corpus 2 (KSC2) is the first industrial-scale open-source Kazakh speech corpus ... contains utterances with the Kazakh-Russian code-switching".
- https://raw.githubusercontent.com/IS2AI/Kazakh_ASR/main/README.md recorded 2026-09-24
"This repository provides the recipe for the paper KSC2"; "Our code builds upon ESPnet".