AI Potluck
Back to Gap Map Model components / Language-specific datasets

Kazakh Speech Corpus 2

Institute of Smart Systems and Artificial Intelligence, Nazarbayev University
open / Overall score: 2.7

Kazakh Speech Corpus 2 (KSC2) provides about 1,200 hours of transcribed Kazakh speech in over 600,000 utterances for speech recognition. It merges the earlier Kazakh Speech Corpus and Kazakh Text-To-Speech 2 with recordings from TV, radio, the Senate, and podcasts, and includes Kazakh-Russian code-switching. The Institute of Smart Systems and Artificial Intelligence (ISSAI) at Nazarbayev University releases it with an ESPnet training recipe.

Openness

4 medium confidence
4.0
license
cc-by-4.0(ISSAI corpus page
access
public(ungated on the Hub, also downloadable from the ISSAI site)
dataset_card
partial(short card with summary and paper link

The audio and transcripts download freely for research and commercial use, with an attribution line the publisher asks commercial products to carry. The Hub card is a short summary, so the details of collection come from the paper and ISSAI's page.

Adoption

2 high confidence
2.0

Hugging Face downloads of the Hub repository, which holds the corpus as split archive parts, so one full download counts several times. Downloads from the ISSAI site are not counted.

Capability

3 medium confidence
3.0

KSC gathers earlier Kazakh corpora with broadcast, Senate and podcast audio into what its publisher calls the first large open Kazakh speech corpus, documented in an Interspeech paper, and several community Whisper fine-tunes for Kazakh are trained on it. At about twelve hundred hours it is small beside English speech corpora of hundreds of thousands of hours, and it serves only recognition in one language, where IndicVoices spans twenty-two.

Verified 2026-09-24