AI Potluck
Back to Gap Map Model components / Language-specific datasets

Khmer ASR Cultural Dataset

Digital Divide Data
open / Overall score: 2.0

Khmer ASR Cultural Dataset is a speech corpus of about 728 hours of Khmer read by native speakers, paired with transcripts, for training speech recognition and related models. Twelve speakers from Cambodia (five women, seven men) recorded sentences written around 61 cultural topics and 1,749 subtopics, averaging eight seconds per clip, with speaker gender, age group and origin city attached. Digital Divide Data built and published it.

Openness

5 high confidence
5.0
license
cc-by-sa-4.0
access
public(ungated on Hugging Face
dataset_card
present(speakers, topics, format and sample rows)

The audio and transcripts download without a gate under a share-alike license that asks users to credit Digital Divide Data. The card describes speakers, topics and file layout, though no paper accompanies it.

Adoption

2 high confidence
2.0

Adoption is Hugging Face downloads of the repository. Copies taken from the Mozilla Data Collective listing are not included.

Capability

2 medium confidence
2.0

The corpus dwarfs OpenSLR's earlier Khmer speech-text set of under four hours, pairing prompted readings with curated transcripts, and community Khmer recognizers built on Whisper and Qwen are trained on it. At about seven hundred hours from twelve speakers reading prepared sentences it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices draws spontaneous speech from tens of thousands of speakers.

Verified 2026-09-24