Khmer ASR Cultural Dataset
Digital Divide DataKhmer ASR Cultural Dataset is a speech corpus of about 728 hours of Khmer read by native speakers, paired with transcripts, for training speech recognition and related models. Twelve speakers from Cambodia (five women, seven men) recorded sentences written around 61 cultural topics and 1,749 subtopics, averaging eight seconds per clip, with speaker gender, age group and origin city attached. Digital Divide Data built and published it.
Openness
5 high confidence- license
- cc-by-sa-4.0
- access
- public(ungated on Hugging Face
- dataset_card
- present(speakers, topics, format and sample rows)
The audio and transcripts download without a gate under a share-alike license that asks users to credit Digital Divide Data. The card describes speakers, topics and file layout, though no paper accompanies it.
- https://huggingface.co/api/datasets/Digital-Divide-Data/khmer-speech-dataset recorded 2026-09-24
gated: false; license:cc-by-sa-4.0
- https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset/raw/main/README.md recorded 2026-09-24
"license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0). Please attribute Digital Divide Data"
Adoption
2 high confidenceAdoption is Hugging Face downloads of the repository. Copies taken from the Mozilla Data Collective listing are not included.
- https://huggingface.co/api/datasets/Digital-Divide-Data/khmer-speech-dataset recorded 2026-09-24
3787 downloads in the trailing 30 days for Digital-Divide-Data/khmer-speech-dataset
Capability
2 medium confidenceThe corpus dwarfs OpenSLR's earlier Khmer speech-text set of under four hours, pairing prompted readings with curated transcripts, and community Khmer recognizers built on Whisper and Qwen are trained on it. At about seven hundred hours from twelve speakers reading prepared sentences it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices draws spontaneous speech from tens of thousands of speakers.
- https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset recorded 2026-09-24
Number of rows: 450,396; Models trained or fine-tuned: seanghay/Qwen3-ASR-0.6B-Khmer, sengtha/whisper-base-khmer
- https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset/raw/main/README.md recorded 2026-09-24
"727.94 hours of manually curated speech-text pairs by native speakers"; "there was only one Khmer speech-text dataset: OpenSLR 42 with 3.97 hours"
Verified 2026-09-24