AI Potluck
Back to Gap Map Organization

Digital Divide Data

foundation · United States

Scores

1 product on the map — 1 open.

Khmer ASR Cultural Dataset

Openness

5 high confidence
5.0
license
cc-by-sa-4.0
access
public(ungated on Hugging Face
dataset_card
present(speakers, topics, format and sample rows)

The audio and transcripts download without a gate under a share-alike license that asks users to credit Digital Divide Data. The card describes speakers, topics and file layout, though no paper accompanies it.

Adoption

2 high confidence
2.0

Adoption is Hugging Face downloads of the repository. Copies taken from the Mozilla Data Collective listing are not included.

Capability

2 medium confidence
2.0

The corpus dwarfs OpenSLR's earlier Khmer speech-text set of under four hours, pairing prompted readings with curated transcripts, and community Khmer recognizers built on Whisper and Qwen are trained on it. At about seven hundred hours from twelve speakers reading prepared sentences it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices draws spontaneous speech from tens of thousands of speakers.