ASLP Lab
lab · ChinaScores
1 product on the map — 1 open.
Openness
5 medium confidence- license
- cc-by-4.0(card Licensing section
- access
- public
- dataset_card
- present(card describes sources, structure, annotation fields and quality checks)
The card releases the corpus under an attribution-only license with no gate. The audio itself is taken from YouTube and state television archives, and the card does not say how rights to it were cleared.
- https://huggingface.co/api/datasets/ASLP-lab/UrduSpeech recorded 2026-09-24
gated: false; no license in cardData
- https://huggingface.co/datasets/ASLP-lab/UrduSpeech/raw/main/README.md recorded 2026-09-24
"This dataset is released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license."; "Sources: YouTube and Pakistan Television (PTV) archival content"
Adoption
2 high confidenceHugging Face downloads of the corpus repository. The repository holds tens of thousands of files, so one download can mean many file fetches.
- https://huggingface.co/api/datasets/ASLP-lab/UrduSpeech recorded 2026-09-24
2889 downloads in the trailing 30 days for ASLP-lab/UrduSpeech
Capability
2 medium confidenceUrduSpeech gives Urdu, code-switched and accented speech from YouTube and television archives, documented in a paper, with a benchmark corrected by native speakers and the rest transcribed by an LLM pipeline, though no model outside its paper names it. At under two hundred hours it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices has people transcribe far more speech.
- https://arxiv.org/abs/2605.17846 recorded 2026-09-24
"a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata"; "LLM-driven pipeline"; "9-hour US-Benchmark set, manually corrected by native annotators"