AI Potluck
Back to Gap Map Model components / Language-specific datasets

UrduSpeech

ASLP Lab
open / Overall score: 2.0

UrduSpeech is an Urdu speech corpus of 156 hours in 71,792 utterances, split into standard Pakistani Urdu, Urdu-English code-switching and Pakistani-accented English. Audio comes from YouTube and Pakistan Television archives across 12 categories such as news, drama and poetry, then transcribed and tagged with 12 paralinguistic attributes through an LLM pipeline built on Gemini 2.5 Pro. ASLP-lab built it, with a separate 9-hour benchmark corrected by native speakers.

Openness

5 medium confidence
5.0
license
cc-by-4.0(card Licensing section
access
public
dataset_card
present(card describes sources, structure, annotation fields and quality checks)

The card releases the corpus under an attribution-only license with no gate. The audio itself is taken from YouTube and state television archives, and the card does not say how rights to it were cleared.

Adoption

2 high confidence
2.0

Hugging Face downloads of the corpus repository. The repository holds tens of thousands of files, so one download can mean many file fetches.

Capability

2 medium confidence
2.0

UrduSpeech gives Urdu, code-switched and accented speech from YouTube and television archives, documented in a paper, with a benchmark corrected by native speakers and the rest transcribed by an LLM pipeline, though no model outside its paper names it. At under two hundred hours it is small beside English speech corpora of hundreds of thousands of hours, and IndicVoices has people transcribe far more speech.

  • https://arxiv.org/abs/2605.17846 recorded 2026-09-24

    "a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata"; "LLM-driven pipeline"; "9-hour US-Benchmark set, manually corrected by native annotators"

Verified 2026-09-24