LaoBench
Beijing Academy of Artificial Intelligence (BAAI)LaoBench is a benchmark for testing language models in Lao, with more than 17,000 expert-curated items across three areas: culturally grounded knowledge, K-12 curriculum questions, and translation among Lao, Chinese and English. Experts wrote the items, and an agent-assisted pipeline verified them. It comes in an open subset and a held-out subset scored through a controlled service. BAAI publishes it.
The Hugging Face card carries only metadata; the description and scale come from the LaoBench paper.
Openness
2 medium confidence- license
- apache-2.0(card metadata
- access
- public(open subset ungated
- dataset_card
- partial(README has YAML metadata only, no prose)
- answers
- held-out(paper: held-out portion enables black-box evaluation via a controlled service
The open subset downloads without a gate under Apache-2.0, while part of the benchmark is kept private and scored by BAAI's service. The repository card has no description, so how the data was made is documented only in the paper.
- https://arxiv.org/abs/2511.11334 recorded 2026-09-24
"It includes open-source and held-out subsets, where the held-out portion enables secure black-box evaluation via a controlled service"
- https://huggingface.co/api/datasets/BAAI/LaoBench recorded 2026-09-24
gated: false; files: Lao-7k/test-00000-of-00001.parquet
- https://huggingface.co/datasets/BAAI/LaoBench/raw/main/README.md recorded 2026-09-24
YAML only: license: apache-2.0; language: lo; pretty_name: LaoBench; size_categories 1K<n<10K
Adoption
1 high confidenceAdoption is Hugging Face downloads of the open subset. Evaluations on the held-out portion go through BAAI's service and are not counted.
- https://huggingface.co/api/datasets/BAAI/LaoBench recorded 2026-09-24
154 downloads in the trailing 30 days for BAAI/LaoBench
Capability
3 medium confidenceLaoBench is documented in a paper and offers expert-written Lao items in knowledge, schooling and translation. No outside leaderboard or model report using it was found, less than half of it is openly downloadable, and SEA-HELM tests Lao among eleven Southeast Asian languages on its own leaderboard.
- https://arxiv.org/abs/2511.11334 recorded 2026-09-24
"LaoBench contains 17,000+ expert-curated samples across three dimensions"
- https://huggingface.co/datasets/BAAI/LaoBench recorded 2026-09-24
Split (1) test · 7k rows
Verified 2026-09-24