AI Potluck
Back to Gap Map Model components / Language-specific datasets

Vietnamese Curated Dataset

VTS NLP (Viettel)
restricted / Overall score: 2.7

Vietnamese Curated Dataset is a pretraining text corpus of about 12.2 million Vietnamese documents assembled from existing open sources: the Vietnamese subsets of C4 and OSCAR 23.01, Vietnamese Wikipedia and the Binhvq news corpus. Viettel Solutions cleaned it with NVIDIA NeMo Curator, applying Unicode normalization, exact deduplication, and heuristic and classifier-based quality filtering.

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(no license in card metadata, card text or API tags
access
public(ungated)
dataset_card
present(sources, curation steps and statistics)

The corpus downloads without a gate and the card lists its sources and cleaning steps, but no license is given for the release. Its component corpora each carry their own terms, which a user would have to trace.

Adoption

2 high confidence
2.0

Hugging Face downloads of the repository, which is the only place the corpus is published.

Capability

3 medium confidence
3.0

The corpus re-filters existing web, OSCAR, Wikipedia and news text into a cleaner Vietnamese pretraining set, documented by its card and an NVIDIA technical post, and Viettel Solutions trained its Llama-based Vietnamese model on it. At roughly fifteen billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and it adds no new text, where SEA-PILE draws fresh crawl text for Vietnamese and eight other languages.

Verified 2026-09-24