AI Potluck
Back to Gap Map Organization

VTS NLP (Viettel)

company · Vietnam

Scores

1 product on the map — 1 closed.

Vietnamese Curated Dataset

Openness

2 medium confidence
2.0
license
not-clearly-stated-on-card(no license in card metadata, card text or API tags
access
public(ungated)
dataset_card
present(sources, curation steps and statistics)

The corpus downloads without a gate and the card lists its sources and cleaning steps, but no license is given for the release. Its component corpora each carry their own terms, which a user would have to trace.

Adoption

2 high confidence
2.0

Hugging Face downloads of the repository, which is the only place the corpus is published.

Capability

3 medium confidence
3.0

The corpus re-filters existing web, OSCAR, Wikipedia and news text into a cleaner Vietnamese pretraining set, documented by its card and an NVIDIA technical post, and Viettel Solutions trained its Llama-based Vietnamese model on it. At roughly fifteen billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and it adds no new text, where SEA-PILE draws fresh crawl text for Vietnamese and eight other languages.