VTS NLP (Viettel)
company · VietnamScores
1 product on the map — 1 closed.
Openness
2 medium confidence- license
- not-clearly-stated-on-card(no license in card metadata, card text or API tags
- access
- public(ungated)
- dataset_card
- present(sources, curation steps and statistics)
The corpus downloads without a gate and the card lists its sources and cleaning steps, but no license is given for the release. Its component corpora each carry their own terms, which a user would have to trace.
- https://huggingface.co/api/datasets/VTSNLP/vietnamese_curated_dataset recorded 2026-09-24
gated: false; no license tag; files are .gitattributes, README.md and data/ only
- https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset/raw/main/README.md recorded 2026-09-24
Card metadata has dataset_info and configs only, no license field; lists C4, OSCAR 23.01, Wikipedia and Binhvq news as sources
Adoption
2 high confidenceHugging Face downloads of the repository, which is the only place the corpus is published.
- https://huggingface.co/api/datasets/VTSNLP/vietnamese_curated_dataset recorded 2026-09-24
1204 downloads in the trailing 30 days for VTSNLP/vietnamese_curated_dataset
Capability
3 medium confidenceThe corpus re-filters existing web, OSCAR, Wikipedia and news text into a cleaner Vietnamese pretraining set, documented by its card and an NVIDIA technical post, and Viettel Solutions trained its Llama-based Vietnamese model on it. At roughly fifteen billion tokens, estimated from its size on disk, it is far below the trillions of tokens in English pretraining corpora, and it adds no new text, where SEA-PILE draws fresh crawl text for Vietnamese and eight other languages.
- https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset recorded 2026-09-24
Number of rows: 12,169,131; Models trained or fine-tuned: VTSNLP/Llama3-ViettelSolutions-8B
- https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset/raw/main/README.md recorded 2026-09-24
"This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator"; dataset_size: 65506190827