SEA-PILE
AI SingaporeSEA-PILE is AI Singapore's pretraining corpus for Southeast Asian languages. Its v2 release holds about 120 billion tokens of web text in Vietnamese, Indonesian, Tamil, Malay, Thai, Tagalog, Khmer, Lao and Burmese, extracted from 24 CommonCrawl snapshots, deduplicated per snapshot and filtered with heuristics and perplexity scoring developed with native speakers. The earlier v1 release is the cleaned mC4 portion of the data behind the first SEA-LION models.
The card for v2 says the public repository holds only the initial open release; the larger pool cited for later SEA-LION models is internal.
Openness
5 high confidence- license
- odc-by(SEA-PILE-v2
- access
- public(both repos ungated)
- dataset_card
- present(per-language token table, pipeline and limitations for v2
Both releases download without a gate under permissive licenses, and the v2 card describes the crawl snapshots and filtering steps. The expanded pool used for later SEA-LION models is not published, so the open release is a subset of what the models saw.
- https://huggingface.co/api/datasets/aisingapore/SEA-PILE-v1 recorded 2026-09-24
gated: false; license:mit
- https://huggingface.co/api/datasets/aisingapore/SEA-PILE-v2 recorded 2026-09-24
gated: false; license:odc-by
- https://huggingface.co/datasets/aisingapore/SEA-PILE-v1/raw/main/README.md recorded 2026-09-24
license: mit; "This repository contains the cleaned mC4 portion of the SEA-LION-Pile"
- https://huggingface.co/datasets/aisingapore/SEA-PILE-v2/raw/main/README.md recorded 2026-09-24
"This dataset is made available under ODC-By 1.0 license; users should also abide by the CommonCrawl ToU"; "This public repository contains the initial 120B/122B token open-source release"
Adoption
2 high confidenceAdoption is Hugging Face downloads summed over the v1 and v2 repositories, most of them for v2. A pretraining corpus is typically fetched once and reused, so repeat downloads understate how widely a copy is used.
- https://huggingface.co/api/datasets/aisingapore/SEA-PILE-v1 recorded 2026-09-24
697 downloads in the trailing 30 days for aisingapore/SEA-PILE-v1
- https://huggingface.co/api/datasets/aisingapore/SEA-PILE-v2 recorded 2026-09-24
2766 downloads in the trailing 30 days for aisingapore/SEA-PILE-v2
Capability
4 medium confidenceSEA-PILE offers filtered web text across nine Southeast Asian languages, more than any other open corpus for the region here, and the SEA-LION models were trained on this line of data, with Vietnamese and Indonesian holding most of the tokens and Khmer, Lao and Burmese under a billion each. At about a hundred and twenty billion tokens it is among the larger non-English pretraining sets, though English web corpora such as FineWeb run to trillions.
- https://huggingface.co/datasets/aisingapore/SEA-PILE-v1/raw/main/README.md recorded 2026-09-24
"SEA-LION-Pile is the pretraining data set for SEA-LION"
- https://huggingface.co/datasets/aisingapore/SEA-PILE-v2 recorded 2026-09-24
Number of rows: 186,809,986; Total file size: 334 GB
- https://huggingface.co/datasets/aisingapore/SEA-PILE-v2/raw/main/README.md recorded 2026-09-24
Token table: Vietnamese 51.4B, Indonesian 41.9B ... Khmer 0.6B, Lao 0.6B, Burmese 0.2B; "extracting text from 24 CommonCrawl snapshots"
Verified 2026-09-24