AI Potluck
Back to Gap Map Model components / Language-specific datasets

SEA-PILE

AI Singapore
open / Overall score: 3.4

SEA-PILE is AI Singapore's pretraining corpus for Southeast Asian languages. Its v2 release holds about 120 billion tokens of web text in Vietnamese, Indonesian, Tamil, Malay, Thai, Tagalog, Khmer, Lao and Burmese, extracted from 24 CommonCrawl snapshots, deduplicated per snapshot and filtered with heuristics and perplexity scoring developed with native speakers. The earlier v1 release is the cleaned mC4 portion of the data behind the first SEA-LION models.

The card for v2 says the public repository holds only the initial open release; the larger pool cited for later SEA-LION models is internal.

Openness

5 high confidence
5.0
license
odc-by(SEA-PILE-v2
access
public(both repos ungated)
dataset_card
present(per-language token table, pipeline and limitations for v2

Both releases download without a gate under permissive licenses, and the v2 card describes the crawl snapshots and filtering steps. The expanded pool used for later SEA-LION models is not published, so the open release is a subset of what the models saw.

Adoption

2 high confidence
2.0

Adoption is Hugging Face downloads summed over the v1 and v2 repositories, most of them for v2. A pretraining corpus is typically fetched once and reused, so repeat downloads understate how widely a copy is used.

Capability

4 medium confidence
4.0

SEA-PILE offers filtered web text across nine Southeast Asian languages, more than any other open corpus for the region here, and the SEA-LION models were trained on this line of data, with Vietnamese and Indonesian holding most of the tokens and Khmer, Lao and Burmese under a billion each. At about a hundred and twenty billion tokens it is among the larger non-English pretraining sets, though English web corpora such as FineWeb run to trillions.

Verified 2026-09-24