AI Potluck
Back to Gap Map Model components / Language-specific datasets

101 Billion Arabic Words

ClusterlabAi
open / Overall score: 3.4

The 101 Billion Arabic Words Dataset is a pretraining corpus of Arabic web text, mixing Modern Standard Arabic and dialects, for training and fine-tuning language models. Arabic content was extracted from Common Crawl WET files, then cleaned and deduplicated, yielding about 101 billion words in 33 million documents. It is curated by the Clusterlab team.

Openness

5 high confidence
5.0
license
apache-2.0(card metadata and the card body ("License: Apache 2.0"))
access
public(ungated on the Hugging Face Hub)
dataset_card
present(card and paper describe the Common Crawl extraction and cleaning)

The full corpus downloads from the Hub without a sign-up, and its card and paper describe how the web text was extracted and cleaned. The text itself is crawled web content, so the terms of the original pages still sit behind the permissive grant on the compilation.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single Hub repository. Downloads of a corpus this size mostly count partial fetches of its shards and cannot show how many training runs used it.

Capability

4 medium confidence
4.0

The corpus gives Arabic a large cleaned and deduplicated Common Crawl collection, documented in a paper whose authors call it the largest Arabic dataset available, and ModernAraBERT is trained on it. At about a hundred billion words it is among the largest single-language corpora outside English and Chinese, though still short of the trillions of tokens in FineWeb or DCLM, and like Sangraha it is a big documented corpus with little outside use recorded.

Verified 2026-09-24