AI Potluck
Back to Gap Map Model components / Language-specific datasets

ArabicWeb24

LightOn
gated / Overall score: 2.4

ArabicWeb24 is a pretraining corpus of Arabic web text from a custom crawl, processed with the datatrove library; the card gives more than 28 billion tokens for the fully cleaned and deduplicated version. A second version skips the sentence deduplication step, and the processing code and small ablation models trained on both versions are released alongside. LightOn built it with INSAT.

The card headline says more than 39 billion tokens while its body says more than 28 billion for the cleaned version; the description uses the body figure.

Openness

3 high confidence
3.0
license
odc-by(card metadata and API tag)
access
auto(automatic approval after agreeing to share contact information)
dataset_card
present(card describes both versions, fields and the processing library)

The corpus carries an attribution-only data license, and access is granted automatically once a user shares contact details on the Hub. The detailed processing steps are in blog posts rather than on the card or in a paper.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single repository, which holds both versions. Each shard fetched from a large corpus counts separately, so the figure is not a count of training runs.

Capability

3 medium confidence
3.0

ArabicWeb24 is a documented Arabic web corpus released with its processing code. No model outside LightOn's own ablations is known to be trained on it, unlike 101 Billion Arabic Words, which ModernAraBERT uses, or naab for Farsi, a larger documented corpus merged from many sources.

  • https://huggingface.co/datasets/lightonai/ArabicWeb24 recorded 2026-09-24

    "more than 28 billion tokens of cleaned and deduplicated Arabic web data from a customized crawl ... processed using the large scale data processing library datatrove"; "Total file size: 471 GB"; models listed: ArabicWeb24-ablation-model-v1, -v5.

Verified 2026-09-24