AI Potluck
Back to Gap Map Model components / Language-specific datasets

naab

Speech and Language Processing Lab, Sharif University of Technology
open / Overall score: 2.7

naab is a cleaned Farsi text corpus for pretraining language models, with about 130 GB, 250 million paragraphs, and 15 billion words. It merges existing Persian corpora, OSCAR-fa, the formerly private AGP corpus, Telegram channels, and LSCP, and a raw version (naab-raw) and the preprocessing code are published beside it. Sharif University of Technology's Speech and Language Processing Lab builds it.

Openness

5 medium confidence
5.0
license
mit(card metadata and API tag
access
public(both repositories ungated)
dataset_card
present(card lists each source corpus and the cleaning process)

The corpus, its raw version and the cleaning code are all downloadable, and the card names every source it merged. The license line on the card ends in a question mark, and the merged sources carry their own terms, so the grant is less certain than the metadata tag suggests.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the cleaned and raw repositories, most of them for the cleaned one. Shard fetches of a large corpus are counted individually, so the figure is not a count of training runs.

Capability

3 medium confidence
3.0

naab is the largest cleaned public Farsi pretraining corpus, compiled from web and messaging sources rather than newly collected and documented in a paper with its raw version and preprocessing code, and a Farsi text-to-text transformer and a Persian ALBERT are trained on it. At about fifteen billion words it is a sizable national corpus, yet small beside the trillions of tokens in English pretraining sets such as FineWeb.

  • https://arxiv.org/abs/2208.13486 recorded 2026-09-24

    Abstract: "the largest publicly available, cleaned, and ready-to-use Farsi textual corpus. naab consists of 130GB of data, comprising over 250 million paragraphs and 15 billion words".

  • https://huggingface.co/datasets/SLPL/naab recorded 2026-09-24

    Hub page "Models trained or fine-tuned on SLPL/naab" lists shekar-ai/albert-base-v2-persian-zwnj-naab-mlm and SLPL/t5-fa.

Verified 2026-09-24