AI Potluck
Back to Gap Map Organization

Speech and Language Processing Lab, Sharif University of Technology

lab · Iran

Scores

1 product on the map — 1 open.

naab

Openness

5 medium confidence
5.0
license
mit(card metadata and API tag
access
public(both repositories ungated)
dataset_card
present(card lists each source corpus and the cleaning process)

The corpus, its raw version and the cleaning code are all downloadable, and the card names every source it merged. The license line on the card ends in a question mark, and the merged sources carry their own terms, so the grant is less certain than the metadata tag suggests.

Adoption

2 high confidence
2.0

Hugging Face downloads summed over the cleaned and raw repositories, most of them for the cleaned one. Shard fetches of a large corpus are counted individually, so the figure is not a count of training runs.

Capability

3 medium confidence
3.0

naab is the largest cleaned public Farsi pretraining corpus, compiled from web and messaging sources rather than newly collected and documented in a paper with its raw version and preprocessing code, and a Farsi text-to-text transformer and a Persian ALBERT are trained on it. At about fifteen billion words it is a sizable national corpus, yet small beside the trillions of tokens in English pretraining sets such as FineWeb.

  • https://arxiv.org/abs/2208.13486 recorded 2026-09-24

    Abstract: "the largest publicly available, cleaned, and ready-to-use Farsi textual corpus. naab consists of 130GB of data, comprising over 250 million paragraphs and 15 billion words".

  • https://huggingface.co/datasets/SLPL/naab recorded 2026-09-24

    Hub page "Models trained or fine-tuned on SLPL/naab" lists shekar-ai/albert-base-v2-persian-zwnj-naab-mlm and SLPL/t5-fa.