AI Potluck
Back to Gap Map Model components / Language-specific datasets

HeDC4

HeNLP
restricted / Overall score: 2.4

HeDC4 (Hebrew Deduplicated and Cleaned Common Crawl Corpus) is a Hebrew pretraining corpus built by cleaning and approximately deduplicating Common Crawl text. Released as a 47.6 GB file of 910,000 documents, it was introduced to train the HeRo and LongHeRo Hebrew encoder models. HeNLP publishes it, with Vitaly Shalumov and Harel Haskey as authors.

Openness

2 high confidence
2.0
license
not-clearly-stated-on-card(no license in the card, the API tags, or any repository file (the repo holds only HeDC4.csv and README.md))
access
public(ungated on the Hub)
dataset_card
partial(two-sentence summary and a citation

The corpus downloads without a sign-up, but no license is stated anywhere, so the terms of reuse are unknown. The card is a two-line summary, and how the crawl was cleaned is explained only in the HeRo paper.

Adoption

1 high confidence
1.0

Hugging Face downloads of the single repository. The figure cannot show how many of the Hebrew models listed as trained on it drew from this copy.

Capability

3 medium confidence
3.0

The Hebrew Deduplicated and Cleaned Common Crawl Corpus was the largest Hebrew pretraining set when its paper introduced it, and HeRo, LongHeRo and a community HebrewGPT model are trained on it, but it draws on Common Crawl alone and its card is two sentences with no stated license. At roughly ten billion tokens, estimated from its size on disk, it is close in size to the Farsi corpus naab, which documents far more, and far below the trillions in English sets such as FineWeb.

  • https://arxiv.org/abs/2304.11077 recorded 2026-09-24

    Abstract: "providing it with the largest so far pre-train dataset HeDC4, a state-of-the-art pre-trained language model HeRo ... and an efficient transformer LongHeRo".

  • https://huggingface.co/datasets/HeNLP/HeDC4 recorded 2026-09-24

    Hub page: "Number of rows: 910,000"; "Total file size: 47.6 GB"; models listed include Slasky/HebrewGPT-1B, HeNLP/HeRo, HeNLP/LongHeRo.

Verified 2026-09-24