HeDC4
HeNLPHeDC4 (Hebrew Deduplicated and Cleaned Common Crawl Corpus) is a Hebrew pretraining corpus built by cleaning and approximately deduplicating Common Crawl text. Released as a 47.6 GB file of 910,000 documents, it was introduced to train the HeRo and LongHeRo Hebrew encoder models. HeNLP publishes it, with Vitaly Shalumov and Harel Haskey as authors.
Openness
2 high confidence- license
- not-clearly-stated-on-card(no license in the card, the API tags, or any repository file (the repo holds only HeDC4.csv and README.md))
- access
- public(ungated on the Hub)
- dataset_card
- partial(two-sentence summary and a citation
The corpus downloads without a sign-up, but no license is stated anywhere, so the terms of reuse are unknown. The card is a two-line summary, and how the crawl was cleaned is explained only in the HeRo paper.
- https://huggingface.co/api/datasets/HeNLP/HeDC4 recorded 2026-09-24
API JSON: "gated": false; no license tag; files .gitattributes, HeDC4.csv, README.md.
- https://huggingface.co/datasets/HeNLP/HeDC4/raw/main/README.md recorded 2026-09-24
Full card: "A Hebrew Deduplicated and Cleaned Common Crawl Corpus. A thoroughly cleaned and approximately deduplicated dataset for unsupervised learning." No license field.
Adoption
1 high confidenceHugging Face downloads of the single repository. The figure cannot show how many of the Hebrew models listed as trained on it drew from this copy.
- https://huggingface.co/api/datasets/HeNLP/HeDC4 recorded 2026-09-24
54 downloads in the trailing 30 days for HeNLP/HeDC4
Capability
3 medium confidenceThe Hebrew Deduplicated and Cleaned Common Crawl Corpus was the largest Hebrew pretraining set when its paper introduced it, and HeRo, LongHeRo and a community HebrewGPT model are trained on it, but it draws on Common Crawl alone and its card is two sentences with no stated license. At roughly ten billion tokens, estimated from its size on disk, it is close in size to the Farsi corpus naab, which documents far more, and far below the trillions in English sets such as FineWeb.
- https://arxiv.org/abs/2304.11077 recorded 2026-09-24
Abstract: "providing it with the largest so far pre-train dataset HeDC4, a state-of-the-art pre-trained language model HeRo ... and an efficient transformer LongHeRo".
- https://huggingface.co/datasets/HeNLP/HeDC4 recorded 2026-09-24
Hub page: "Number of rows: 910,000"; "Total file size: 47.6 GB"; models listed include Slasky/HebrewGPT-1B, HeNLP/HeRo, HeNLP/LongHeRo.
Verified 2026-09-24