TituLM Bangla corpus
HishabThe TituLM Bangla corpus is a Bangla pretraining text corpus of about 20 billion tokens in 31 million documents. Most of it is Common Crawl text extracted with Trafilatura and filtered on 20 quality signals, and the rest is English news machine-translated into Bangla plus romanized Bangla, both produced with a fine-tuned NLLB model. Hishab released it alongside its TituLLMs Bangla language models.
Openness
5 medium confidence- license
- cc-by-4.0(Hugging Face card metadata
- access
- public(Hugging Face gated: false)
- dataset_card
- present(card describes extraction, filtering thresholds, translation and per-category counts)
The corpus downloads without a gate and its card metadata sets CC BY 4.0. The card explains how each part was extracted, filtered and translated, but says nothing about the rights in the crawled pages.
- https://huggingface.co/api/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
Card: Common Crawl Filtered 24.3M documents, 14.80B tokens; Translated 1.47B; Romanized 3.87B; total 31.21M documents, 20.14B tokens; Gopher-style quality thresholds.
Adoption
2 high confidenceHugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.
- https://huggingface.co/api/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
1660 downloads in the trailing 30 days for hishab/titulm-bangla-corpus
Capability
3 medium confidenceThe TituLM corpus is a large, documented Bangla pretraining set, but its sources do not say the TituLLMs models were trained on this release, and only a small third-party model lists it. About a quarter of its tokens are machine-translated or romanized, and Sangraha, a documented corpus across 22 Indic languages, is far larger.
- https://arxiv.org/abs/2502.11187 recorded 2026-09-24
Abstract: "To train TituLLMs, we collected a pretraining dataset of approximately ~37 billion tokens."
- https://huggingface.co/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
"This dataset is associated with the paper TituLLMs"; "Models trained or fine-tuned on" lists spitfire4794/Alo-70m-Base and Alo-50M-Base-Retied.
Verified 2026-09-24