Hishab
company · BangladeshScores
1 product on the map — 1 open.
Openness
5 medium confidence- license
- cc-by-4.0(Hugging Face card metadata
- access
- public(Hugging Face gated: false)
- dataset_card
- present(card describes extraction, filtering thresholds, translation and per-category counts)
The corpus downloads without a gate and its card metadata sets CC BY 4.0. The card explains how each part was extracted, filtered and translated, but says nothing about the rights in the crawled pages.
- https://huggingface.co/api/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
"gated": false; cardData license "cc-by-4.0".
- https://huggingface.co/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
Card: Common Crawl Filtered 24.3M documents, 14.80B tokens; Translated 1.47B; Romanized 3.87B; total 31.21M documents, 20.14B tokens; Gopher-style quality thresholds.
Adoption
2 high confidenceHugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.
- https://huggingface.co/api/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
1660 downloads in the trailing 30 days for hishab/titulm-bangla-corpus
Capability
3 medium confidenceThe TituLM corpus is a large, documented Bangla pretraining set, but its sources do not say the TituLLMs models were trained on this release, and only a small third-party model lists it. About a quarter of its tokens are machine-translated or romanized, and Sangraha, a documented corpus across 22 Indic languages, is far larger.
- https://arxiv.org/abs/2502.11187 recorded 2026-09-24
Abstract: "To train TituLLMs, we collected a pretraining dataset of approximately ~37 billion tokens."
- https://huggingface.co/datasets/hishab/titulm-bangla-corpus recorded 2026-09-24
"This dataset is associated with the paper TituLLMs"; "Models trained or fine-tuned on" lists spitfire4794/Alo-70m-Base and Alo-50M-Base-Retied.