AI Potluck
Back to Gap Map Model components / Language-specific datasets

TituLM Bangla corpus

Hishab
open / Overall score: 2.7

The TituLM Bangla corpus is a Bangla pretraining text corpus of about 20 billion tokens in 31 million documents. Most of it is Common Crawl text extracted with Trafilatura and filtered on 20 quality signals, and the rest is English news machine-translated into Bangla plus romanized Bangla, both produced with a fine-tuned NLLB model. Hishab released it alongside its TituLLMs Bangla language models.

Openness

5 medium confidence
5.0
license
cc-by-4.0(Hugging Face card metadata
access
public(Hugging Face gated: false)
dataset_card
present(card describes extraction, filtering thresholds, translation and per-category counts)

The corpus downloads without a gate and its card metadata sets CC BY 4.0. The card explains how each part was extracted, filtered and translated, but says nothing about the rights in the crawled pages.

Adoption

2 high confidence
2.0

Hugging Face downloads of the single corpus repository. A download does not show whether the data was used to train a released model.

Capability

3 medium confidence
3.0

The TituLM corpus is a large, documented Bangla pretraining set, but its sources do not say the TituLLMs models were trained on this release, and only a small third-party model lists it. About a quarter of its tokens are machine-translated or romanized, and Sangraha, a documented corpus across 22 Indic languages, is far larger.

Verified 2026-09-24