AI Potluck
Model components / Dataset Processing Tools

TxT360 (pipeline)

LLM360

Construction pipeline for the TxT360 open pretraining corpus: global deduplication and source processing across 99 Common Crawl snapshots and 14 curated sources. Distinct from the released TxT360 dataset.

LLM360's TxT360 pipeline code, distinct from the dataset on the Hub; a reference and replication drop, last pushed December 2024. The README badge asserts a license the repository ships no file for, and the GitHub API reports none, which is why openness abstains rather than scoring it. Verified 2026-08-13 via the GitHub README and the repository API.

Openness

2 low confidence
2.0
license
none-declared(the README badge asserts Apache-2.0 but the repo ships no LICENSE file and the GitHub license endpoint 404s)
source
public(LLM360/TxT360)

The README badge asserts Apache-2.0 but the repo ships no LICENSE file, so there is no actual redistribution grant: the GitHub license endpoint returns 404, the repo API reports a null license, and the badge links to a LICENSE in a different repository. The score follows that rather than the badge. Code you can read but cannot freely run scores 2, source-available, on the software scale, which has rungs at 1, 2, 4 and 5 and none at 3. The license is recorded as none-declared rather than as a sentence, so that the tier lookup reads a license name; no tier covers a grant nobody made.

Adoption

1 medium confidence
1.0

25 GitHub stars on LLM360/TxT360, read from the repository API, which bands at <1K stars, level 1 on the stars fallback scale. That scale caps at level 3, because a star is not a use. The value here is the 5.7T-token corpus the scripts build rather than the tool's own adoption.

Capability

3 medium confidence
3.0

Replication-oriented construction scripts rather than an evolving tool.

Verified 2026-08-13