AI Potluck
Model components / Dataset Processing Tools

TxT360 (pipeline)

LLM360

Construction pipeline for the TxT360 open pretraining corpus: global deduplication and source processing across 99 Common Crawl snapshots and 14 curated sources. Distinct from the released TxT360 dataset.

LLM360 TxT360 pipeline code (distinct from the dataset on HF); reference/replication drop, last push 2024-12. README asserts Apache-2.0 but the repo ships NO LICENSE file (GitHub reports null) -- scored open_source on the asserted license + LLM360's open track record, with the missing file flagged.

Openness

2 low confidence
2.0
license
Apache-2.0 asserted in README only, no LICENSE file in repo (GitHub reports null)
source
public(LLM360/TxT360)

The README badge asserts Apache-2.0 but the repo ships no LICENSE file, so there is no actual redistribution grant. Re-checked 2026-07-30: still no LICENSE file, and the GitHub license endpoint returns 404 while the repo API reports a null license. Scored on that rather than on the badge. Score corrected from 3 to 2 on 2026-07-30. The software ladder has rungs at 1, 2, 4 and 5 and none at 3, so 3/source_available was not a pair any rule could produce; the ladder scores "you can read it and not run it freely" at 2.

Adoption

1 medium confidence
1.0

~25 GitHub stars on the pipeline repo; the value is the 5.7T-token corpus it builds, not the tool's own adoption.

Capability

3 medium confidence
3.0

Replication-oriented construction scripts rather than an evolving tool.

Unchanged since 2026-07-30 (last edited, not re-checked)