TxT360 (pipeline)
LLM360Construction pipeline for the TxT360 open pretraining corpus: global deduplication and source processing across 99 Common Crawl snapshots and 14 curated sources. Distinct from the released TxT360 dataset.
LLM360 TxT360 pipeline code (distinct from the dataset on HF); reference/replication drop, last push 2024-12. README asserts Apache-2.0 but the repo ships NO LICENSE file (GitHub reports null) -- scored open_source on the asserted license + LLM360's open track record, with the missing file flagged.
Openness
2 low confidence- license
- Apache-2.0 asserted in README only, no LICENSE file in repo (GitHub reports null)
- source
- public(LLM360/TxT360)
The README badge asserts Apache-2.0 but the repo ships no LICENSE file, so there is no actual redistribution grant. Re-checked 2026-07-30: still no LICENSE file, and the GitHub license endpoint returns 404 while the repo API reports a null license. Scored on that rather than on the badge. Score corrected from 3 to 2 on 2026-07-30. The software ladder has rungs at 1, 2, 4 and 5 and none at 3, so 3/source_available was not a pair any rule could produce; the ladder scores "you can read it and not run it freely" at 2.
- https://github.com/LLM360/TxT360 recorded 2026-07-21
README asserts Apache-2.0; no LICENSE file in repo (GitHub reports null license)
Adoption
1 medium confidence~25 GitHub stars on the pipeline repo; the value is the 5.7T-token corpus it builds, not the tool's own adoption.
- https://github.com/LLM360/TxT360 recorded 2026-07-21
25 stars
Capability
3 medium confidenceReplication-oriented construction scripts rather than an evolving tool.
- https://github.com/LLM360/TxT360 recorded 2026-07-21
txt360-pipeline repo feature set confirmed
Unchanged since 2026-07-30 (last edited, not re-checked)