TxT360
LLM360TxT360 is LLM360's globally deduplicated pretraining corpus combining 99 Common Crawl snapshots with 14 curated high-quality sources (FreeLaw, PG-19, arXiv, Wikipedia, StackExchange, and more) under ODC-BY. It trains K2-V2 together with the TxT360-Midas mid-training and 3efforts SFT extensions.
Family bucket: TxT360 plus the Midas mid-training and 3efforts SFT extensions.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, documented pipeline on GitHub.
- https://huggingface.co/datasets/LLM360/TxT360 recorded 2026-07-04
ODC-BY license, ungated, dataset card
Adoption
2 high confidence4,218 monthly downloads on Hugging Face, graded on the training-corpus bands.
- https://huggingface.co/api/datasets/LLM360/TxT360 recorded 2026-07-04
downloads field = 4,218 (30-day)
Capability
4 medium confidenceA self-reported controlled ablation beats FineWeb (8x8B MoE, 1.5T tokens) but without published numeric scores; trains K2 / K2-V2 (LLM360, 2024-25).
- https://huggingface.co/datasets/LLM360/TxT360 recorded 2026-07-04
5.7T deduplicated tokens upsampled to 15T+, claimed to outperform FineWeb 15T; trains K2/K2-V2
Unchanged since 2026-07-04 (last edited, not re-checked)