TxT360
LLM360LLM360's globally deduplicated pretraining corpus, combining 99 Common Crawl snapshots with 14 curated sources including FreeLaw, PG-19, arXiv, Wikipedia and StackExchange. It holds about 5 trillion deduplicated tokens and supports a recipe reaching beyond 15 trillion, and trains K2-V2 together with its mid-training and SFT extensions.
A family record covering TxT360 with the Midas mid-training and 3efforts SFT extensions. The LLM360 path now redirects to the IFM organization. Verified 2026-08-13 via the TxT360 dataset card on Hugging Face.
Openness
5 high confidence- license
- ODC-BY
- gated
- false
- dataset_card
- present
ODC-BY, ungated, documented pipeline on GitHub. The LLM360 path redirects to the IFM org on the Hub, same repo.
- https://huggingface.co/datasets/LLM360/TxT360 recorded 2026-08-13
ODC-BY license, ungated, dataset card
Adoption
2 high confidence9,230 downloads in the trailing 30 days for IFM/TxT360, which bands at 1K-10K, level 2 on the dataset adoption scale.
- https://huggingface.co/api/datasets/LLM360/TxT360 recorded 2026-08-13
9,230 downloads in the trailing 30 days for IFM/TxT360
Capability
4 medium confidenceA self-reported controlled ablation beats FineWeb (8x8B MoE, 1.5T tokens) but without published numeric scores; trains K2 / K2-V2 (LLM360, 2024-25). The card describes ~5T globally deduplicated tokens and a mixing recipe reaching 15T+, and the K2-V2 collection lists it.
- https://huggingface.co/datasets/LLM360/TxT360 recorded 2026-08-13
Card - ~5T globally deduplicated tokens across 99 Common Crawl snapshots and 14 non-web sources, a recipe reaching 15T+ tokens, and a 1.5T-token 8x8B MoE ablation against FineWeb reported without numeric scores; listed in the K2-V2 collection
Verified 2026-08-13