AI Potluck
Model components / Training & synthetic datasets

TxT360

LLM360

LLM360's globally deduplicated pretraining corpus, combining 99 Common Crawl snapshots with 14 curated sources including FreeLaw, PG-19, arXiv, Wikipedia and StackExchange. It holds about 5 trillion deduplicated tokens and supports a recipe reaching beyond 15 trillion, and trains K2-V2 together with its mid-training and SFT extensions.

A family record covering TxT360 with the Midas mid-training and 3efforts SFT extensions. The LLM360 path now redirects to the IFM organization. Verified 2026-08-13 via the TxT360 dataset card on Hugging Face.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, documented pipeline on GitHub. The LLM360 path redirects to the IFM org on the Hub, same repo.

Adoption

2 high confidence
2.0

9,230 downloads in the trailing 30 days for IFM/TxT360, which bands at 1K-10K, level 2 on the dataset adoption scale.

Capability

4 medium confidence
4.0

A self-reported controlled ablation beats FineWeb (8x8B MoE, 1.5T tokens) but without published numeric scores; trains K2 / K2-V2 (LLM360, 2024-25). The card describes ~5T globally deduplicated tokens and a mixing recipe reaching 15T+, and the K2-V2 collection lists it.

  • https://huggingface.co/datasets/LLM360/TxT360 recorded 2026-08-13

    Card - ~5T globally deduplicated tokens across 99 Common Crawl snapshots and 14 non-web sources, a recipe reaching 15T+ tokens, and a 1.5T-token 8x8B MoE ablation against FineWeb reported without numeric scores; listed in the K2-V2 collection

Verified 2026-08-13