AI Potluck
Model components / Training & synthetic datasets

TxT360

LLM360

TxT360 is LLM360's globally deduplicated pretraining corpus combining 99 Common Crawl snapshots with 14 curated high-quality sources (FreeLaw, PG-19, arXiv, Wikipedia, StackExchange, and more) under ODC-BY. It trains K2-V2 together with the TxT360-Midas mid-training and 3efforts SFT extensions.

Family bucket: TxT360 plus the Midas mid-training and 3efforts SFT extensions.

Openness

5 high confidence
5.0
license
ODC-BY
gated
false
dataset_card
present

ODC-BY, ungated, documented pipeline on GitHub.

Adoption

2 high confidence
2.0

4,218 monthly downloads on Hugging Face, graded on the training-corpus bands.

Capability

4 medium confidence
4.0

A self-reported controlled ablation beats FineWeb (8x8B MoE, 1.5T tokens) but without published numeric scores; trains K2 / K2-V2 (LLM360, 2024-25).

Unchanged since 2026-07-04 (last edited, not re-checked)