Swiss AI Initiative
labOpenness profile
1 product on the map — 1 open.
Openness
5 high confidence- weights
- open(Apache-2.0, BF16 safetensors)
- data
- open(full pretraining-data reconstruction scripts, github.com/swiss-ai/pretrain-data)
- code
- open(Megatron-LM training recipe + reconstruction scripts)
- checkpoints
- open(intermediate checkpoints on HF branches)
- license
- Apache-2.0(OSI)
Among the most transparent large-model releases to date: open weights, open data (reconstruction scripts rather than a raw dump, to respect opt-out/robots.txt and strip personal data), open training code/recipe, intermediate checkpoints, a technical report (arXiv 2509.14233), and EU AI Act + Swiss data-protection/copyright compliance docs. All Apache-2.0. Sits in the OLMo/Pythia full-openness tier.
- https://api.github.com/repos/swiss-ai/pretrain-data recorded 2026-07-30
`private: false`, `archived: false`, license `Apache-2.0`, description "Pretraining data reconstruction scripts for Apertus". The reconstruction scripts are published, which is what satisfies data:open by reconstruction rather than by a released corpus.
- https://huggingface.co/swiss-ai/Apertus-70B-2509 recorded 2026-07-30
Model card: "Fully open model: open weights + open data + full training details including all data and training recipes"; "Training Framework: Megatron-LM" linking github.com/swiss-ai/Megatron-LM; "The training intermediate checkpoints are available on the different branches of this same repository"; license apache-2.0; BF16 safetensors.
- https://ethz.ch/en/news-and-events/eth-news/news/2025/09/press-release-apertus-a-fully-open-transparent-multilingual-language-model.html recorded 2026-07-30
ETH Zurich press release, title "Apertus: a fully open, transparent, multilingual language model", dated 2 September 2025 (EPFL/ETH/CSCS). Corroborates the release framing; establishes no dimension on its own, since a press release describes rather than exhibits the artifacts.
- https://github.com/swiss-ai/pretrain-data recorded 2026-07-30
Repository page for the pretraining-data reconstruction scripts; "reconstruction" appears throughout the tree and README.