AI Potluck
Model components / Base / pretrained models

Marin

Stanford CRFM

Stanford CRFM's Marin, a fully-open 'open lab' foundation model (Marin 8B Base/Instruct and 32B Base), built in JAX/Levanter with bit-for-bit reproducibility. Code, data, experiments, hyperparameters and training logs are all documented and released.

~12.7T-token training run with full data/code/logs and bit-reproducible training. Fully open but below OLMo 3 on fully-open evals; modest adoption. The distinguishing move is procedural rather than about the artifacts - every experiment is tracked as a public GitHub issue, so the record of how the model was reached is open as well as the recipe. Verified 2026-08-13 via the Marin announcement post, the marin-8b-base HF model card and the marin-community/marin repository.

Openness

5 high confidence
5.0
weights
open(Apache-2.0)
data
open(documented mixtures)
code
open(marin + Levanter, JAX)
checkpoints
open
reproducibility
bit-for-bit
license
Apache-2.0(OSI)

Stanford CRFM 'open lab': Apache-2.0 weights plus documented code, data, experiments, hyperparameters and training logs, with bit-for-bit reproducible training.

Adoption

2 low confidence
2.0

12,168 downloads in the trailing 30 days across the two declared artifacts (marin-community/marin-8b-base 8,604; marin-community/marin-8b-instruct 3,564), which bands at level 2 (10K-100K) on the model adoption scale. A research-stage fully-open project, on the map as a full-openness exemplar rather than for its popularity.

Capability

3 medium confidence
3.0

Mid-tier capability for a fully-open model at 8B and 32B, a step below OLMo 3. Marin's own announcement puts 8B Base ahead of Llama 3.1 8B Base on 14 of 19 evaluations, and 8B Instruct ahead of OLMo 2 but short of Llama 3.1 Tulu, while Ai2's Olmo 3 post claims the strongest performance among fully open base models - so placing Marin one step under OLMo is what both vendors' own pages say.

Verified 2026-08-13