AI Potluck
Model components / Fine-tuned / chat models

Devstral

Mistral AI

Mistral's open-weight agentic coding model family, purpose-built for software-engineering tasks. The Small 2 instruct variant is a 24B model that runs on a single consumer GPU at a 256K context, and the larger 123B variant targets frontier-class coding performance.

All three axes were rescored onto the governing Devstral 2 release on 2026-08-14, having been read previously on the superseded Small 1.1. Its two SKUs are not licensed alike, so the release resolves most-restrictive. Verified 2026-08-14 via both model cards.

Openness

3 high confidence
3.0
weights
open(Apache-2.0, on HF)
data
closed
code
partial(inference/serving
license
Modified-MIT(non-OSI, resolved most-restrictive across the governing release's SKUs (Devstral 2 123B))+Apache-2.0(OSI, Devstral Small 2 24B SKU only)

Open weights, with the training and post-training data and the pipeline undisclosed, so open weights rather than open source. The tier ships Devstral 2 (2512), which is what the product declares as its artifacts, and openness is read on that current release, resolved to the most restrictive license across the SKUs the release distributes. The two 2512 SKUs are not licensed alike: Devstral Small 2 24B reads "This model is licensed under the Apache 2.0 License", while Devstral 2 123B reads "This model is licensed under a Modified MIT License" and is tagged as an other license on the Hub. These are sizes of one release rather than substitutable alternatives, so the more restrictive of the two governs and the recorded license is permissive but non-OSI. The score is 3 either way - with data closed and code partial the ladder falls through to open weights - but the recorded license must not claim OSI terms the flagship SKU does not offer. Searching the Hub newest-first confirms 2512 is the latest first-party Devstral; everything newer is a community quantization or LoRA.

Adoption

3 high confidence
3.0

A specialist coding-agent model with real but non-mass-market uptake, in the OpenHands and coding-agent niche rather than the chat mainstream. A tier bands on its strongest release, and the two declared artifacts are the Devstral 2 (2512) pair, which sum to 221,596 Hugging Face downloads in the trailing 30 days (Devstral-Small-2-24B-Instruct-2512 188,901; Devstral-2-123B-Instruct-2512 32,695) - inside 100K-1M. The superseded 2507 SKU reads 97,969, a level lower on its own.

Capability

4 high confidence
4.0

A tier is scored on its strongest release, which here is Devstral 2 (2512): its card reports SWE-bench Verified 72.2% for the 123B and 68.0% for Devstral Small 2 24B. 4 rather than 5 because the frontier open coders still sit clearly above it on the same benchmark - Kimi at 80.2% and GLM at SWE-bench Pro 58.4 both score 5 here - and 72.2% lands beside Qwen3-Coder, also 4. It is also above the 69.4% Ornith reports at 4, so 4 is the floor this evidence supports rather than a generous reading. The superseded Devstral Small 1.1 (2507) card reads SWE-bench Verified 53.6%, a fair mid-tier placement for that release but not the one the tier scores on. No peer at this level has been checked recently enough to serve as a comparison anchor, so none is recorded.

Verified 2026-08-14