AI Potluck
Model components / Evaluation code

Compar:IA

DINUM / beta.gouv.fr (French Government)

Blind model-comparison arena created by the French government's betagouv and DINUM with the Ministry of Culture. A user types a prompt, two anonymous models reply, and the user votes; the identities and each model's estimated energy cost are revealed afterward. Its purpose is building large non-English human-preference datasets, and it deploys with Docker on a single server.

A second instance, AI-arenaen, runs in Denmark. Verified 2026-08-13 via the betagouv/ComparIA repository README and its LICENSE.

Openness

5 high confidence
5.0
license
Apache-2.0(OSI
source
public
core-gated
ungated(publicly funded, README documents full Docker self-hosting)

Permissive OSI license with full public source code. The README bills the project as 'The open-source LLM arena. Deploy it for your organisation, sector, or language', with Docker self-hosting instructions for the whole platform and no paid tier to withhold anything for.

Adoption

3 high confidence
3.0

The repository's 74 GitHub stars are beside the point here: compar:IA is a hosted French government LLM arena, and its traction is the data it has collected. The arXiv paper reports over 600,000 free-form prompts and over 250,000 preference votes as of 2026-02-07, about 89% of it in French. That collected data, rather than any repository or user count, is what the level rests on.

Capability

4 high confidence
4.0

A broad arena feature set comparable to LMArena, plus environmental tracking and dataset generation: blind pairwise voting, a public leaderboard, open datasets published to Hugging Face, data.gouv.fr and Mozilla, Docker self-hosting, and EcoLogits maintained on the roadmap. It sits level with LMArena.

  • https://github.com/betagouv/ComparIA recorded 2026-06-22

    Feature list: model comparison, voting, EcoLogits, dataset publishing, self-hosting

  • https://raw.githubusercontent.com/betagouv/ComparIA/develop/README.md recorded 2026-08-13

    'compar:IA is a blind arena for large language models. You type a prompt, two anonymous models reply, and you vote for the answer you prefer'; teaching aims include 'the energy cost of generative AI'; the first public leaderboard shipped Nov 2025; datasets are published to Hugging Face, data.gouv.fr and the Mozilla Data Collective; a second instance, AI-arenaen, runs in Denmark; Docker self-hosting documented; an EcoLogits update is listed on the roadmap.

Verified 2026-08-12