Skip to content

Latest commit

 

History

7,001 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
compar:IA logo

compar:IA

An open-source LLM arena for your organisation, sector, or language.

Collect human votes, compare models through real use, and publish open datasets for any language or sector.


License Hugging Face datasets Paper DPG Badge

Try the arena · Leaderboard · Walkthrough · Datasets · Deploy your own · Roadmap · Contribute


What is compar:IA?

compar:IA is an LLM arena. Enter a prompt and two anonymous models respond. Vote for the answer you prefer or skip the vote; the model names are revealed only afterwards. The French public arena is free and does not require an account.

The model catalogue lists each model's origin and technical characteristics. Where enough technical data is available, compar:IA also shows an EcoLogits estimate of its energy use.

Frame 15928

Walkthrough video

compar-ia-walkthrough.mp4

What can you use it for?

  • Raise awareness about differences between models, including bias, openness, and energy consumption.
  • Rank models through real-world use rather than laboratory benchmarks, for a specific use case and language.
  • Publish open datasets of prompts, votes, and reactions for research, training, and fine-tuning.

Leaderboard and open data

The public leaderboard converts blind votes into Bradley-Terry scores with 95% confidence intervals. It measures the preferences collected in the arena, not objective model quality. The methodology, calculations, and source data are public.

The comparia-fr-arena dataset is published under the Etalab Open Licence 2.0 and CC BY 4.0.

Project history

  • Oct 2024: The French government launched comparia.beta.gouv.fr, a public LLM arena.
  • Mar 2025: The arena reached 50,000 votes, and the first dataset was published on Hugging Face.
  • Nov 2025: The first public leaderboard was released, and compar:IA was recognized as a Digital Public Good. A second instance, Denmark's AI-arenaen, also went live.
  • Jun 2026: The project passed 700,000 conversations and 250,000 votes (about 89% in French), with more than 130 models tested and several datasets published.
  • Sept 2026: compar:IA 2.0 was released with message history, personal leaderboards, and an admin panel. Companies, sectors, and language communities can now deploy their own instance.

Active instances

Instance Region Live Datasets
compar:IA, French Government 🇫🇷 comparia.beta.gouv.fr Hugging Face, data.gouv.fr, Mozilla
AI-arenaen, Denmark 🇩🇰 ai-arenaen.dk Coming soon
Yours? 🌍 Deploy one Your own

Deploy your own

Host compar:IA with your own models, language, datasets, and leaderboard.

Self-host with Docker: Run the platform on a single server with automatic HTTPS from Caddy. Follow the self-hosting guide, then configure your models, languages and branding in the admin panel.

Local development:

cp .env.example .env   # configure your environment variables
make install           # install all dependencies
source .env

make dev-backend       # backend  -> http://localhost:8008
make dev-frontend      # frontend -> http://localhost:5173

See CONTRIBUTING.md for the full setup: instances, Docker, the database, testing, and translations.

Roadmap

In progress

  • Agentic tool use
  • OIDC SSO
  • Improved style control
Done
  • Authentication (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Style control, #532 (🇫🇷 Ministry of Culture)
  • Prompt moderation, #542 (🇫🇷 Ministry of Culture)
  • Improved model cards (🇫🇷 Ministry of Culture)
  • Live use-case mapping (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Message history (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Socio-demographic data collection (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Back-office management (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • New voting system (🇪🇺 ALT-EDIC, 🇫🇷 Ministry of Culture)
  • Web search (🇪🇺 ALT-EDIC)
  • Separation of all platforms into separate instances (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Ranking consolidation and internationalization (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Language / platform-specific model support (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Gradio to FastAPI migration (🇫🇷 Ministry of Culture, 🇫🇷 DINUM, 🇪🇺 ALT-EDIC)
  • EcoLogits update (🇪🇺 ALT-EDIC, 🇫🇷 DINUM)
  • Dataset publishing pipeline v1 (🇫🇷 DINUM, 🇫🇷 Ministry of Culture)
  • Leaderboard v1 (🇫🇷 DINUM, 🇫🇷 Ministry of Culture, with 🇫🇷 PEReN)
  • Archived models (🇫🇷 DINUM, 🇫🇷 Ministry of Culture)
  • Blog section (🇫🇷 DINUM, 🇫🇷 Ministry of Culture)
  • Internationalization foundations (🇫🇷 DINUM, 🇫🇷 Ministry of Culture)
  • compar:IA v1 (🇫🇷 DINUM, 🇫🇷 Ministry of Culture)

Contribute

compar:IA is a digital common. You can support it by running an instance, funding the work, contributing code or translations, or sharing research and ideas.

  • Run an instance: Each deployment can produce an open dataset and benchmark for its language, sector, or organisation.
  • Fund the project: compar:IA is funded by ALT-EDIC, DINUM, and the French Ministry of Culture. New partners and funders help cover infrastructure, add languages, and keep the project independent. Contact contact@comparia.beta.gouv.fr.
  • Contribute code or translations: Bug fixes, features, translations, and documentation can be submitted through a pull request.
  • Share ideas or report issues: Start or join a thread in GitHub Discussions.
  • Report a vulnerability: Write to us first, not in a public issue. See SECURITY.md.
  • Research and partnerships: For academic work, media enquiries, partnerships, or other forms of support, get in touch.

Built by

DINUM, the French Ministry of Culture, and ALT-EDIC 🇪🇺, with AI-arenaen (Denmark), PIX, PEReN, and other contributors.


About

Open source LLM arena created by the French Government

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

85 stars

Watchers

2 watching

Forks

Releases

Used by

Contributors

Languages