Battle Arena
Two anonymous models answer your prompt β vote for the better response. Rankings update live via Elo. Runs 100% in your browser.
Leaderboard
Models ranked by Elo rating from battle votes. Bars show the 95% confidence interval.
| # | Model | Arena Score | 95% CI | Votes | Organization | License |
|---|
Elo: K=32 Β· ties and "both bad" count as draws Β· ratings seeded with priors, adjusted by real votes Β· CI via Wilson approximation on win-rate. Methodology inspired by lmarena/arena-rank.
Models
The catalog available for battle. Click Battle to challenge a model against a random opponent.
Battle History
Your past battles and votes (stored in your browser).
π Distillation
Your best-of collection β responses you saved during battles. Inspired by arena-ai's distillation idea: collect the best evaluated responses as training data for your own model.
Settings
Safety & evaluation features extracted from the Sarus Arena framework, plus data management.
- AB-testing battle β side-by-side anonymous comparison (arena-ai / lmarena core)
- User feedback evaluation β votes feed an Elo rating system
- Formula-based evaluation β measurable response metrics: length, structure, code blocks, readability
- LLM as a judge β auto-judge scores helpfulness, clarity, accuracy & formatting (heuristic simulation)
- PII removal β email / phone / IP redaction
- Guardrailing β profanity detection & censoring
- Evaluation-based routing β "Route to winner" sends follow-ups to the top-rated model
- Distillation β save the best evaluated responses as a dataset
Everything lives in your browser's localStorage β no server, no account, no tracking.
An educational re-implementation inspired by lmarena.ai (LMSYS Chatbot Arena), github.com/lmarena and the Sarus Arena framework (arena-ai/arena). Not affiliated with either project. All model personas are simulated.