LMArena (text)
https://lmarena.ai/leaderboard
- Kind
- leaderboard
- Maintainer
- LMArena (formerly LMSYS)
- Feeds axes
- human_preference
- Measures
- Which model's answers ordinary people prefer, when shown two anonymous responses side by side and asked to pick — scored like a chess rating, where beating strong opponents raises your number.
- Does not measure
- Whether answers are actually correct on hard problems, performance on long projects, anything a voter can't judge in a quick read.
- Known issues
- Long, confident, nicely formatted answers win votes even when they're wrong; vendors quietly test unreleased models here to tune for votes; the scoring method was tightened in Aug 2026, shuffling rankings.
- Trust
- medium
- Last reviewed
- 2026-08-16
LMArena shows visitors two answers to the same question, hides which model wrote each, and asks "which is better?" Millions of these votes are turned into a rating using the same math as chess ratings: beat highly-rated opponents often and your rating climbs.
It's the best available measure of "which model do people like talking to" — which matters for a daily assistant — but likability and correctness are different things. A model can win votes by being confident and well-formatted while being subtly wrong.
How to read a score: gaps under about 15 points are a statistical tie. Use the "style-controlled" version of the ranking, which tries to correct for pretty formatting winning votes. And treat sudden ranking jumps right after a scoring-method change (like the August 2026 one) as an artifact of the change until they persist for a few weeks.