AI Change Tracker

LMArena (text)

https://lmarena.ai/leaderboard

Kind
leaderboard
Maintainer
LMArena (formerly LMSYS)
Feeds axes
human_preference
Measures
Which model's answers ordinary people prefer, when shown two anonymous responses side by side and asked to pick — scored like a chess rating, where beating strong opponents raises your number.
Does not measure
Whether answers are actually correct on hard problems, performance on long projects, anything a voter can't judge in a quick read.
Known issues
Long, confident, nicely formatted answers win votes even when they're wrong; vendors quietly test unreleased models here to tune for votes; the scoring method was tightened in Aug 2026, shuffling rankings.
Trust
medium
Last reviewed
2026-08-16

LMArena shows visitors two answers to the same question, hides which model wrote each, and asks "which is better?" Millions of these votes are turned into a rating using the same math as chess ratings: beat highly-rated opponents often and your rating climbs.

It's the best available measure of "which model do people like talking to" — which matters for a daily assistant — but likability and correctness are different things. A model can win votes by being confident and well-formatted while being subtly wrong.

How to read a score: gaps under about 15 points are a statistical tie. Use the "style-controlled" version of the ranking, which tries to correct for pretty formatting winning votes. And treat sudden ranking jumps right after a scoring-method change (like the August 2026 one) as an artifact of the change until they persist for a few weeks.

← All methods