AI Change Tracker

Artificial Analysis Intelligence Index

https://artificialanalysis.ai/models

Kind
aggregator
Maintainer
Artificial Analysis
Feeds axes
reasoning
Measures
General problem-solving ability — a blend of reasoning, knowledge, math, and coding tests, all run by the same independent team on each model's public service, so every model faces identical conditions.
Does not measure
Long multi-step projects, taste and judgment in writing, performance on very long documents, anything requiring the model to use tools.
Known issues
The recipe for blending the tests changes over time, so scores from different index versions aren't directly comparable.
Trust
high
Last reviewed
2026-08-01

Artificial Analysis is an independent measurement firm — think of them as a Consumer Reports for AI models. They run a fixed set of tests against each model's public service themselves, rather than trusting vendor claims, and combine the results into a single number. Because every model takes the same tests under the same conditions, comparing two models' index scores is meaningful in a way that comparing two vendors' press releases is not.

How to read a score: treat it as a ranking tool, not a grade. A 5-point gap means one model is clearly a tier above; a 1–2 point gap is within the noise of measurement. Check that two scores come from the same index version before comparing them — the recipe gets revised periodically.

The same firm also measures output speed, price, and performance on multi-step tasks; this site uses those measurements for the speed, cost, and agentic columns.

← All methods