AI Change Tracker

Models

Frontier and near-frontier AI models, closed and open weights.

L = lab-reported Β· A = aggregator Β· tap a value for all sources
Entity Hard reasoning Coding Agentic / tool use Output speed Cost ↓ Long context Runs well locally on 64GB Apple Silicon Human preference Velocity
GPT-5.2
openai Β· closed
73A
73indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder.
76.8
76.8percent
SWE-bench Verified↗ · as of 2026-08-20
Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β€” the closest thing to watching it do a programmer's job.
Seed placeholder.
60.1
60.1percent
Terminal-Bench↗ · as of 2026-08-18
Drops the model into a computer's command line with a task like "get this software building again" and checks whether the end result actually works β€” a test of getting real multi-step work done, not answering questions.
Seed placeholder.
112A
112tokens_per_secondA
artificial-analysis-speed↗ · as of 2026-08-24
Seed placeholder.
17.5A
17.5usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder.
84
84percent
Fiction.liveBench↗ · as of 2026-08-11
Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β€” advertised capacity and real comprehension often differ a lot.
Seed placeholder β€” 192k bucket.
no
no
owner-notes↗ · as of 2026-08-25
Closed weights; API only.
1475A
1475eloA
LMArena (text)β†— Β· as of 2026-08-18
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
Seed placeholder.
β€”
Claude Opus 4.8
anthropic Β· closed
71A
71indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder β€” replace with live value on first weekly run.
78.2L
78.2percentL
SWE-bench Verified↗ · as of 2026-08-20
Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β€” the closest thing to watching it do a programmer's job.
Seed placeholder. Lab-reported; independent run pending.
62.4
62.4percent
Terminal-Bench↗ · as of 2026-08-18
Drops the model into a computer's command line with a task like "get this software building again" and checks whether the end result actually works β€” a test of getting real multi-step work done, not answering questions.
Seed placeholder.
61A
61tokens_per_secondA
artificial-analysis-speed↗ · as of 2026-08-24
Seed placeholder.
30A
30usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder.
91
91percent
Fiction.liveBench↗ · as of 2026-08-11
Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β€” advertised capacity and real comprehension often differ a lot.
Seed placeholder β€” 192k bucket.
no
no
owner-notes↗ · as of 2026-08-25
Closed weights; API only.
1468A
1468eloA
LMArena (text)β†— Β· as of 2026-08-18
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
Seed placeholder.
β€”
Gemini 3 Pro
google-deepmind Β· closed
70A
70indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder.
72.1
72.1percent
SWE-bench Verified↗ · as of 2026-08-20
Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β€” the closest thing to watching it do a programmer's job.
Seed placeholder.
β€”
138A
138tokens_per_secondA
artificial-analysis-speed↗ · as of 2026-08-24
Seed placeholder.
9A
9usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder.
93
93percent
Fiction.liveBench↗ · as of 2026-08-11
Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β€” advertised capacity and real comprehension often differ a lot.
Seed placeholder β€” 192k bucket; 2M window largest tracked.
no
no
owner-notes↗ · as of 2026-08-25
Closed weights; API only.
1471A
1471eloA
LMArena (text)β†— Β· as of 2026-08-18
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
Seed placeholder.
β€”
DeepSeek V4 new
deepseek Β· open_weights
68A
68indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder β€” within 5 points of frontier at release.
71.5L
71.5percentL
SWE-bench Verified↗ · as of 2026-08-24
Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β€” the closest thing to watching it do a programmer's job.
Seed placeholder. Lab-reported at release; independent run pending.
β€”
84A
84tokens_per_secondA
artificial-analysis-speed↗ · as of 2026-08-24
Seed placeholder β€” first-party API.
2.2A
2.2usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder.
β€”
no
no
owner-notes↗ · as of 2026-08-25
Open weights but ~600B MoE β€” does not fit 64GB even at 4-bit.
β€” +3800/7d
Claude Opus 5 new
anthropic Β· closed
63A
63indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Adaptive Reasoning, Max Effort variant, ranked #1 of 178 models evaluated; Xhigh also 63, High 61. Index v4.1.1 (9 evals) β€” a different scale from the pre-v4.1.1 values stored on entity files.
β€” β€” β€” β€” β€” β€”
1492A
1492eloA
LMArena (text)β†— Β· as of 2026-08-31
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
claude-opus-5-high, #7 on the text overall board; 1493 on the 2026-08-26 fetch.
1493eloA
LMArena (text)β†— Β· as of 2026-08-26
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
claude-opus-5-high variant on the 2026-08-26 board fetch; other axes pending weekly runs.
β€”
Claude Fable 5 new
anthropic Β· closed
62A
62indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Adaptive Reasoning, Max Effort, Opus 4.8 Fallback variant; ranked #3. Index v4.1.1.
β€” β€” β€” β€” β€” β€”
1507A
1507eloA
LMArena (text)β†— Β· as of 2026-08-31
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
Still #1 on the text overall board; 1508 on the 2026-08-26 fetch.
1508eloA
LMArena (text)β†— Β· as of 2026-08-26
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
Top of the text board in the 2026-08-26 fetch; other axes pending weekly runs.
β€”
GPT-5.6 new
openai Β· closed
61A
61indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
GPT-5.6 Sol (max), ranked #5 β€” first index entry for this entity; answers the open hard-problems flag's request for a GPT-5.6 reading. Index v4.1.1.
β€” β€” β€” β€” β€” β€” β€” β€”
Kimi K3 new
moonshot Β· open_weights
60A
60indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Kimi K3 (max) β€” highest-ranked open-weights model of 98 open-weights entries, 3 points off the top model. Index v4.1.1.
β€” β€” β€” β€” β€” β€”
1489A
1489eloA
LMArena (text)β†— Β· as of 2026-08-31
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
kimi-k3-max, #10 and the top open-weights entry on the text board; unchanged in value from the 2026-08-26 fetch, as-of refreshed.
1489eloA
LMArena (text)β†— Β· as of 2026-08-26
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
kimi-k3-max variant, tenth place on the 2026-08-26 board fetch; reported ~2.8T-parameter MoE β€” not locally runnable.
β€”
Qwen3 Coder 32B
qwen Β· open_weights
54A
54indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder.
58.9
58.9percent
SWE-bench Verified↗ · as of 2026-08-20
Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β€” the closest thing to watching it do a programmer's job.
Seed placeholder β€” best open score in its size class.
β€”
24
24tokens_per_second
owner-notes↗ · as of 2026-08-10
Seed placeholder β€” 4-bit MLX on M-series 64GB, owner-measured.
0.9A
0.9usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder β€” hosted API price; local is free.
β€”
yes
yes
owner-notes↗ · as of 2026-08-10
4-bit quant fits in ~20GB; comfortable on 64GB.
β€” β€”
Qwen3 8B
qwen Β· open_weights
38A
38indexA
An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β€” useful for comparing models against each other, not as an absolute measure.
Seed placeholder.
β€” β€”
68
68tokens_per_second
owner-notes↗ · as of 2026-08-10
Seed placeholder β€” 4-bit MLX on M-series 64GB, owner-measured.
0.2A
0.2usd_per_million_blendedA
artificial-analysis-pricing↗ · as of 2026-08-24
Seed placeholder β€” hosted API price; local is free.
β€”
yes
yes
owner-notes↗ · as of 2026-08-10
4-bit quant fits in ~6GB; fast enough for drafts and tooling.
β€” β€”
Gemini 3.7 Flash new
google-deepmind Β· closed
β€” β€” β€” β€” β€” β€” β€”
1490A
1490eloA
LMArena (text)β†— Β· as of 2026-08-31
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
gemini-3.7-flash-high, #9; unchanged in value from the 2026-08-26 fetch, as-of refreshed.
1490eloA
LMArena (text)β†— Β· as of 2026-08-26
Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β€” it tells you which answers people like, which is not the same as which are right.
gemini-3.7-flash-high variant on the 2026-08-26 board fetch; introductory pricing through 2026-12-31 per Google's announcement.
β€”
Qwen3.8-27B new
qwen Β· open_weights
β€”
1595A
1595eloA
lmarena-webdev↗ · as of 2026-08-26
WebDev board, 2026-08-26 fetch (top closed entry sat at 1691). No SWE-bench Verified entry yet.
β€” β€” β€” β€” β€” β€” β€”