Models
Frontier and near-frontier AI models, closed and open weights.
L = lab-reported Β·
A = aggregator Β· tap a value for all sources
| Entity | Hard reasoning | Coding | Agentic / tool use | Output speed | Cost β | Long context | Runs well locally on 64GB Apple Silicon | Human preference | Velocity |
|---|---|---|---|---|---|---|---|---|---|
| GPT-5.2 openai Β· closed | 73A73indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-18 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder. | 76.876.8percent SWE-bench Verifiedβ
Β· as of 2026-08-20 Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β the closest thing to watching it do a programmer's job. Seed placeholder. | 60.160.1percent Terminal-Benchβ
Β· as of 2026-08-18 Drops the model into a computer's command line with a task like "get this software building again" and checks whether the end result actually works β a test of getting real multi-step work done, not answering questions. Seed placeholder. | 112A | 17.5A | 8484percent Fiction.liveBenchβ
Β· as of 2026-08-11 Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β advertised capacity and real comprehension often differ a lot. Seed placeholder β 192k bucket. | no | 1475A1475eloA LMArena (text)β
Β· as of 2026-08-18 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. Seed placeholder. | β |
| Claude Opus 4.8 anthropic Β· closed | 71A71indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-18 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder β replace with live value on first weekly run. | 78.2L78.2percentL SWE-bench Verifiedβ
Β· as of 2026-08-20 Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β the closest thing to watching it do a programmer's job. Seed placeholder. Lab-reported; independent run pending. | 62.462.4percent Terminal-Benchβ
Β· as of 2026-08-18 Drops the model into a computer's command line with a task like "get this software building again" and checks whether the end result actually works β a test of getting real multi-step work done, not answering questions. Seed placeholder. | 61A | 30A | 9191percent Fiction.liveBenchβ
Β· as of 2026-08-11 Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β advertised capacity and real comprehension often differ a lot. Seed placeholder β 192k bucket. | no | 1468A1468eloA LMArena (text)β
Β· as of 2026-08-18 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. Seed placeholder. | β |
| Gemini 3 Pro google-deepmind Β· closed | 70A70indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-18 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder. | 72.172.1percent SWE-bench Verifiedβ
Β· as of 2026-08-20 Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β the closest thing to watching it do a programmer's job. Seed placeholder. | β | 138A | 9A | 9393percent Fiction.liveBenchβ
Β· as of 2026-08-11 Tests whether a model that claims to handle book-length input actually keeps the story straight all the way through, or quietly loses the plot β advertised capacity and real comprehension often differ a lot. Seed placeholder β 192k bucket; 2M window largest tracked. | no | 1471A1471eloA LMArena (text)β
Β· as of 2026-08-18 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. Seed placeholder. | β |
| DeepSeek V4 new deepseek Β· open_weights | 68A68indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-24 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder β within 5 points of frontier at release. | 71.5L71.5percentL SWE-bench Verifiedβ
Β· as of 2026-08-24 Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β the closest thing to watching it do a programmer's job. Seed placeholder. Lab-reported at release; independent run pending. | β | 84A84tokens_per_secondA artificial-analysis-speedβ
Β· as of 2026-08-24 Seed placeholder β first-party API. | 2.2A | β | no | β | +3800/7d |
| Claude Opus 5 new anthropic Β· closed | 63A63indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-31 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Adaptive Reasoning, Max Effort variant, ranked #1 of 178 models evaluated; Xhigh also 63, High 61. Index v4.1.1 (9 evals) β a different scale from the pre-v4.1.1 values stored on entity files. | β | β | β | β | β | β | 1492A1492eloA LMArena (text)β
Β· as of 2026-08-31 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. claude-opus-5-high, #7 on the text overall board; 1493 on the 2026-08-26 fetch. 1493eloA LMArena (text)β
Β· as of 2026-08-26 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. claude-opus-5-high variant on the 2026-08-26 board fetch; other axes pending weekly runs. | β |
| Claude Fable 5 new anthropic Β· closed | 62A62indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-31 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Adaptive Reasoning, Max Effort, Opus 4.8 Fallback variant; ranked #3. Index v4.1.1. | β | β | β | β | β | β | 1507A1507eloA LMArena (text)β
Β· as of 2026-08-31 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. Still #1 on the text overall board; 1508 on the 2026-08-26 fetch. 1508eloA LMArena (text)β
Β· as of 2026-08-26 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. Top of the text board in the 2026-08-26 fetch; other axes pending weekly runs. | β |
| GPT-5.6 new openai Β· closed | 61A61indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-31 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. GPT-5.6 Sol (max), ranked #5 β first index entry for this entity; answers the open hard-problems flag's request for a GPT-5.6 reading. Index v4.1.1. | β | β | β | β | β | β | β | β |
| Kimi K3 new moonshot Β· open_weights | 60A60indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-31 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Kimi K3 (max) β highest-ranked open-weights model of 98 open-weights entries, 3 points off the top model. Index v4.1.1. | β | β | β | β | β | β | 1489A1489eloA LMArena (text)β
Β· as of 2026-08-31 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. kimi-k3-max, #10 and the top open-weights entry on the text board; unchanged in value from the 2026-08-26 fetch, as-of refreshed. 1489eloA LMArena (text)β
Β· as of 2026-08-26 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. kimi-k3-max variant, tenth place on the 2026-08-26 board fetch; reported ~2.8T-parameter MoE β not locally runnable. | β |
| Qwen3 Coder 32B qwen Β· open_weights | 54A54indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-18 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder. | 58.958.9percent SWE-bench Verifiedβ
Β· as of 2026-08-20 Gives the model real bug reports from actual software projects and checks whether its fix passes the project's own tests β the closest thing to watching it do a programmer's job. Seed placeholder β best open score in its size class. | β | 2424tokens_per_second owner-notesβ
Β· as of 2026-08-10 Seed placeholder β 4-bit MLX on M-series 64GB, owner-measured. | 0.9A0.9usd_per_million_blendedA artificial-analysis-pricingβ
Β· as of 2026-08-24 Seed placeholder β hosted API price; local is free. | β | yes | β | β |
| Qwen3 8B qwen Β· open_weights | 38A38indexA Artificial Analysis Intelligence Indexβ
Β· as of 2026-08-18 An independent firm runs the same battery of tests on every model and blends the results into one overall smarts score β useful for comparing models against each other, not as an absolute measure. Seed placeholder. | β | β | 6868tokens_per_second owner-notesβ
Β· as of 2026-08-10 Seed placeholder β 4-bit MLX on M-series 64GB, owner-measured. | 0.2A0.2usd_per_million_blendedA artificial-analysis-pricingβ
Β· as of 2026-08-24 Seed placeholder β hosted API price; local is free. | β | yes | β | β |
| Gemini 3.7 Flash new google-deepmind Β· closed | β | β | β | β | β | β | β | 1490A1490eloA LMArena (text)β
Β· as of 2026-08-31 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. gemini-3.7-flash-high, #9; unchanged in value from the 2026-08-26 fetch, as-of refreshed. 1490eloA LMArena (text)β
Β· as of 2026-08-26 Thousands of people vote blind on which of two anonymous answers they prefer, producing a popularity rating β it tells you which answers people like, which is not the same as which are right. gemini-3.7-flash-high variant on the 2026-08-26 board fetch; introductory pricing through 2026-12-31 per Google's announcement. | β |
| Qwen3.8-27B new qwen Β· open_weights | β | 1595A1595eloA lmarena-webdevβ
Β· as of 2026-08-26 WebDev board, 2026-08-26 fetch (top closed entry sat at 1691). No SWE-bench Verified entry yet. | β | β | β | β | β | β | β |