AI Change Tracker

Verdicts

The owner's current pick per use case — judgment, not data. Scores live in Compare; how they're measured lives in Methods.

Models

High-volume extraction / classification

changed 9d ago

DeepSeek V4 decided 2026-08-26

Tests demonstrate the performance at a fraction of the cost

Evidence: cost · artificial-analysis-pricing · 2026-08-24speed · artificial-analysis-speed · 2026-08-24

Runner-up: Gemini 3 Pro

Weighs: cost, speed

History
  • 2026-04-15 → 2026-08-26: Gemini 3 Pro — Replaced by deepseek-v4 (flag flag-2026-08-24-bulk-cheap).

Methods behind this:

Agentic coding

changed 25d ago

Claude Opus 4.8 decided 2026-08-10

Top of SWE-bench Verified and Terminal-Bench simultaneously, and in my own agentic sessions it recovers from bad states instead of digging in. That recovery behavior is the deciding factor over the raw scores.

Evidence: coding · swe-bench-verified · 2026-08-20agentic · terminal-bench · 2026-08-18long_context · fiction-livebench · 2026-08-11

Runner-up: GPT-5.2

Weighs: coding, agentic, long_context

History
  • 2026-03-02 → 2026-08-10: claude-opus-4-5 — Superseded by 4.8; coding delta of 9 points.

Methods behind this: Aider polyglot leaderboard, Fiction.liveBench, SWE-bench Verified, Terminal-Bench

Default assistant for daily work

Claude Opus 4.8 decided 2026-06-01

Best writing quality and judgment per interaction, and speed is acceptable since the 4.8 serving improvements. The GPT-5.2 price cut narrows the cost argument but hasn't changed what I reach for first.

Evidence: reasoning · artificial-analysis-intelligence-index · 2026-08-18speed · artificial-analysis-speed · 2026-08-24human_preference · lmarena-text · 2026-08-18cost · artificial-analysis-pricing · 2026-08-24

Runner-up: Gemini 3 Pro

Weighs: reasoning, speed, human_preference, cost

History
  • 2026-01-20 → 2026-06-01: claude-opus-4-5 — Superseded by 4.8.

Methods behind this: Artificial Analysis Intelligence Index, LMArena (text)

Hardest reasoning problems, cost no object

GPT-5.2 decided 2026-07-14

Highest AA intelligence index of the current crop and the most reliable on multi-step math and proof-style problems in my own hard-problem scratch tests. Opus 4.8 is close and wins when the problem needs very long context, but 5.2's reasoning consistency edges it for pure difficulty.

Evidence: reasoning · artificial-analysis-intelligence-index · 2026-08-18long_context · fiction-livebench · 2026-08-11

Runner-up: Claude Opus 4.8

Weighs: reasoning, long_context

History
  • 2026-01-20 → 2026-07-14: claude-opus-4-5 — GPT-5.2 overtook on the reasoning index and held it for two months.

Methods behind this: Artificial Analysis Intelligence Index, Fiction.liveBench

Fast local model for drafts and tooling

Qwen3 8B decided 2026-07-05

68 tok/s at 4-bit leaves it fast enough for draft generation, summarization, and tooling glue, and it fits alongside the 32B model in memory. Anything smaller degrades too much on instruction following.

Evidence: local_64gb · owner-notes · 2026-08-10speed · owner-notes · 2026-08-10

Weighs: local_64gb, speed

Methods behind this:

Best model that runs well on 64GB Apple Silicon

Qwen3 Coder 32B decided 2026-07-05

Best model that genuinely fits and runs at usable speed on the 64GB machine. Clearly ahead of everything else local on coding; reasoning is acceptable for offline work. DeepSeek V4 is open but far too large to run locally.

Evidence: local_64gb · owner-notes · 2026-08-10reasoning · artificial-analysis-intelligence-index · 2026-08-18coding · swe-bench-verified · 2026-08-20

Weighs: local_64gb, reasoning, coding

Methods behind this: Aider polyglot leaderboard, Artificial Analysis Intelligence Index, SWE-bench Verified