Verdicts
The owner's current pick per use case — judgment, not data. Scores live in Compare; how they're measured lives in Methods.
Models
High-volume extraction / classification
changed 9d agoDeepSeek V4 decided 2026-08-26
Tests demonstrate the performance at a fraction of the cost
Runner-up: Gemini 3 Pro
Weighs: cost, speed
History
- 2026-04-15 → 2026-08-26: Gemini 3 Pro — Replaced by deepseek-v4 (flag flag-2026-08-24-bulk-cheap).
Methods behind this:
Agentic coding
changed 25d agoClaude Opus 4.8 decided 2026-08-10
Top of SWE-bench Verified and Terminal-Bench simultaneously, and in my own agentic sessions it recovers from bad states instead of digging in. That recovery behavior is the deciding factor over the raw scores.
Runner-up: GPT-5.2
Weighs: coding, agentic, long_context
History
- 2026-03-02 → 2026-08-10: claude-opus-4-5 — Superseded by 4.8; coding delta of 9 points.
Methods behind this: Aider polyglot leaderboard, Fiction.liveBench, SWE-bench Verified, Terminal-Bench
Default assistant for daily work
Claude Opus 4.8 decided 2026-06-01
Best writing quality and judgment per interaction, and speed is acceptable since the 4.8 serving improvements. The GPT-5.2 price cut narrows the cost argument but hasn't changed what I reach for first.
Runner-up: Gemini 3 Pro
Weighs: reasoning, speed, human_preference, cost
History
- 2026-01-20 → 2026-06-01: claude-opus-4-5 — Superseded by 4.8.
Methods behind this: Artificial Analysis Intelligence Index, LMArena (text)
Hardest reasoning problems, cost no object
GPT-5.2 decided 2026-07-14
Highest AA intelligence index of the current crop and the most reliable on multi-step math and proof-style problems in my own hard-problem scratch tests. Opus 4.8 is close and wins when the problem needs very long context, but 5.2's reasoning consistency edges it for pure difficulty.
Runner-up: Claude Opus 4.8
Weighs: reasoning, long_context
History
- 2026-01-20 → 2026-07-14: claude-opus-4-5 — GPT-5.2 overtook on the reasoning index and held it for two months.
Methods behind this: Artificial Analysis Intelligence Index, Fiction.liveBench
Fast local model for drafts and tooling
Qwen3 8B decided 2026-07-05
68 tok/s at 4-bit leaves it fast enough for draft generation, summarization, and tooling glue, and it fits alongside the 32B model in memory. Anything smaller degrades too much on instruction following.
Weighs: local_64gb, speed
Methods behind this:
Best model that runs well on 64GB Apple Silicon
Qwen3 Coder 32B decided 2026-07-05
Best model that genuinely fits and runs at usable speed on the 64GB machine. Clearly ahead of everything else local on coding; reasoning is acceptable for offline work. DeepSeek V4 is open but far too large to run locally.
Weighs: local_64gb, reasoning, coding
Methods behind this: Aider polyglot leaderboard, Artificial Analysis Intelligence Index, SWE-bench Verified