Changes
Everything that moved: significant events, score updates on current entities, verdict and confidence changes. Full provenance is the repo's git history.
- 2026-09-04 thesis "Agentic replacement of workers is overstated and overhyped as efforts …" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 thesis "Agentic deception, collusion and transcript falsification will continu…" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 thesis "The enterprise AI ROI gap persists despite increased agentic spending…" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 thesis "Nvidia's chips coupled with HuggingFace acquisition will significantly…" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 thesis "OpenAI's predictions are real or marketing? They currently (Sept 3 202…" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 thesis "OpenAI's future success could be limited by senior infrastructure lead…" confidence → 0.5 (Initial — set on adoption; owner to adjust.)
- 2026-09-04 category Category activated: Personal agent frameworks
- 2026-09-03 event OpenAI launches GPT-6 Astra, priced at Fable parity, with disputed benchmark claims and a bumpy rollout
- 2026-09-02 event Report: Anthropic could IPO as soon as September or October, raising up to $100B; OpenAI seen pushed to 2027
- 2026-09-01 event Anthropic launches Claude Fable 5.1 and Mythos 5.1 with cache price cut, removed data retention limits
- 2026-08-31 event METR/Redwood report reveals OpenAI eval agents used deception, collusion to attack Hugging Face
- 2026-08-31 score Claude Fable 5 · reasoning → 62 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-31 score Claude Fable 5 · human_preference → 1507 elo (lmarena-text, aggregator)
- 2026-08-31 score Claude Opus 5 · reasoning → 63 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-31 score Claude Opus 5 · human_preference → 1492 elo (lmarena-text, aggregator)
- 2026-08-31 score Gemini 3.7 Flash · human_preference → 1490 elo (lmarena-text, aggregator)
- 2026-08-31 score GPT-5.6 · reasoning → 61 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-31 score Kimi K3 · reasoning → 60 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-31 score Kimi K3 · human_preference → 1489 elo (lmarena-text, aggregator)
- 2026-08-28 event Meta AI business leader Clara Shih departs to launch nonprofit after concluding agents already collapsed entry-level roles
- 2026-08-28 event Federal judge rules Trump administration illegally labeled Anthropic a supply-chain risk
- 2026-08-27 event Salesforce and Anthropic launch 'Claudeforce,' embedding Claude across Salesforce's enterprise products
- 2026-08-27 event Nvidia warns of supply bottlenecks into 2028 despite record data-center revenue
- 2026-08-27 event Mystery model 'Ox Alpha' confirmed as GLM-5.3-Flash; independent quantization and serving benchmarks emerge
- 2026-08-27 event Security researcher reports an 80%-success prompt-injection bypass of Claude Code's Auto Mode safety layer
- 2026-08-26 event Nvidia to acquire Hugging Face for roughly $13B
- 2026-08-26 event Meta settles child-safety suits with US states for up to $17.1B, adds teen usage limits
- 2026-08-26 event Z.ai formally launches GLM-5.3-Flash, revealing it as the previously teased 'Ox Alpha'
- 2026-08-26 event Anthropic signs $45B compute deal with Nscale
- 2026-08-26 score Claude Fable 5 · human_preference → 1508 elo (lmarena-text, aggregator)
- 2026-08-26 score Claude Opus 5 · human_preference → 1493 elo (lmarena-text, aggregator)
- 2026-08-26 score Gemini 3.7 Flash · human_preference → 1490 elo (lmarena-text, aggregator)
- 2026-08-26 score Kimi K3 · human_preference → 1489 elo (lmarena-text, aggregator)
- 2026-08-26 score Qwen3.8-27B · coding → 1595 elo (lmarena-webdev, aggregator)
- 2026-08-26 verdict models/bulk-cheap → deepseek-v4
- 2026-08-26 verdict models/bulk-cheap: gemini-3-pro retired — Replaced by deepseek-v4 (flag flag-2026-08-24-bulk-cheap).
- 2026-08-26 thesis "Google's need to protect its consumer/ads core will keep it from leadi…" confidence → 0.6 (Initial — drafted from owner's stated view; owner to adjust.)
- 2026-08-26 thesis "A lab leader's established pattern — especially the apologize-without-…" confidence → 0.7 (Initial — drafted from owner's stated view; owner to adjust.)
- 2026-08-26 category Category activated: Labs
- 2026-08-25 event OpenAI's top data center executive departs amid continued senior-leadership exits
- 2026-08-25 event McKinsey's 2026 State of AI survey: enterprise conviction and agentic AI scaling outpace measurable ROI
- 2026-08-25 score Claude Opus 4.8 · local_64gb → no (owner-notes, independent)
- 2026-08-25 score DeepSeek V4 · local_64gb → no (owner-notes, independent)
- 2026-08-25 score Gemini 3 Pro · local_64gb → no (owner-notes, independent)
- 2026-08-25 score GPT-5.2 · local_64gb → no (owner-notes, independent)
- 2026-08-25 category Category activated: Models
- 2026-08-24 event OpenAI leadership signals internal AGI declaration timeline for late 2026
- 2026-08-24 event GPT-5.6 rolls out to Kiro and appears as OpenAI's default in its own docs, with no verdict-holder gpt-5-2 benchmark in sight
- 2026-08-24 event DeepSeek releases V4 open weights under MIT license
- 2026-08-24 score Claude Opus 4.8 · speed → 61 tokens_per_second (artificial-analysis-speed, aggregator)
- 2026-08-24 score Claude Opus 4.8 · cost → 30 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 score DeepSeek V4 · reasoning → 68 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-24 score DeepSeek V4 · coding → 71.5 percent (swe-bench-verified, lab)
- 2026-08-24 score DeepSeek V4 · speed → 84 tokens_per_second (artificial-analysis-speed, aggregator)
- 2026-08-24 score DeepSeek V4 · cost → 2.2 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 score Gemini 3 Pro · speed → 138 tokens_per_second (artificial-analysis-speed, aggregator)
- 2026-08-24 score Gemini 3 Pro · cost → 9 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 score GPT-5.2 · speed → 112 tokens_per_second (artificial-analysis-speed, aggregator)
- 2026-08-24 score GPT-5.2 · cost → 17.5 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 score Qwen3 8B · cost → 0.2 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 score Qwen3 Coder 32B · cost → 0.9 usd_per_million_blended (artificial-analysis-pricing, aggregator)
- 2026-08-24 thesis "Open-weights models will remain within one generation of the closed fr…" confidence → 0.7 (DeepSeek V4 within 5 AA points)
- 2026-08-20 event Claude Opus 4.8 posts 78.2 on SWE-bench Verified
- 2026-08-20 score Claude Opus 4.8 · coding → 78.2 percent (swe-bench-verified, lab)
- 2026-08-20 score Gemini 3 Pro · coding → 72.1 percent (swe-bench-verified, independent)
- 2026-08-20 score GPT-5.2 · coding → 76.8 percent (swe-bench-verified, independent)
- 2026-08-20 score Qwen3 Coder 32B · coding → 58.9 percent (swe-bench-verified, independent)
- 2026-08-18 event OpenAI cuts GPT-5.2 API prices by 40%
- 2026-08-18 score Claude Opus 4.8 · reasoning → 71 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-18 score Claude Opus 4.8 · agentic → 62.4 percent (terminal-bench, independent)
- 2026-08-18 score Claude Opus 4.8 · human_preference → 1468 elo (lmarena-text, aggregator)
- 2026-08-18 score Gemini 3 Pro · reasoning → 70 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-18 score Gemini 3 Pro · human_preference → 1471 elo (lmarena-text, aggregator)
- 2026-08-18 score GPT-5.2 · reasoning → 73 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-18 score GPT-5.2 · agentic → 60.1 percent (terminal-bench, independent)
- 2026-08-18 score GPT-5.2 · human_preference → 1475 elo (lmarena-text, aggregator)
- 2026-08-18 score Qwen3 8B · reasoning → 38 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-18 score Qwen3 Coder 32B · reasoning → 54 index (artificial-analysis-intelligence-index, aggregator)
- 2026-08-11 score Claude Opus 4.8 · long_context → 91 percent (fiction-livebench, independent)
- 2026-08-11 score Gemini 3 Pro · long_context → 93 percent (fiction-livebench, independent)
- 2026-08-11 score GPT-5.2 · long_context → 84 percent (fiction-livebench, independent)
- 2026-08-10 score Qwen3 8B · speed → 68 tokens_per_second (owner-notes, independent)
- 2026-08-10 score Qwen3 8B · local_64gb → yes (owner-notes, independent)
- 2026-08-10 score Qwen3 Coder 32B · speed → 24 tokens_per_second (owner-notes, independent)
- 2026-08-10 score Qwen3 Coder 32B · local_64gb → yes (owner-notes, independent)
- 2026-08-10 verdict models/coding-agent → claude-opus-4-8
- 2026-08-10 verdict models/coding-agent: claude-opus-4-5 retired — Superseded by 4.8; coding delta of 9 points.
- 2026-08-05 thesis "By end of 2027, a model that runs well on 64GB Apple Silicon will matc…" confidence → 0.45 (Qwen agent runtime closes part of the scaffold gap)
- 2026-07-14 verdict models/hard-problems → gpt-5-2
- 2026-07-14 verdict models/hard-problems: claude-opus-4-5 retired — GPT-5.2 overtook on the reasoning index and held it for two months.
- 2026-07-05 verdict models/local-fast → qwen3-8b
- 2026-07-05 verdict models/local-heavy → qwen3-coder-32b
- 2026-06-01 verdict models/daily-driver → claude-opus-4-8
- 2026-06-01 verdict models/daily-driver: claude-opus-4-5 retired — Superseded by 4.8.
- 2026-05-01 thesis "By end of 2027, a model that runs well on 64GB Apple Silicon will matc…" confidence → 0.4 (Initial)
- 2026-04-01 thesis "Open-weights models will remain within one generation of the closed fr…" confidence → 0.6 (Initial)