Open-weights models will remain within one generation of the closed frontier through end of 2027.
confidence 0.7 Β· horizon 2027-12-31 Β· status open
Falsifiers
- No open-weights model within 10 AA index points of frontier for 6 consecutive months.
- Major open-weights lab converts to closed releases.
Evidence
- +1 DeepSeek opens weights for DeepSeek-V4-Flash-Vision-Exp, adding vision to the V4 Flash line β An open, vision-capable DeepSeek V4 variant broadens open-weights feature parity with closed frontier multi-modal models, though the release is unverified independently.
- +1 Z.ai open-weights the full GLM-5.3 model (744B/40B active), positioned for agentic coding and cyber defense β A second Z.ai open-weights release (full GLM-5.3, 744B/40B active) within the same week extends the open frontier, though the claims are curator-relayed and unverified, capping the weight.
- +1 Tencent releases Hunyuan Hy4-preview, a 770B open-weights MoE claimed to lead SWE-bench Pro and jump on WebDev arena β A fourth Chinese lab (Tencent) reportedly fielding a top-tier open MoE would broaden the open-weights field beyond DeepSeek/Qwen/Z.ai, but claims are social-media-relayed and unverified, capping the weight.
- +1 Mystery model 'Ox Alpha' confirmed as GLM-5.3-Flash; independent quantization and serving benchmarks emerge β Independent quantization and serving benchmarks (Unsloth, Baseten, Databricks, Together AI) corroborate GLM-5.3-Flash's price-performance beyond the initial lab announcement, reinforcing the narrow open/closed gap.
- +1 Z.ai confirmed as the lab behind Ox Alpha, a mysterious open model reportedly topping leaderboards β A reportedly leaderboard-topping new open model would extend the open-weights field beyond DeepSeek/Qwen, but the claim is unverified and weights are not yet released, so weight is capped low.
- +2 Z.ai formally launches GLM-5.3-Flash, revealing it as the previously teased 'Ox Alpha' β An MIT-licensed 320B/18B-active model that independent measurement flagged for strong intelligence-per-dollar extends the open-weights frontier beyond DeepSeek/Qwen, though the Opus-4.8-parity coding claim is lab-reported and unverified.
- +1 Qwen releases Qwen3.8-Flash-Next as an early preview of Qwen4's architecture β Qwen open-sourcing an early Qwen4-architecture preview broadens the open field even before independent benchmarks exist.
- +2 DeepSeek releases V4 open weights under MIT license β Smallest open/closed gap on record
- +1 Mistral raises β¬4B, commits to keeping flagship weights open β β¬4B round tied to an explicit open-weights commitment through 2027.
Confidence history
- 2026-04-01 β 0.6 β Initial
- 2026-08-24 β 0.7 β DeepSeek V4 within 5 AA points