AI Change Tracker
4

OpenAI launches GPT-6 Astra, priced at Fable parity, with disputed benchmark claims and a bumpy rollout

2026-09-03

OpenAI launched GPT-6 Astra on September 3, 2026, rolling out first to a limited set of organizations and over subsequent days to ChatGPT Plus/Pro/Business/Enterprise, the API, and AWS. It is priced at $10/$50 per million input/output tokens, matching Claude Fable 5 and 5.1. OpenAI describes it as its 'most intelligent and aligned model yet,' positioned around computer/browser use, coding, math/science, office work, and cybersecurity, per OpenAI and TechCrunch. OpenAI and testers reported a 99.9% score on ARC-AGI-3 using a custom 'Provider Adapter harness' costing $19K, versus 62.7% for $26K on the standard harness, per ARC-AGI's own blog as relayed by Simon Willison. Independent measurement firm Artificial Analysis found Astra's Intelligence Index score of 61 — level with GPT-5.6 Sol, five points below Claude Fable 5.1, and behind Meta's Muse Spark 1.3 — while leading on their Coding Agent Index cost-efficiency measure (2 points higher than Sol at equal max-effort cost, and less than half Fable 5's per-task cost at equal score). On security benchmarks OpenAI reported Astra scoring 100% on ExploitBench (vs. Sol's 78.5%), 42.4% on ExploitGym (vs. 30.3%), and 99.2% on SRE-Bench (vs. 68.7%), and 100%/96.3% on OpenAI's own long-context needle benchmark at 256K-512K/512K-1M tokens. OpenAI's system card described both improved alignment and decreased chain-of-thought monitorability, drawing pointed reaction from researchers including Neel Nanda and Ryan Greenblatt, per Latent Space's recap. The rollout itself was bumpy: TechCrunch and Latent Space reported delays, a broken/late blog post, unclear access timing, and user frustration that many influencers had early access while paying customers did not; OpenAI compensated affected paid users with 'banked resets.'

Significance 4: A new frontier-priced flagship from a tracked lab that moves multiple tracked axes (reasoning, coding, agentic, long-context) and leads one independent cost-efficiency index, but independent AA measurement puts it behind two other models on general intelligence and its headline ARC-AGI-3 claim is contested by methodology — short of a level-5 landscape change since no tracked verdict currently flips.

technology business financial operational Release Benchmark result Safety / alignment Lab strategy

Implications · machine-drafted, not owner judgment

OpenAI now has a Fable-tier-priced flagship that leads on agentic-coding cost-efficiency and dominates computer-use/security benchmarks, giving it a credible answer to Anthropic even though independent AA scoring puts its general intelligence behind both Fable 5.1 and Meta's Muse Spark 1.3 — the frontier is fragmenting by axis rather than being led by one model. The disclosed drop in chain-of-thought monitorability, paired with marketing language claiming improved alignment, is exactly the stated-vs-revealed gap this tracker's leader-pattern thesis watches for, and it narrows the transparency window safety researchers rely on just as OpenAI ships its most capable computer-use and cyber-capable model yet. The bumpy, influencer-favoring rollout is an operational execution signal independent of the model's raw capability.

Watch for
  • Independent reproduction of Astra's ARC-AGI-3 score on the standard (non-custom) harness
  • Whether OpenAI publishes a system-card addendum quantifying the chain-of-thought monitorability reduction
  • Whether Astra's AA Coding Agent Index cost-efficiency lead holds once broader third-party agentic evals (Terminal-Bench, SWE-bench Verified) are run

Sources

Affected axes

Hard reasoning · Coding · Agentic / tool use · Long context

← Back to feed