OpenAI launches GPT-6 Astra, priced at Fable parity, with disputed benchmark claims and a bumpy rollout
2026-09-03
OpenAI launched GPT-6 Astra on September 3, 2026, rolling out first to a limited set of organizations and over subsequent days to ChatGPT Plus/Pro/Business/Enterprise, the API, and AWS. It is priced at $10/$50 per million input/output tokens, matching Claude Fable 5 and 5.1. OpenAI describes it as its 'most intelligent and aligned model yet,' positioned around computer/browser use, coding, math/science, office work, and cybersecurity, per OpenAI and TechCrunch. OpenAI and testers reported a 99.9% score on ARC-AGI-3 using a custom 'Provider Adapter harness' costing $19K, versus 62.7% for $26K on the standard harness, per ARC-AGI's own blog as relayed by Simon Willison. Independent measurement firm Artificial Analysis found Astra's Intelligence Index score of 61 — level with GPT-5.6 Sol, five points below Claude Fable 5.1, and behind Meta's Muse Spark 1.3 — while leading on their Coding Agent Index cost-efficiency measure (2 points higher than Sol at equal max-effort cost, and less than half Fable 5's per-task cost at equal score). On security benchmarks OpenAI reported Astra scoring 100% on ExploitBench (vs. Sol's 78.5%), 42.4% on ExploitGym (vs. 30.3%), and 99.2% on SRE-Bench (vs. 68.7%), and 100%/96.3% on OpenAI's own long-context needle benchmark at 256K-512K/512K-1M tokens. OpenAI's system card described both improved alignment and decreased chain-of-thought monitorability, drawing pointed reaction from researchers including Neel Nanda and Ryan Greenblatt, per Latent Space's recap. The rollout itself was bumpy: TechCrunch and Latent Space reported delays, a broken/late blog post, unclear access timing, and user frustration that many influencers had early access while paying customers did not; OpenAI compensated affected paid users with 'banked resets.'
Significance 4: A new frontier-priced flagship from a tracked lab that moves multiple tracked axes (reasoning, coding, agentic, long-context) and leads one independent cost-efficiency index, but independent AA measurement puts it behind two other models on general intelligence and its headline ARC-AGI-3 claim is contested by methodology — short of a level-5 landscape change since no tracked verdict currently flips.
Implications · machine-drafted, not owner judgment
OpenAI now has a Fable-tier-priced flagship that leads on agentic-coding cost-efficiency and dominates computer-use/security benchmarks, giving it a credible answer to Anthropic even though independent AA scoring puts its general intelligence behind both Fable 5.1 and Meta's Muse Spark 1.3 — the frontier is fragmenting by axis rather than being led by one model. The disclosed drop in chain-of-thought monitorability, paired with marketing language claiming improved alignment, is exactly the stated-vs-revealed gap this tracker's leader-pattern thesis watches for, and it narrows the transparency window safety researchers rely on just as OpenAI ships its most capable computer-use and cyber-capable model yet. The bumpy, influencer-favoring rollout is an operational execution signal independent of the model's raw capability.
- Independent reproduction of Astra's ARC-AGI-3 score on the standard (non-custom) harness
- Whether OpenAI publishes a system-card addendum quantifying the chain-of-thought monitorability reduction
- Whether Astra's AA Coding Agent Index cost-efficiency lead holds once broader third-party agentic evals (Terminal-Bench, SWE-bench Verified) are run
Sources
- curator GPT-6 Astra retrieved 2026-09-04
- press OpenAI launches Astra, its powerful (and controversial) new model retrieved 2026-09-04
- curator [AINews] GPT-6 Astra: OpenAI's biggest LLM launch of all time retrieved 2026-09-04
- curator GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour retrieved 2026-09-04
- curator An Interview with OpenAI President Greg Brockman About Astra and Alignment retrieved 2026-09-04
- curator Google gets away with it (PLUS: GPT-6 Astra and 'the AGI era') retrieved 2026-09-04
Affected axes
Hard reasoning · Coding · Agentic / tool use · Long context