GPT-5.5
OpenAI's latest agentic flagship for coding, research, computer-use workflows, and long multi-step knowledge work.
OpenAI's frontier answer to Fable 5.1 — computer-use and agentic-coding leader at $10/$50.
Computer and browser use, long-horizon agentic coding, and frontier math and science work
You are cost-sensitive or latency-bound — GPT-5.6 Sol is a fifth of the price — or you need it today in an Enterprise workspace, where it is off by default, or for offensive-security work, which it refuses outside OpenAI Daybreak.
Compare every model's knowledge cutoff, max output, and context window.
Released September 3, 2026. API ID gpt-6-astra; rolling out over the coming days to ChatGPT Plus, Pro, Business and Enterprise (usage inside existing allowances; GPT-6 Astra Pro for Pro/Business/Enterprise; Enterprise off by default), the OpenAI API, Microsoft Azure and Amazon Bedrock. Standard API pricing $10/$50 per 1M tokens; Fast mode is up to 2x speed at 2x price; cache reads and writes have separate rates. Model docs list 1,050,000 context, 128,000 max output, knowledge cutoff April 30, 2026, reasoning efforts up to 'max'. Published launch numbers (Astra / GPT-5.6 Sol / Fable 5.1 / Opus 5): OSWorld 2.0 72.6 / 65.7 / — / 70.2; Terminal-Bench 4.0 57.9 / 37.3 / 55.8 / 52.3; Terminal-Bench Science 0.1 64.6 / 22.4 / 52.6 / 30.0; FrontierMath Tier 4 v2 97.6 / 83.0 / 87.8 / 73.2; GPQA Diamond 96.0 / 94.6 / 93.7 / 93.7; Humanity's Last Exam w/ tools 57.2 / — / 65.0 / 63.6; AutomationBench 41.4 / 18.1 / 31.4 / 26.9; DeepSWE v1.1 74.1 / 72.7 / 67.4 / 73.7; ARC-AGI-2 95.0 / 92.5 / 90.0 / 90.4; ARC-AGI-3 99.9 (OpenAI responses-API harness; ARC Prize's stateless runs score far lower) / 7.8 / — / 30.2; ExploitBench 100.0 / 78.5 / — / 70; SRE-Bench 88.0 / 55.9; Artificial Analysis Intelligence Index v4.1.1 61.2 / 60.9 / 65.7 / 63.1. Meets the Critical threshold for cybersecurity under OpenAI's Preparedness Framework; advanced cyber workflows gated behind OpenAI Daybreak. All figures from OpenAI's launch post and model docs, verified September 4, 2026.
OSWorld 2.0 72.6% vs 65.7% for GPT-5.6 Sol, in roughly 47% less time per task — the new computer-use ceiling
Terminal-Bench 4.0 57.9% — ahead of Claude Fable 5.1 (55.8%), Opus 5 (52.3%) and GPT-5.6 Sol (37.3%)
FrontierMath Tier 4 (v2) 97.6% vs Fable 5.1's 87.8%; Terminal-Bench Science 64.6% vs 52.6%; GPQA Diamond 96.0%
1.05M context with 96.3% on OpenAI MRCR 8-needle at 512K–1M (Sol: 73.8%) and 128K max output
OpenAI's lowest misaligned-outcome rates to date: 2.4% on its computer-use safety benchmark vs 22.0% for Sol, 0% scope-creep on impossible cyber tasks vs 48%
$10/$50 per 1M — five times GPT-5.6 Sol's $2/$10 and the same premium as Claude Fable 5.1; Fast mode doubles it again
Trails Claude Fable 5.1 on Humanity's Last Exam with tools (57.2% vs 65.0%) and on the Artificial Analysis Intelligence Index (61.2 vs 65.7)
OpenAI published no SWE-bench Verified or SWE-bench Pro figure at launch, so it does not appear on our SWE-bench leaderboard
Rolling out over days, not instantly: Enterprise access is off by default, advanced cyber tasks are refused outside OpenAI Daybreak, and the safety layer can pause or stop legitimate agent runs
OpenAI's own system card finds its written reasoning harder to monitor than Sol's
What people actually use GPT-6 Astra for.
Computer-use agents that fill forms, update CRMs, run QA in a browser — 72.6% on OSWorld 2.0 at ~40 minutes per task vs Sol's 75
Agentic coding in Codex with cross-context notes — 57.9% Terminal-Bench 4.0, ahead of Claude Fable 5.1 (55.8%)
Frontier math and science — 97.6% FrontierMath Tier 4 and 64.6% Terminal-Bench Science, both well clear of every rival OpenAI tested
Computer and browser use, long-horizon agentic coding, and frontier math and science work. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
OpenAI's latest agentic flagship for coding, research, computer-use workflows, and long multi-step knowledge work.
Anthropic's September 1, 2026 frontier release and the new capability ceiling for coding, agents, and scientific work. Base pricing is unchanged at $10/$50, but cache reads dropped 75% to $0.25/1M — roughly 25% cheaper on typical workloads and up to 45% cheaper on agentic ones. 1M context, 128K output, adaptive thinking always on.
Anthropic's new Mythos-class flagship and the most capable coding model anyone can use — 80.3% SWE-Bench Pro, an 11-point jump over Opus 4.8. 1M context, 128K output, native parallel subagents. Released June 9, 2026.
Pricing moves, ranking shifts, and capability updates.
OpenAI released GPT-6 Astra on September 3, 2026, two days after Anthropic's Claude Fable 5.1, calling it its most intelligent and aligned model. The headline is computer use: 72.6% on OSWorld 2.0 in roughly 47% less time per task than GPT-5.6 Sol (65.7%), plus 92.7% on ScreenSpot-Pro. On agentic coding it posts 57.9% on Terminal-Bench 4.0, ahead of Fable 5.1 (55.8%), Opus 5 (52.3%) and Sol (37.3%), and on frontier math it saturates FrontierMath Tier 4 at 97.6% against Fable 5.1's 87.8%. It is not a clean sweep: Fable 5.1 still leads Humanity's Last Exam with tools (65.0% vs 57.2%) and the Artificial Analysis Intelligence Index (65.7 vs 61.2), and OpenAI published no SWE-bench figure. Standard API pricing is $10 per million input tokens and $50 per million output — five times GPT-5.6 Sol and identical to Fable 5.1 — with a Fast mode at 2x speed for 2x price. Model docs list a 1,050,000-token context, 128,000 max output and an April 30, 2026 knowledge cutoff. Availability: rolling out over the coming days to ChatGPT Plus, Pro, Business and Enterprise (inside existing allowances, Enterprise off by default, a GPT-6 Astra Pro tier for Pro/Business/Enterprise), the OpenAI API as gpt-6-astra, Microsoft Azure and Amazon Bedrock. Astra meets the Critical cybersecurity threshold under OpenAI's Preparedness Framework — 100% on ExploitBench — so advanced cyber tasks are refused outside the OpenAI Daybreak program, and OpenAI notes its written reasoning is harder to monitor than Sol's. Every figure is from OpenAI's launch post and model documentation, verified September 4, 2026.
View modelGPT-6 Astra costs $10 per million input tokens and $50 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $200.00 at list price, before any batch or caching discounts.
GPT-6 Astra has a 1.1M tokens context window, with up to 128k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
GPT-6 Astra's training data runs through April 30, 2026, and the model was released on September 3, 2026. For anything after that date it needs web search or documents in the prompt.
GPT-6 Astra is best for computer and browser use, long-horizon agentic coding, and frontier math and science work. It is a strong fit when that workflow matters more than the tradeoffs around premium pricing and deliberate speed.
You are cost-sensitive or latency-bound — GPT-5.6 Sol is a fifth of the price — or you need it today in an Enterprise workspace, where it is off by default, or for offensive-security work, which it refuses outside OpenAI Daybreak.
Grok 4.5 is the lower-cost option to compare first when you want a similar workflow fit with less token spend.
GPT-5.5 is the better pick when response time matters more than maximum depth or premium quality.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.