Anthropic's September 1, 2026 release puts real distance between Claude Fable 5.1 and OpenAI's flagship on the agentic benchmarks both labs report. Fable 5.1 scores 55.8% on Terminal-Bench 4.0 against GPT-5.6 Sol's 37.3%, 52.6% on Terminal-Bench-Science 0.1 against 22.4%, 73.4% on CursorBench 3.2.0 against 67.2%, and 1853 on GDPval-AA v2 against 1711. GPT-5.6 Sol is cheaper on input ($5 against $10) and leads on its own strengths — 94.6% GPQA Diamond science reasoning, 90.4% BrowseComp agentic browsing, and ultra mode's sub-agent orchestration. The honest split: pick Sol for research browsing and OpenAI-native tooling, pick Fable 5.1 for long autonomous engineering and scientific agent work.
AnthropicPremium
Claude Fable 5.1
New frontier leader — better than Fable 5 on every published benchmark, and cheaper to run.
Winner
VS
OpenAIPremium
GPT-5.6 Sol
Best OpenAI flagship — leads terminal coding and agentic browsing.
At a glance
Claude Fable 5.1
GPT-5.6 Sol
Input cost / 1M tokens
$$10.00/1M
$$5.00/1M
Output cost / 1M tokens
$$50.00/1M
$$30.00/1M
Context window
1M tokens
1.1M tokens
Speed
Deliberate
Deliberate
Price tier
Premium
Premium
Benchmarks
SWE-bench (coding)
—
96.2%
Arena Elo
—
—
MMLU
—
—
How they compare
Which model wins for each use case — and why.
Agentic codingClaude Fable 5.1 wins
Terminal-Bench 4.0: 55.8% for Fable 5.1 against 37.3% for GPT-5.6 Sol — an 18.5-point gap on the benchmark both labs quote.
Agentic scienceClaude Fable 5.1 wins
Terminal-Bench-Science 0.1: 52.6% against 22.4%, more than double.
Knowledge workClaude Fable 5.1 wins
GDPval-AA v2 scores Fable 5.1 at 1853 against Sol's 1711.
Science Q&A and browsingGPT-5.6 Sol wins
GPT-5.6 Sol posts 94.6% on GPQA Diamond and 90.4% on BrowseComp — the strongest published agentic-browsing result.
Input priceGPT-5.6 Sol wins
Sol is $5/$30 against Fable 5.1's $10/$50, though Sol adds a long-context surcharge ($10/$45) above 272K tokens and ultra mode costs 2–3x.
Ecosystem fitGPT-5.6 Sol wins
If your stack is Codex, ChatGPT rollout, or OpenAI-native computer use, Sol removes an integration boundary that benchmarks do not capture.
Which should you pick?
Pick Claude Fable 5.1 if…
Long autonomous coding and research agents are the workload — the gap is 18–30 points there
You want the strongest published knowledge-work score (GDPval-AA v2 1853)
Your agent loops replay large cached prompts, where $0.25/1M cache reads change the economics
You are free to pick the best model rather than the one your tooling assumes
For most workflows, Claude Fable 5.1 is the stronger choice.
The new frontier leader, and the rare upgrade that costs less than the model it replaces. Base $10/$50 pricing did not move, but a 75% cut to cache reads makes typical workloads ~25% cheaper than Fable 5 and heavily agentic ones ~45% cheaper — while more than doubling Fable 5 on agentic science and beating Opus 5 on agentic coding. If you were already on Fable 5, switching is a straight win. If you were on Opus 5 for the price, it still costs 2x on uncached tokens.
The case for each model
What each one is genuinely good at, where it falls down, and when we would steer you away from it.
Anthropic's September 1, 2026 frontier release and the new capability ceiling for coding, agents, and scientific work. Base pricing is unchanged at $10/$50, but cache reads dropped 75% to $0.25/1M — roughly 25% cheaper on typical workloads and up to 45% cheaper on agentic ones. 1M context, 128K output, adaptive thinking always on.
Input
$10.00/1M
Output
$50.00/1M
Context
1M tokens
Speed
Deliberate
What people actually use it for
Agentic scientific research — 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%
Long-running autonomous coding agents that plan, run, and debug across a whole repository
Cache-heavy agent loops where the 75% cache-read cut ($1.00 → $0.25 per 1M) is the real saving
Where it wins
52.6% Terminal-Bench-Science 0.1 — 2.1x Fable 5 (24.7%), well clear of Opus 5 (29.0%) and GPT-5.6 Sol (22.4%)
55.8% Terminal-Bench 4.0 agentic coding, ahead of Opus 5 (52.3%) and Fable 5 (42.0%)
Cache reads cut 75% to $0.25/1M — ~25% cheaper for typical use, ~45% for heavily agentic work
Independent Vals AI evaluation ranks it #1 of 51 on the Vals Index, #1 on LiveCodeBench (90.5%) and MMLU Pro (92.4%)
1M-token context and 128K max output at standard rates, with adaptive thinking always enabled
Where it falls down
Base rates are still $10/$50 per 1M — double Claude Opus 5 for anything that is not cache-heavy
Anthropic published no SWE-bench Verified or SWE-bench Pro figure for 5.1 at launch
On Claude Pro it only runs on pay-as-you-go credits; Max includes it up to 50% of weekly limits
Deliberate, high-latency profile — the wrong choice for interactive, latency-bound apps
Skip it if
You need low latency, or your workload is uncached and cost-sensitive — Claude Opus 5 is half the base price and within a few points on agentic coding.
Our verdict
The new frontier leader, and the rare upgrade that costs less than the model it replaces. Base $10/$50 pricing did not move, but a 75% cut to cache reads makes typical workloads ~25% cheaper than Fable 5 and heavily agentic ones ~45% cheaper — while more than doubling Fable 5 on agentic science and beating Opus 5 on agentic coding. If you were already on Fable 5, switching is a straight win. If you were on Opus 5 for the price, it still costs 2x on uncached tokens.
Released September 1, 2026 alongside Claude Mythos 5.1, the first update to the Mythos-class line since Fable 5 on June 9. API ID claude-fable-5-1; generally available on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. Published launch numbers (Fable 5.1 / Fable 5 / Opus 5 / GPT-5.6 Sol): Terminal-Bench-Science 0.1 52.6 / 24.7 / 29.0 / 22.4; Terminal-Bench 4.0 55.8 / 42.0 / 52.3 / 37.3; CursorBench 3.2.0 73.4 / 70.5 / 70.0 / 67.2; AutomationBench 31.4 / 17.1 / 26.9 / 19.6; OSWorld 2.0 strict 41.7 / 36.1 / 39.6; Humanity's Last Exam (no tools) 60.9 / 57.8 / 56.6; GDPval-AA v2 1853 / 1723 / 1824 / 1711. GDPval-AA v2 is rescaled from the v1 numbers quoted on the Fable 5 page and is not directly comparable to them.
The flagship of OpenAI's GPT-5.6 family — its most capable reasoning and agentic-coding model, with an 'ultra' mode that spawns sub-agents for long autonomous workflows.
Input
$5.00/1M
Output
$30.00/1M
Context
1.1M tokens
Speed
Deliberate
What people actually use it for
Tool-heavy terminal coding — 88.8% on Terminal-Bench 2.1 (91.9% in ultra mode, a state of the art)
Agentic web research at 90.4% BrowseComp with top-tier GPQA Diamond science reasoning (94.6%)
Long autonomous workflows using ultra mode's sub-agent orchestration
Where it wins
Terminal-Bench 2.1 leader at 88.8% (91.9% ultra) — the top OpenAI agentic-coding result
94.6% GPQA Diamond and 90.4% BrowseComp — frontier science reasoning and agentic browsing
Artificial Analysis Coding Agent Index leader at 80 points, near Fable 5 intelligence at roughly one-third the cost
Where it falls down
Trails Claude Opus 5 badly on repository-level engineering (SWE-bench Pro 64.6% vs 79.2%)
Long-context surcharge ($10/$45 above 272K) and 2–3x ultra-mode costs stack up fast
Skip it if
Repo-level coding is the main job — Opus 5 leads SWE-bench Pro by ~15 points — or you're cost-sensitive (Terra is 60% cheaper at 1–4 points off).
Our verdict
OpenAI's strongest model and the terminal-workflow leader. Sol beats everything on Terminal-Bench and agentic browsing, but Claude Opus 5 remains the better pick for repository-level software engineering.
First frontier model family to clear a customer-by-customer US government review: limited preview June 26, full public release July 9, 2026. Pricing $5/$30 ($10/$45 above 272K context). Knowledge cutoff Feb 16, 2026.
Frequently asked questions
Is Claude Fable 5.1 better than GPT-5.6 Sol?
On the agentic benchmarks both labs report, yes and by a wide margin: Terminal-Bench 4.0 55.8% against 37.3%, Terminal-Bench-Science 0.1 52.6% against 22.4%, CursorBench 3.2.0 73.4% against 67.2%, GDPval-AA v2 1853 against 1711. GPT-5.6 Sol leads on GPQA Diamond science Q&A (94.6%) and BrowseComp agentic browsing (90.4%).
Which is cheaper, Fable 5.1 or GPT-5.6 Sol?
Sol is cheaper on paper at $5/$30 against $10/$50, but it adds a long-context surcharge of $10/$45 above 272K tokens and ultra mode costs 2–3x. Fable 5.1 reads cache at $0.25 per 1M, so cache-heavy agent workloads can close or reverse the gap.
Which model is better for autonomous agents?
Fable 5.1, on the published evidence. It leads Terminal-Bench 4.0 by 18.5 points and AutomationBench by 11.8 (31.4% against 19.6%), the two benchmarks that most directly measure long multi-step workflows completing without intervention.
Do they have the same context window?
Close. Fable 5.1 offers 1M tokens at standard rates; GPT-5.6 Sol offers about 1.05M but charges a higher rate ($10/$45) above 272K tokens, so long-context work is meaningfully more expensive on Sol.
Can I use both?
Many teams do, and it is usually the cheapest answer. Route agentic engineering and scientific research to Fable 5.1, and web research or OpenAI-native computer-use tasks to Sol.