Claude Fable 5
Anthropic's new Mythos-class flagship and the most capable coding model anyone can use — 80.3% SWE-Bench Pro, an 11-point jump over Opus 4.8. 1M context, 128K output, native parallel subagents. Released June 9, 2026.
Best Chinese flagship — beats GPT-5.6 Sol on coding, #2 globally for vision.
Multimodal and vision-heavy workloads at scale
You need independently verified benchmarks or Western data residency.
Compare every model's knowledge cutoff, max output, and context window.
Announced August 3, 2026 on Alibaba Cloud Model Studio; open weights promised a week after launch. $2/$6 is first-party Model Studio pricing; cache reads from $0.17/1M. Announcement moved Alibaba stock +7% in Hong Kong.
SWE-bench Pro 67.7 — ahead of GPT-5.6 Sol and close to Claude Opus 4.8
#2 globally on Arena.AI vision (behind only a Claude Fable 5 variant); #1 Chinese model for text
First Alibaba open-weights release at this scale — 2.4T MoE at $2/$6 per 1M
Well behind Claude Fable 5 on SWE-bench Pro (67.7 vs 80.0) and behind several Anthropic models on text rankings
No independent third-party benchmarks at GA — early claims are largely Alibaba-reported
What people actually use Qwen 3.8 Max for.
Agentic coding — 67.7 SWE-bench Pro, ahead of GPT-5.6 Sol (64.6) and near Claude Opus 4.8 (69.2)
Vision-heavy pipelines: image and video understanding ranked #2 globally on Arena.AI
Large-scale deployments where 95B active params keep inference cost moderate
The nearest models people weigh against it, and what actually separates them.
vs Claude Fable 5 — Against Claude Fable 5 (Anthropic), Qwen 3.8 Max runs about 87% cheaper per token and answers faster. Take Qwen 3.8 Max unless you specifically need what Claude Fable 5 does better.
vs Claude Fable 5.1 — Against Claude Fable 5.1 (Anthropic), Qwen 3.8 Max runs about 87% cheaper per token and answers faster. Take Qwen 3.8 Max unless you specifically need what Claude Fable 5.1 does better.
vs Claude Opus 4.7 — Against Claude Opus 4.7 (Anthropic), Qwen 3.8 Max runs about 73% cheaper per token and answers faster. Take Qwen 3.8 Max unless you specifically need what Claude Opus 4.7 does better.
Price History
→0% since Aug 7
38 data points · tracked daily since Aug 7, 2026
Multimodal and vision-heavy workloads at scale. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Anthropic's new Mythos-class flagship and the most capable coding model anyone can use — 80.3% SWE-Bench Pro, an 11-point jump over Opus 4.8. 1M context, 128K output, native parallel subagents. Released June 9, 2026.
Anthropic's September 1, 2026 frontier release and the new capability ceiling for coding, agents, and scientific work. Base pricing is unchanged at $10/$50, but cache reads dropped 75% to $0.25/1M — roughly 25% cheaper on typical workloads and up to 45% cheaper on agentic ones. 1M context, 128K output, adaptive thinking always on.
Anthropic's previous Opus flagship, now superseded by Opus 4.8. Still the second-best coding model publicly available at the same $5/$25 price.
Qwen 3.8 Max costs $2 per million input tokens and $6 per million output tokens on the API, with cached input at $0.25 per million. A month of 10M input and 2M output tokens runs about $32.00 at list price, before any batch or caching discounts.
Qwen 3.8 Max has a 1M tokens context window, with up to 128k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
Qwen 3.8 Max is best for multimodal and vision-heavy workloads at scale. It is a strong fit when that workflow matters more than the tradeoffs around balanced pricing and balanced speed.
You need independently verified benchmarks or Western data residency.
GPT-5.1-Codex-Max (OpenAI) at $1.25/1M/1M input against Qwen 3.8 Max's $2.00/1M/1M. The strongest choice for serious software engineering work, provided you can absorb the output-side pricing. Compare it first if Qwen 3.8 Max's pricing is the thing stopping you.
Claude Fable 5 — deliberate against Qwen 3.8 Max's balanced, with 1M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.