UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

JUST ANNOUNCED · JUN 26GPT-5.6 — OpenAI's new Sol, Terra & Luna modelsLimited preview; verified benchmarks land here as OpenAI publishes them.FIRST LOOK →
UPDATED 2026-08-2739 PUBLIC RELEASES YTDSCORED ON 5 BENCHMARKS0 PAID RANKINGS

40 new flagship models.
One year-defining release.

Every major model that shipped in 2026 — ranked, benchmarked, and dated. Scrub the timeline to see how the field reshaped itself month by month.

The 2026 release timeline.

JAN 01 ─── DEC 31
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
NOV
DEC
Claude Haiku 4.5
DeepSeek V4
Llama 4 Scout
Llama 4 Maverick
Grok 4
Gemini 3.1 Pro
GPT-5.4
Llama 5
Muse Spark
Claude Opus 4.7
Mistral Medium 3.5
GPT-5.5
Mistral Medium 3.1
Gemini 3.5 Flash
Qwen 3.7 Max
Claude Opus 4.8
Claude Fable 5
Kimi K2.7 Code
GLM-5.2
GPT-5.6 Sol
GPT-5.6 Terra
Claude Sonnet 5
Grok 4.5
GPT-5.6 Luna
Kimi K3
DeepSeek V4-Pro
Gemini 3.6 Flash
Gemini 3.5 Flash-Lite
Claude Opus 5
DeepSeek V4-Flash
Qwen 3.8 Max
Muse Glimmer 30B
Grok 4.6
Gemini 3.7 Flash
GLM-5.3
GLM-5.3 Flash
Qwen 3.8 Flash
Claude Fable 5.1
GPT-6 Astra
2026-07-2450 DAYS AGO
Claude Opus 5
ANTHROPIC · PREMIUM
The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.
SWE-BENCH
79.2
CONTEXT
1M
$/M IN
$5
READ FULL REPORT →
01 / 06

The three picks that shaped the year.

EDITOR'S VERDICT
BEST OVERALL · 2026
Claude Opus 5
ANTHROPIC · PREMIUM

SWE-bench Verified leader (~96%) and Terminal-Bench 89.1% at Opus 4.8's exact $5/$25 price — within half a point of Fable 5 on agentic coding at half the cost. Fable 5 keeps the absolute ceiling; Opus 5 is the one to actually ship on.

RELEASED
07-24
$/M IN
$5
CONTEXT
1M
BEST VALUE
GPT-5.6 Luna
OPENAI · BUDGET

After the ~80% July price cut: scores within 2–4 points of GPT-5.6 Sol at $0.20/M input. The budget disruptor of the year.

RELEASED
07-09
$/M IN
$0.2
CONTEXT
1M
BEST OPEN-WEIGHTS
DeepSeek V4-Pro
DEEPSEEK · OPEN-WEIGHTS

80.6% SWE-bench Verified (self-reported), MIT license, 1M context, $0.87/M output — an order of magnitude cheaper than the closed frontier.

RELEASED
07-20
$/M IN
$0.44
CONTEXT
1M
02 / 06

Benchmark battle. Pick the lens.

5 BENCHMARKS · 32 MODELS
01Claude Fable 5
80.3$10/M
02Claude Opus 5
79.2$5/M
03Kimi K3
70.0$3/M
04Claude Opus 4.8
69.2$5/M
05Qwen 3.8 Max
67.7$2/M
06Claude Sonnet 5
66.0$3/M
07DeepSeek V4-Pro
65.0$0.44/M
08Grok 4.5
64.7$2/M
09GPT-5.6 Sol
64.6$5/M
10Claude Opus 4.7
64.3$5/M
11GPT-5.6 Terra
63.4$2/M
12GPT-5.6 Luna
62.7$0.2/M
13Qwen 3.8 Flash
62.5$0.16/M
14GLM-5.2
62.1$1.4/M
15Llama 5
62.0—
16Qwen 3.7 Max
60.6$2.5/M
17GPT-5.5
60.1$5/M
18Muse Spark
60.0$1.25/M
19Kimi K2.7 Code
60.0$0.95/M
20Mistral Medium 3.5
58.0$1.5/M
21DeepSeek V4-Flash
58.0$0.14/M
22GPT-5.4
57.7$2.5/M
23Gemini 3.6 Flash
57.0$1.5/M
24Grok 4
56.0$2/M
25Gemini 3.5 Flash
55.1$1.5/M
26Gemini 3.5 Flash-Lite
54.2$0.3/M
27Gemini 3.1 Pro
53.2$1.25/M
28DeepSeek V4
52.8$0.27/M
29Mistral Medium 3.1
49.0$1/M
30Llama 4 Maverick
47.2$0.6/M
31Llama 4 Scout
41.0$0.3/M
32Claude Haiku 4.5
28.4$0.8/M
03 / 06

Quality per dollar. The full chart.

SCATTER · 38 MODELS
Claude Haiku 4.5
DeepSeek V4
Llama 4 Scout
Llama 4 Maverick
Grok 4
Gemini 3.1 Pro
GPT-5.4
Claude Opus 4.7
GPT-5.5
Mistral Medium 3.1
Claude Opus 4.8
Claude Fable 5.1
GPT-6 Astra
Claude Fable 5
Muse Spark
Mistral Medium 3.5
Gemini 3.5 Flash
Qwen 3.7 Max
Kimi K2.7 Code
GLM-5.2
GPT-5.6 Sol
GPT-5.6 Terra
Claude Sonnet 5
Grok 4.5
GPT-5.6 Luna
Kimi K3
DeepSeek V4-Pro
Gemini 3.6 Flash
Gemini 3.5 Flash-Lite
Claude Opus 5
DeepSeek V4-Flash
Qwen 3.8 Max
Muse Glimmer 30B
Grok 4.6
Gemini 3.7 Flash
GLM-5.3
GLM-5.3 Flash
Qwen 3.8 Flash
$/M INPUT →

The Pareto frontier of 2026.

Each dot is one model. Up means better quality, right means more expensive. The top-left edge is the Pareto frontier — every dot under it is strictly dominated.

ON THE FRONTIER
Claude Fable 5 (top quality) · Claude Opus 4.8 (best value at the top) · Grok 4 (best value) · Llama 4 Maverick (best free option)
04 / 06

Who shipped what. Year-to-date.

10 PROVIDERS · 37 RELEASES
Anthropic
6
RELEASES · YTD
OpenAI
5
RELEASES · YTD
Google
5
RELEASES · YTD
Meta
5
RELEASES · YTD
xAI
3
RELEASES · YTD
DeepSeek
3
RELEASES · YTD
Alibaba
3
RELEASES · YTD
Z.ai
3
RELEASES · YTD
Mistral
2
RELEASES · YTD
Moonshot
2
RELEASES · YTD
05 / 06

Every release. One card each.

43 CARDS · NEWEST FIRST
MODEL OF THE YEAR
2026-09-039D AGO · RELEASE
GPT-6 Astra
OPENAIPREMIUM

OpenAI's first GPT-6 model and its answer to Fable 5.1, two days later and at the same $10/$50. The new ceiling on computer use (72.6% OSWorld 2.0, ~47% less time per task than Sol), agentic coding (57.9% Terminal-Bench 4.0 vs Fable 5.1's 55.8%) and frontier math (97.6% FrontierMath Tier 4). Fable 5.1 keeps Humanity's Last Exam (65.0% vs 57.2%). No SWE-bench figure published.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$10.00/M
Output cost$50.00/M
Context1.05M
+ PROS
  • +72.6% OSWorld 2.0 — the computer-use ceiling, in ~47% less time than Sol
  • +57.9% Terminal-Bench 4.0, ahead of Fable 5.1 (55.8%) and Opus 5 (52.3%)
  • +97.6% FrontierMath Tier 4 and 96.0% GPQA Diamond
  • +1.05M context, 128K output, April 30 2026 cutoff
– CONS
  • –$10/$50 — five times GPT-5.6 Sol
  • –Trails Fable 5.1 on Humanity's Last Exam and the AA Intelligence Index
  • –No published SWE-bench score; Enterprise access off by default
VIEW FULL REPORT →
MODEL OF THE YEAR
2026-09-0111D AGO · RELEASE
Claude Fable 5.1
ANTHROPICPREMIUM

The new frontier ceiling, and the rare upgrade that costs less than what it replaces. 52.6% on Terminal-Bench-Science 0.1 against Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0 against Opus 5's 52.3%. Base pricing is unchanged at $10/$50, but cache reads fell 75% to $0.25/1M — about 25% cheaper on typical workloads and up to 45% on agentic ones. Anthropic published no SWE-bench figure at launch.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
  • +52.6% Terminal-Bench-Science 0.1 — 2.1x Fable 5's 24.7%
  • +55.8% Terminal-Bench 4.0, ahead of Opus 5 (52.3%)
  • +Cache reads cut 75% to $0.25/1M — cheaper than Fable 5 to run
  • +Vals AI ranks it #1 of 51 overall and #1 on LiveCodeBench (90.5%)
– CONS
  • –$10/$50 base — still double Claude Opus 5
  • –No published SWE-bench Verified or Pro score at launch
VIEW FULL REPORT →
2026-09-0111D AGO · RESTRICTED
Claude Mythos 5.1
ANTHROPICFRONTIERNOT AVAILABLE

Fable 5.1 with the safeguards lifted — the highest agentic-coding score Anthropic has published (60.9% Terminal-Bench 4.0), restricted to vetted US cybersecurity and life-sciences organisations. Identical price, context and cache discount to Fable 5.1.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost—
Output cost—
Context1M
+ PROS
  • +60.9% Terminal-Bench 4.0 — Anthropic's highest published figure
  • +Same $10/$50 pricing and $0.25/1M cache reads as Fable 5.1
  • +Safeguards lifted for advanced security and life-sciences research
– CONS
  • –Not generally available — vetted trusted access only
  • –No advantage over Fable 5.1 for ordinary work
VIEW FULL REPORT →
2026-08-2617D AGO · RELEASE
GLM-5.3 Flash
Z.AIOPEN-WEIGHTS

The first natively multimodal GLM-5 — a 320B-A18B MoE with vision and video built in, MIT weights, a 1M context, and a fifteen-cent price. The cheapest serious multimodal model on the market.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$0.15/M
Output cost$0.50/M
Context1M
+ PROS
  • +$0.15/$0.50 — halved again through Sep 9 launch promo
  • +Native vision + video, not a bolt-on
  • +MIT weights, 18B active params, 1M context
– CONS
  • –No published SWE-bench score
  • –Volume model, not a ceiling model — frontier reasoning is elsewhere
VIEW FULL REPORT →
2026-08-2617D AGO · RELEASE
Qwen 3.8 Flash
ALIBABABUDGET

Alibaba's Qwen4 architecture preview: a 125B MoE with just 6B active per token that posts SWE-bench Pro 62.5 and GPQA Diamond 91.7 — roughly GLM-5.2-class coding at a tenth of the running cost.

SWE
63
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$0.16/M
Output cost$0.47/M
Context991K
+ PROS
  • +SWE-bench Pro 62.5 at $0.16/M input
  • +6B active params — cheap to host, very fast
  • +Open-weight Qwen3.8-Flash-Next variant previews the Qwen4 design
– CONS
  • –Architecture preview, not the flagship — that's still Qwen 3.8 Max
  • –Self-reported model-card numbers, no independent evals yet
VIEW FULL REPORT →
2026-08-1429D AGO · RELEASE
GLM-5.3
Z.AIOPEN-WEIGHTS

The clear GLM-5.2 upgrade at the exact same price — huge agentic gains (DeepSWE 46.2 → 66.9, Terminal-Bench 3.0 4.6 → 28.3) plus a first: 84.5% on CyberGym, narrowly ahead of Claude Mythos 5's 83.8% on security.

SWE
—
TERM
88
MMLU
—
VIS
—
MATH
—
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
  • +Beats Claude Mythos 5 on CyberGym (84.5 vs 83.8)
  • +Same $1.40/$4.40 price as GLM-5.2
  • +Terminal-Bench 2.1 88.2, GPQA Diamond 91.7
– CONS
  • –No published SWE-bench Verified score
  • –Access initially routed through the GLM Coding Plan ($18/mo)
VIEW FULL REPORT →
2026-08-1330D AGO · RELEASE
Gemini 3.7 Flash
GOOGLEBALANCED

Google's best coding score per dollar — 80.8% SWE-bench Verified and GPQA Diamond 94.8% at half the price of 3.6 Flash, shipped three weeks after it. The catch: the price doubles to $1.50/$7.50 on January 1, 2027.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$0.75/M
Output cost$3.75/M
Context1M
+ PROS
  • +80.8% SWE-bench Verified at $0.75/M input
  • +GPQA Diamond 94.8% — expert reasoning near the frontier
  • +Big SWE jumps over 3.6 Flash: DeepSWE 49.0 → 65.3
– CONS
  • –Introductory price expires Dec 31, 2026 — then doubles
  • –No SWE-bench Pro or Terminal-Bench figures published yet
  • –Makes Google's own 3.6 Flash look overpriced 3 weeks after launch
VIEW FULL REPORT →
2026-08-1231D AGO · RELEASE
Grok 4.6
XAIBALANCED

xAI's long-horizon agent play: finishes agent tasks in roughly half the turns of rivals, so real cost per completed task lands well under the sticker price. AA Intelligence Index 61 — 4th overall.

SWE
—
TERM
88
MMLU
—
VIS
—
MATH
—
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
  • +~2x turn efficiency on long agent runs
  • +Terminal-Bench 2.1 88.4 — ahead of GLM-5.3
  • +Same $2/$6 price as Grok 4.5
– CONS
  • –No published SWE-bench score of any kind
  • –Price doubles to $4/$12 for prompts ≥200K
  • –GPT-5.6 Sol Max beats it on DeepSWE and Terminal-Bench 3.0
VIEW FULL REPORT →
2026-08-0934D AGO · RELEASE
Muse Glimmer 30B
METAOPEN-WEIGHTS

Meta's return to genuine open source — the first model from Meta Superintelligence Labs, a 30B dense Apache 2.0 release built for always-on agents. 76.0% SWE-bench Verified from a model that runs on a 24GB GPU.

SWE
—
TERM
—
MMLU
—
VIS
—
MATH
—
Input cost$0.35/M
Output cost$1.50/M
Context131K
+ PROS
  • +Apache 2.0 — no commercial restrictions at all
  • +76.0% SWE-bench Verified at 30B dense
  • +Runs locally: 24GB GPU with quantization
– CONS
  • –Qwen3.6-27B beats it on several agent tests in Meta's own table
  • –131K context is small next to the 1M field
  • –No first-party API — third-party hosting or self-host only
VIEW FULL REPORT →
2026-08-0340D AGO · RELEASE
Qwen 3.8 Max
ALIBABABALANCED

Alibaba's 2.4T MoE flagship — a legitimate SWE-bench Pro upset over GPT-5.6 Sol (67.7 vs 64.6) and #2 globally for vision. The strongest Chinese multimodal model yet.

SWE
68
TERM
76
MMLU
90
VIS
92
MATH
91
Input cost$2.00/M
Output cost$6.00/M
Context1M
+ PROS
  • +Beats GPT-5.6 Sol on SWE-bench Pro
  • +#2 globally on vision, behind only a Fable 5 variant
  • +First Alibaba open-weights release at this scale
– CONS
  • –No independent benchmarks at GA yet
  • –Trails Fable 5 badly on the coding ceiling
VIEW FULL REPORT →
2026-07-3143D AGO · RELEASE
DeepSeek V4-Flash
DEEPSEEKOPEN-WEIGHTS

The best cheap agent engine of 2026: Terminal-Bench 82.7 at $0.14/M input, with MIT weights self-hostable in ~110 GB. Nothing touches its agentic capability per dollar.

SWE
58
TERM
83
MMLU
86
VIS
38
MATH
90
Input cost$0.14/M
Output cost$0.28/M
Context1M
+ PROS
  • +Terminal-Bench 82.7 rivals models 30x its price
  • +Beats the V4-Pro preview on all nine agent benchmarks
  • +2,500 concurrent requests, 1M context
– CONS
  • –Text-only
  • –Several headline numbers from DeepSeek's own evals
VIEW FULL REPORT →
2026-07-2450D AGO · RELEASE
Claude Opus 5
ANTHROPICPREMIUM

The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.

SWE
79
TERM
89
MMLU
93
VIS
88
MATH
96
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +SWE-bench Verified leader (~96%) as of August 2026
  • +Terminal-Bench 2.1 89.1% — edges GPT-5.6 Sol
  • +Within 0.5% of Fable 5 on CursorBench at half the cost
– CONS
  • –Fable 5 still holds the absolute frontier ceiling
  • –Fast mode doubles pricing to $10/$50
VIEW FULL REPORT →
2026-07-2153D AGO · RELEASE
Gemini 3.6 Flash
GOOGLEBALANCED

The best Google model for agents: beats 3.5 Flash on every tested benchmark with ~17% fewer output tokens and cheaper output. The default Gemini pick until Gemini 4 lands.

SWE
57
TERM
80
MMLU
89
VIS
88
MATH
90
Input cost$1.50/M
Output cost$7.50/M
Context1M
+ PROS
  • +OSWorld-Verified 83.0% — best computer use in class
  • +~17% fewer output tokens + lower price compound
  • +~280–304 tokens/sec
– CONS
  • –Point release, not a generational leap
  • –No Pro tier in this generation
VIEW FULL REPORT →
2026-07-2153D AGO · RELEASE
Gemini 3.5 Flash-Lite
GOOGLEBUDGET

The fastest model in Google's lineup at 350 tokens/sec, with real agentic chops at $0.30/M. GPT-5.6 Luna beats it on price-per-benchmark; Flash-Lite answers with speed and full multimodal input.

SWE
54
TERM
54
MMLU
84
VIS
82
MATH
84
Input cost$0.30/M
Output cost$2.50/M
Context1M
+ PROS
  • +350 tokens/sec — fastest in the 3.5 line
  • +Huge jump over 3.1 Flash-Lite
  • +1M context with video/audio/PDF input
– CONS
  • –Trails full Flash on hard agentic work
  • –Luna wins pure price-per-benchmark math
VIEW FULL REPORT →
2026-07-2153D AGO · RESTRICTED
Gemini 3.5 Flash Cyber
GOOGLERESTRICTEDNOT AVAILABLE

A cybersecurity fine-tune of 3.5 Flash that found 55 confirmed V8 issues vs 47 for the base model. Gated to governments and trusted partners via the CodeMender pilot — no public access or pricing.

SWE
55
TERM
76
MMLU
88
VIS
85
MATH
88
Input cost—
Output cost—
Context1M
+ PROS
  • +Beat base 3.5 Flash and Opus 4.6 at vulnerability discovery
  • +Found RCE bugs in Google's own APIs within 2 hours
– CONS
  • –Not publicly available — pilot only
  • –No published pricing
VIEW FULL REPORT →
2026-07-2054D AGO · RELEASE
DeepSeek V4-Pro
DEEPSEEKOPEN-WEIGHTS

The 1.6T MoE open-weights flagship: 80.6% SWE-bench Verified (self-reported) at an order of magnitude below closed-frontier pricing. Independent agentic scores land lower — but the value is still absurd.

SWE
65
TERM
78
MMLU
90
VIS
40
MATH
95
Input cost$0.44/M
Output cost$0.87/M
Context1M
+ PROS
  • +Top open-weights coding score at release
  • +Codeforces 3206 — elite competitive coding
  • +MIT license, 1M context, $0.87/M output
– CONS
  • –Independent evals land below self-reported numbers
  • –Text-only; surge pricing during Beijing hours
VIEW FULL REPORT →
2026-07-1658D AGO · RELEASE
Kimi K3
MOONSHOTPREMIUM

The closest Chinese challenger to the frontier — #4 of all models on AA Intelligence Index, ahead of Opus 4.8. At 2.8T parameters, it's the largest open-weight release in history.

SWE
70
TERM
88
MMLU
92
VIS
85
MATH
94
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
  • +#4 overall on aggregate intelligence
  • +Largest open-weight model ever (2.8T, weights July 26)
  • +Terminal-Bench 2.0 88.3 (Moonshot-reported)
– CONS
  • –Always-on thinking burns output tokens; slow
  • –Signups paused July 19 over GPU capacity
VIEW FULL REPORT →
2026-07-0965D AGO · RELEASE
GPT-5.6 Luna
OPENAIBUDGET

The budget disruptor of 2026. After the ~80% July 30 price cut, Luna posts scores within 2–4 points of Sol at 1/25th the output cost. Just don't trust it past 512K tokens.

SWE
63
TERM
85
MMLU
90
VIS
80
MATH
91
Input cost$0.20/M
Output cost$1.20/M
Context1M
+ PROS
  • +GPQA Diamond 92.3% at $0.20/M input
  • +1.05M context at commodity pricing
  • +Within ~2–4 points of Sol on most benchmarks
– CONS
  • –Long-context recall collapses past 512K (41.3%)
  • –Text and image input only
VIEW FULL REPORT →
2026-07-0866D AGO · RELEASE
Grok 4.5
XAIBALANCED

xAI's coding-first model on the 1.5T V9 base. Wins SWE Marathon (29% pass@1, ahead of Opus 4.8 and Fable 5) and solves SWE tasks with ~4.2x fewer tokens than Opus 4.8.

SWE
65
TERM
78
MMLU
88
VIS
78
MATH
89
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
  • +Top SWE Marathon score for long sessions
  • +~60–80% cheaper per solved task than rivals
  • +Exceptional token efficiency
– CONS
  • –Raw ceiling trails the frontier (64.7% SWE-bench Pro)
  • –500K context; price doubles at ≥200K prompt
VIEW FULL REPORT →
2026-06-3074D AGO · RELEASE
Claude Sonnet 5
ANTHROPICPREMIUM

The most agentic Sonnet yet and the new claude.ai default — 72.7% SWE-bench Verified (vs 62.3% for Sonnet 4.6) at the same price, with a $2/$10 promo through August 31.

SWE
66
TERM
80
MMLU
91
VIS
86
MATH
92
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
  • +Big agentic jump over Sonnet 4.6 at the same price
  • +78.5% OSWorld-Verified computer use
  • +Promo pricing $2/$10 through Aug 31, 2026
– CONS
  • –20+ points behind Opus 5 on SWE-bench Verified
  • –New tokenizer raises effective cost ~1.0–1.35x
VIEW FULL REPORT →
2026-06-2678D AGO · RELEASE
GPT-5.6 Sol
OPENAIPREMIUM

OpenAI's flagship — Terminal-Bench 2.1 leader at 88.8% (91.9% in ultra mode with sub-agents) and the first frontier model to clear a US government review. Opus 5 still owns repo-level coding.

SWE
65
TERM
89
MMLU
93
VIS
90
MATH
95
Input cost$5.00/M
Output cost$30.00/M
Context1M
+ PROS
  • +Terminal-Bench 2.1 leader (88.8%, 91.9% ultra)
  • +GPQA Diamond 94.6%, BrowseComp 90.4%
  • +Ultra mode spawns sub-agents for long workflows
– CONS
  • –SWE-bench Pro 64.6% vs Opus 5's 79.2%
  • –Long-context surcharge $10/$45 above 272K
VIEW FULL REPORT →
2026-06-2678D AGO · RELEASE
GPT-5.6 Terra
OPENAIBALANCED

The sensible OpenAI default: within 1–4 points of Sol on core benchmarks at 60% lower output cost after the July 30 price cut. Makes GPT-5.5's rate card look obsolete.

SWE
63
TERM
87
MMLU
92
VIS
88
MATH
93
Input cost$2.00/M
Output cost$12.00/M
Context1M
+ PROS
  • +~97% of Sol's benchmark line at $2/$12
  • +Long-context recall nearly matches Sol
  • +Cut ~20% on July 30, 2026
– CONS
  • –Big gap to Sol on computer use
  • –Well behind Opus 5 on repo-level coding
VIEW FULL REPORT →
2026-06-1688D AGO · RELEASE
GLM-5.2
Z.AIOPEN-WEIGHTS

The top open-weights coding model of mid-2026 — beats GPT-5.5 on SWE-bench Pro at roughly a sixth of the cost, MIT-licensed, with two reasoning-effort modes.

SWE
62
TERM
74
MMLU
85
VIS
60
MATH
86
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
  • +SWE-bench Pro 62.1 — ahead of GPT-5.5
  • +MIT license; third-party hosts from ~$0.75/M
  • +Coding plan from ~$12.60/mo effective
– CONS
  • –AA Index 51 vs Opus 5's 61 — clear frontier gap
  • –No launch-day benchmarks from Z.ai itself
VIEW FULL REPORT →
2026-06-1292D AGO · RELEASE
Kimi K2.7 Code
MOONSHOTBUDGET

Open-weight 1T MoE coding specialist with only 32B active params — fast, cheap to serve, and ~30% more token-efficient than its predecessor.

SWE
60
TERM
75
MMLU
82
VIS
55
MATH
83
Input cost$0.95/M
Output cost$4.00/M
Context256K
+ PROS
  • ++21.8% over K2.6 on coding evals with fewer thinking tokens
  • +Modified MIT license, weights on Hugging Face
  • +$0.95/M input at coding-specialist quality
– CONS
  • –256K context limits big-repo agent work
  • –General reasoning lags the frontier
VIEW FULL REPORT →
MODEL OF THE YEAR
2026-06-0995D AGO · RELEASE
Claude Fable 5
ANTHROPICPREMIUM

The new global #1. 80.3% SWE-Bench Pro is an 11-point leap over Opus 4.8 (69.2%) — the biggest single-release jump of 2026. 1932 GDPval-AA, 1M context, native parallel subagents. Costs 2x Opus 4.8 ($10/$50), so reserve it for the hardest agentic and engineering work.

SWE
80
TERM
86
MMLU
94
VIS
90
MATH
97
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
  • +80.3% SWE-Bench Pro — new #1, +11 pts over Opus 4.8
  • +1932 GDPval-AA, ahead of Opus 4.8 (1890)
  • +Mythos-class capability, generally available
  • +1M context + native parallel subagents
– CONS
  • –$10/$50 — double Opus 4.8's price
  • –Deliberate pace; not for latency-sensitive apps
VIEW FULL REPORT →
2026-06-0995D AGO · RESTRICTED
Claude Mythos 5
ANTHROPICFRONTIERNOT AVAILABLE

The same model as Fable 5 with safeguards lifted in high-risk areas — restricted to vetted partners for advanced cybersecurity and research. For everyone else, Fable 5 is identical at $10/$50 with standard safety controls.

SWE
80
TERM
86
MMLU
94
VIS
90
MATH
97
Input cost—
Output cost—
Context1M
+ PROS
  • +Tied with Fable 5 as the highest public coding score (80.3% SWE-Bench Pro)
  • +Safeguards lifted for advanced security and research
  • +Same 1M context + 1932 GDPval-AA as Fable 5
– CONS
  • –Not generally available — vetted partners only
  • –Most teams should use Fable 5 instead
VIEW FULL REPORT →
2026-05-27108D AGO · RELEASE
Claude Opus 4.8
ANTHROPICPREMIUM

The best-value premium model. 69.2% SWE-Bench Pro and 1890 Elo at $5/$25 — but Claude Fable 5 (80.3%) now leads the frontier at 2x the price, so Opus 4.8 is the smarter default for most premium work.

SWE
69
TERM
83
MMLU
93
VIS
99
MATH
96
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +69.2% SWE-Bench Pro at $5/$25 — best value at the top
  • +1890 Arena Elo (67% win rate vs GPT-5.5)
  • +Native parallel subagents built in
– CONS
  • –Superseded by Fable 5 (80.3%) on raw coding
  • –Deliberate speed — not for latency-sensitive apps
VIEW FULL REPORT →
2026-05-20115D AGO · RELEASE
Qwen 3.7 Max
ALIBABABALANCED

Alibaba's first closed-weight flagship — an agent-first model built for ~35-hour autonomous runs. Cracked the global top 5 on agentic evals at launch; superseded by Qwen 3.8 Max in August.

SWE
61
TERM
70
MMLU
89
VIS
82
MATH
90
Input cost$2.50/M
Output cost$7.50/M
Context1M
+ PROS
  • +GPQA Diamond 92.4 — near the top of the field
  • +Built for marathon agent sessions
  • +90% cached-input discount
– CONS
  • –API-only, no open weights — a break from Qwen tradition
  • –Qwen 3.8 Max is stronger and cheaper
VIEW FULL REPORT →
2026-05-19116D AGO · RELEASE
Gemini 3.5 Flash
GOOGLEBALANCED

Google's I/O headliner: a Flash-tier model that beat Gemini 3.1 Pro on agentic benchmarks at ~278 tokens/sec. Superseded by 3.6 Flash two months later.

SWE
55
TERM
76
MMLU
88
VIS
87
MATH
89
Input cost$1.50/M
Output cost$9.00/M
Context1M
+ PROS
  • +Beat 3.1 Pro on agentic/coding despite Flash tier
  • +~4x faster than comparable frontier models
  • +Default model in the Gemini app
– CONS
  • –3.6 Flash beats it on every tested benchmark, cheaper
  • –No 3.5 Pro sibling ever shipped
VIEW FULL REPORT →
2026-05-12123D AGO · RELEASE
Mistral Medium 3.1
MISTRALBALANCED

Europe's strongest open release of 2026. A clean middle option for teams that need a non-US model.

SWE
49
TERM
62
MMLU
86
VIS
74
MATH
82
Input cost$1.00/M
Output cost$4.50/M
Context256K
+ PROS
  • +EU-hosted option
  • +Apache 2.0 license
  • +Good speed
– CONS
  • –Below frontier on coding
  • –Smaller ecosystem
VIEW FULL REPORT →
2026-04-30135D AGO · RELEASE
GPT-5.5
OPENAIPREMIUM

Best for agentic, computer-use, and Codex workflows. The right pick if your stack is already OpenAI-native.

SWE
60
TERM
83
MMLU
91
VIS
86
MATH
93
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +Top Terminal-Bench at 82.7%
  • +Best computer-use ability
  • +1M context
– CONS
  • –Vision lags Opus 4.7
  • –More expensive than 5.4 for marginal gains
VIEW FULL REPORT →
2026-04-29136D AGO · RELEASE
Mistral Medium 3.5
MISTRALBALANCED

77.6% SWE-bench Verified from a 128B dense open-weight model with vision — within ~2 points of Sonnet 4.6 at half the price, and actually self-hostable.

SWE
58
TERM
68
MMLU
87
VIS
80
MATH
84
Input cost$1.50/M
Output cost$7.50/M
Context256K
+ PROS
  • +Strongest dense open-weights coding score at release
  • +Single-checkpoint text + vision
  • +Modified MIT license, EU-hosted option
– CONS
  • –256K context lags the 1M frontier norm
  • –Sparse benchmark disclosure at launch
VIEW FULL REPORT →
2026-04-16149D AGO · RELEASE
Claude Opus 4.7
ANTHROPICPREMIUM

Was #1 on SWE-Bench Pro at 64.3% — now superseded by Opus 4.8 (69.2%) at the same price. Vision accuracy 98.5%, strong agentic recall. Still fully supported.

SWE
64
TERM
78
MMLU
92
VIS
99
MATH
94
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +SWE-Bench Pro 64.3% — still top-tier
  • +Vision accuracy 98.5%
  • +1M context with sharp recall
– CONS
  • –Opus 4.8 is strictly better at the same price
  • –New tokenizer can raise effective cost by 35%
VIEW FULL REPORT →
2026-04-11154D AGO · DISCLOSED
Claude Mythos
ANTHROPICINTERNALNOT AVAILABLE

Anthropic's most powerful internal model. Found thousands of zero-days autonomously. Not released publicly.

SWE
87
TERM
93
MMLU
96
VIS
99
MATH
98
Input cost—
Output cost—
Context—
+ PROS
  • +Frontier of frontier
  • +Autonomous capabilities reported
– CONS
  • –Not available — disclosed only
VIEW FULL REPORT →
2026-04-08157D AGO · RELEASE
Llama 5
METAOPEN-WEIGHTS

600B open weights, a 5M-token context window, and Recursive Self-Improvement that re-checks its own reasoning mid-task. No major host has published verified per-token pricing yet — it joins our ranked catalog when one does.

SWE
62
TERM
70
MMLU
90
VIS
78
MATH
88
Input cost—
Output cost—
Context5M
+ PROS
  • +5M context — largest of any 2026 frontier release
  • +Open weights under the Meta Open License
  • +Self-corrects reasoning without external feedback
– CONS
  • –No verified hosted API pricing yet
  • –Needs 8x H100 minimum to self-host
VIEW FULL REPORT →
2026-04-08157D AGO · RELEASE
Muse Spark
METABALANCED

Meta's first closed frontier model — GPT-5.5-tier intelligence with the broadest multimodal input (text, image, video, audio, PDF) at $1.25/$4.25. The surprise value play of the year.

SWE
60
TERM
80
MMLU
89
VIS
88
MATH
87
Input cost$1.25/M
Output cost$4.25/M
Context1M
+ PROS
  • +#5 overall on GDPval-AA v2, ahead of Opus 4.8 on agentic tasks
  • +Video, audio, and PDF input — broadest of any model
  • +~$0.40/task — undercuts OpenAI and Anthropic
– CONS
  • –AA Intelligence Index 54 trails Opus 5 (61) and Sol (59)
  • –Public API only since July — young ecosystem
VIEW FULL REPORT →
2026-03-22174D AGO · RELEASE
GPT-5.4
OPENAIPREMIUM

OpenAI's best price/quality. Pair with Claude Opus for hybrid stacks — they're complementary, not competitive.

SWE
58
TERM
80
MMLU
90
VIS
84
MATH
92
Input cost$2.50/M
Output cost$15.00/M
Context272K
+ PROS
  • +7× cheaper than Opus 4.7
  • +Top-tier reasoning
  • +Mature tools/agents
– CONS
  • –Context capped at 272K
  • –Vision lags Gemini
VIEW FULL REPORT →
2026-03-04192D AGO · RELEASE
Gemini 3.1 Pro
GOOGLEPREMIUM

Research workhorse. 2M context, native multimodality, and the best-priced premium model in the directory.

SWE
53
TERM
67
MMLU
90
VIS
90
MATH
90
Input cost$1.25/M
Output cost$10.00/M
Context2M
+ PROS
  • +2M context
  • +Best research score
  • +Best price-per-quality at premium tier
– CONS
  • –Lags top tier on raw coding
  • –Stuck inside Google's tooling
VIEW FULL REPORT →
2026-02-18206D AGO · RELEASE
Grok 4
XAIBALANCED

Strong coding value at 2M context. Underrated at this price tier. The contrarian voice helps in research.

SWE
56
TERM
70
MMLU
86
VIS
72
MATH
87
Input cost$2.00/M
Output cost$10.00/M
Context2M
+ PROS
  • +2M context at $2/M input
  • +Strong reasoning
  • +Real-time X data integration
– CONS
  • –Writing voice is uneven
  • –Smaller ecosystem
VIEW FULL REPORT →
2026-02-04220D AGO · RELEASE
Llama 4 Scout
METABUDGET

10M context is the headline. Useful for indexing entire codebases but accuracy degrades past 1M.

SWE
41
TERM
56
MMLU
82
VIS
70
MATH
78
Input cost$0.30/M
Output cost$1.20/M
Context10M
+ PROS
  • +10M context window — by far the largest
  • +Open weights
  • +Cheap
– CONS
  • –Long-context accuracy thins out past 1M
  • –Below frontier on reasoning
VIEW FULL REPORT →
2026-02-04220D AGO · RELEASE
Llama 4 Maverick
METABUDGET

Biggest open-weight leap of 2026. Competitive with GPT-5.4 on general tasks at a quarter of the price.

SWE
47
TERM
60
MMLU
85
VIS
76
MATH
84
Input cost$0.60/M
Output cost$2.40/M
Context256K
+ PROS
  • +Open weights at near-frontier quality
  • +Fast
  • +Strong math
– CONS
  • –256K context lags Scout
  • –No native multimodality
VIEW FULL REPORT →
2026-01-22233D AGO · RELEASE
DeepSeek V4
DEEPSEEKOPEN-WEIGHTS

Open-weights, $0.27/M input, beats GPT-4o on coding. Quietly the most disruptive release of January.

SWE
53
TERM
64
MMLU
84
VIS
60
MATH
88
Input cost$0.27/M
Output cost$1.10/M
Context128K
+ PROS
  • +Cheapest serious code model
  • +Open weights — self-hostable
  • +Strong math
– CONS
  • –Data residency questions for some teams
  • –Vision is weak
VIEW FULL REPORT →
2026-01-08247D AGO · RELEASE
Claude Haiku 4.5
ANTHROPICBUDGET

Fast, cheap, surprisingly capable. The cheapest model in the lineup that you can actually ship behind a feature flag.

SWE
28
TERM
41
MMLU
79
VIS
68
MATH
72
Input cost$0.80/M
Output cost$4.00/M
Context200K
+ PROS
  • +96-score speed — fastest in directory
  • +Cheapest serious model at $0.80/M input
  • +Vision matches mid-tier from 2025
– CONS
  • –SWE-Bench Pro under 30%
  • –Context capped at 200K
VIEW FULL REPORT →
FAQ / END

Frequently actually asked.

6 ENTRIES
+

Keep going. More guides.

RELATED
DIRECTORY
All AI models

Every model currently tracked, with live pricing.

COMPARE
Opus 5 vs GPT-5.6 Sol

The 2026 flagship head-to-head: coding, agentic, and price.

GUIDE
Best AI for coding

The 2026 winner by use case and budget.

PRICING
Price history

Track every API price move since launch.

DATA
Benchmark scores

Raw numbers across every benchmark we cite.

GUIDE
Best cheap AI

Free + $20 chatbots ranked by Value Index.