UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansCost CalculatorWhich AI is Cheapest?

Company

About UseRightAIContactWhat ChangedAll ModelsDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

JUST ANNOUNCED · JUN 26GPT-5.6 — OpenAI's new Sol, Terra & Luna modelsLimited preview; verified benchmarks land here as OpenAI publishes them.FIRST LOOK →
UPDATED 2026-08-0631 PUBLIC RELEASES YTDSCORED ON 5 BENCHMARKS0 PAID RANKINGS

31 new flagship models.
One year-defining release.

Every major model that shipped in 2026 — ranked, benchmarked, and dated. Scrub the timeline to see how the field reshaped itself month by month.

The 2026 release timeline.

JAN 01 ─── DEC 31
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
NOV
DEC
Claude Haiku 4.5
DeepSeek V4
Llama 4 Scout
Llama 4 Maverick
Grok 4
Gemini 3.1 Pro
GPT-5.4
Llama 5
Muse Spark
Claude Opus 4.7
Mistral Medium 3.5
GPT-5.5
Mistral Medium 3.1
Gemini 3.5 Flash
Qwen 3.7 Max
Claude Opus 4.8
Claude Fable 5
Kimi K2.7 Code
GLM-5.2
GPT-5.6 Sol
GPT-5.6 Terra
Claude Sonnet 5
Grok 4.5
GPT-5.6 Luna
Kimi K3
DeepSeek V4-Pro
Gemini 3.6 Flash
Gemini 3.5 Flash-Lite
Claude Opus 5
DeepSeek V4-Flash
Qwen 3.8 Max
2026-07-2415 DAYS AGO
Claude Opus 5
ANTHROPIC · PREMIUM
The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.
SWE-BENCH
79.2
CONTEXT
1M
$/M IN
$5
READ FULL REPORT →
01 / 06

The three picks that shaped the year.

EDITOR'S VERDICT
BEST OVERALL · 2026
Claude Opus 5
ANTHROPIC · PREMIUM

SWE-bench Verified leader (~96%) and Terminal-Bench 89.1% at Opus 4.8's exact $5/$25 price — within half a point of Fable 5 on agentic coding at half the cost. Fable 5 keeps the absolute ceiling; Opus 5 is the one to actually ship on.

RELEASED
07-24
$/M IN
$5
CONTEXT
1M
BEST VALUE
GPT-5.6 Luna
OPENAI · BUDGET

After the ~80% July price cut: scores within 2–4 points of GPT-5.6 Sol at $0.20/M input. The budget disruptor of the year.

RELEASED
07-09
$/M IN
$0.2
CONTEXT
1M
BEST OPEN-WEIGHTS
DeepSeek V4-Pro
DEEPSEEK · OPEN-WEIGHTS

80.6% SWE-bench Verified (self-reported), MIT license, 1M context, $0.87/M output — an order of magnitude cheaper than the closed frontier.

RELEASED
07-20
$/M IN
$0.44
CONTEXT
1M
02 / 06

Benchmark battle. Pick the lens.

5 BENCHMARKS · 31 MODELS
01Claude Fable 5
80.3$10/M
02Claude Opus 5
79.2$5/M
03Kimi K3
70.0$3/M
04Claude Opus 4.8
69.2$5/M
05Qwen 3.8 Max
67.7$2/M
06Claude Sonnet 5
66.0$3/M
07DeepSeek V4-Pro
65.0$0.44/M
08Grok 4.5
64.7$2/M
09GPT-5.6 Sol
64.6$5/M
10Claude Opus 4.7
64.3$5/M
11GPT-5.6 Terra
63.4$2/M
12GPT-5.6 Luna
62.7$0.2/M
13GLM-5.2
62.1$1.4/M
14Llama 5
62.0—
15Qwen 3.7 Max
60.6$2.5/M
16GPT-5.5
60.1$5/M
17Muse Spark
60.0$1.25/M
18Kimi K2.7 Code
60.0$0.95/M
19Mistral Medium 3.5
58.0$1.5/M
20DeepSeek V4-Flash
58.0$0.14/M
21GPT-5.4
57.7$2.5/M
22Gemini 3.6 Flash
57.0$1.5/M
23Grok 4
56.0$2/M
24Gemini 3.5 Flash
55.1$1.5/M
25Gemini 3.5 Flash-Lite
54.2$0.3/M
26Gemini 3.1 Pro
53.2$1.25/M
27DeepSeek V4
52.8$0.27/M
28Mistral Medium 3.1
49.0$1/M
29Llama 4 Maverick
47.2$0.6/M
30Llama 4 Scout
41.0$0.3/M
31Claude Haiku 4.5
28.4$0.8/M
03 / 06

Quality per dollar. The full chart.

SCATTER · 30 MODELS
Claude Haiku 4.5
DeepSeek V4
Llama 4 Scout
Llama 4 Maverick
Grok 4
Gemini 3.1 Pro
GPT-5.4
Claude Opus 4.7
GPT-5.5
Mistral Medium 3.1
Claude Opus 4.8
Claude Fable 5
Muse Spark
Mistral Medium 3.5
Gemini 3.5 Flash
Qwen 3.7 Max
Kimi K2.7 Code
GLM-5.2
GPT-5.6 Sol
GPT-5.6 Terra
Claude Sonnet 5
Grok 4.5
GPT-5.6 Luna
Kimi K3
DeepSeek V4-Pro
Gemini 3.6 Flash
Gemini 3.5 Flash-Lite
Claude Opus 5
DeepSeek V4-Flash
Qwen 3.8 Max
$/M INPUT →

The Pareto frontier of 2026.

Each dot is one model. Up means better quality, right means more expensive. The top-left edge is the Pareto frontier — every dot under it is strictly dominated.

ON THE FRONTIER
Claude Fable 5 (top quality) · Claude Opus 4.8 (best value at the top) · Grok 4 (best value) · Llama 4 Maverick (best free option)
04 / 06

Who shipped what. Year-to-date.

10 PROVIDERS · 31 RELEASES
Anthropic
6
RELEASES · YTD
OpenAI
5
RELEASES · YTD
Google
4
RELEASES · YTD
Meta
4
RELEASES · YTD
DeepSeek
3
RELEASES · YTD
xAI
2
RELEASES · YTD
Mistral
2
RELEASES · YTD
Alibaba
2
RELEASES · YTD
Moonshot
2
RELEASES · YTD
Z.ai
1
RELEASES · YTD
05 / 06

Every release. One card each.

34 CARDS · NEWEST FIRST
2026-08-035D AGO · RELEASE
Qwen 3.8 Max
ALIBABABALANCED

Alibaba's 2.4T MoE flagship — a legitimate SWE-bench Pro upset over GPT-5.6 Sol (67.7 vs 64.6) and #2 globally for vision. The strongest Chinese multimodal model yet.

SWE
68
TERM
76
MMLU
90
VIS
92
MATH
91
Input cost$2.00/M
Output cost$6.00/M
Context1M
+ PROS
  • +Beats GPT-5.6 Sol on SWE-bench Pro
  • +#2 globally on vision, behind only a Fable 5 variant
  • +First Alibaba open-weights release at this scale
– CONS
  • –No independent benchmarks at GA yet
  • –Trails Fable 5 badly on the coding ceiling
VIEW FULL REPORT →
2026-07-318D AGO · RELEASE
DeepSeek V4-Flash
DEEPSEEKOPEN-WEIGHTS

The best cheap agent engine of 2026: Terminal-Bench 82.7 at $0.14/M input, with MIT weights self-hostable in ~110 GB. Nothing touches its agentic capability per dollar.

SWE
58
TERM
83
MMLU
86
VIS
38
MATH
90
Input cost$0.14/M
Output cost$0.28/M
Context1M
+ PROS
  • +Terminal-Bench 82.7 rivals models 30x its price
  • +Beats the V4-Pro preview on all nine agent benchmarks
  • +2,500 concurrent requests, 1M context
– CONS
  • –Text-only
  • –Several headline numbers from DeepSeek's own evals
VIEW FULL REPORT →
2026-07-2415D AGO · RELEASE
Claude Opus 5
ANTHROPICPREMIUM

The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.

SWE
79
TERM
89
MMLU
93
VIS
88
MATH
96
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +SWE-bench Verified leader (~96%) as of August 2026
  • +Terminal-Bench 2.1 89.1% — edges GPT-5.6 Sol
  • +Within 0.5% of Fable 5 on CursorBench at half the cost
– CONS
  • –Fable 5 still holds the absolute frontier ceiling
  • –Fast mode doubles pricing to $10/$50
VIEW FULL REPORT →
2026-07-2118D AGO · RELEASE
Gemini 3.6 Flash
GOOGLEBALANCED

The best Google model for agents: beats 3.5 Flash on every tested benchmark with ~17% fewer output tokens and cheaper output. The default Gemini pick until Gemini 4 lands.

SWE
57
TERM
80
MMLU
89
VIS
88
MATH
90
Input cost$1.50/M
Output cost$7.50/M
Context1M
+ PROS
  • +OSWorld-Verified 83.0% — best computer use in class
  • +~17% fewer output tokens + lower price compound
  • +~280–304 tokens/sec
– CONS
  • –Point release, not a generational leap
  • –No Pro tier in this generation
VIEW FULL REPORT →
2026-07-2118D AGO · RELEASE
Gemini 3.5 Flash-Lite
GOOGLEBUDGET

The fastest model in Google's lineup at 350 tokens/sec, with real agentic chops at $0.30/M. GPT-5.6 Luna beats it on price-per-benchmark; Flash-Lite answers with speed and full multimodal input.

SWE
54
TERM
54
MMLU
84
VIS
82
MATH
84
Input cost$0.30/M
Output cost$2.50/M
Context1M
+ PROS
  • +350 tokens/sec — fastest in the 3.5 line
  • +Huge jump over 3.1 Flash-Lite
  • +1M context with video/audio/PDF input
– CONS
  • –Trails full Flash on hard agentic work
  • –Luna wins pure price-per-benchmark math
VIEW FULL REPORT →
2026-07-2118D AGO · RESTRICTED
Gemini 3.5 Flash Cyber
GOOGLERESTRICTEDNOT AVAILABLE

A cybersecurity fine-tune of 3.5 Flash that found 55 confirmed V8 issues vs 47 for the base model. Gated to governments and trusted partners via the CodeMender pilot — no public access or pricing.

SWE
55
TERM
76
MMLU
88
VIS
85
MATH
88
Input cost—
Output cost—
Context1M
+ PROS
  • +Beat base 3.5 Flash and Opus 4.6 at vulnerability discovery
  • +Found RCE bugs in Google's own APIs within 2 hours
– CONS
  • –Not publicly available — pilot only
  • –No published pricing
VIEW FULL REPORT →
2026-07-2019D AGO · RELEASE
DeepSeek V4-Pro
DEEPSEEKOPEN-WEIGHTS

The 1.6T MoE open-weights flagship: 80.6% SWE-bench Verified (self-reported) at an order of magnitude below closed-frontier pricing. Independent agentic scores land lower — but the value is still absurd.

SWE
65
TERM
78
MMLU
90
VIS
40
MATH
95
Input cost$0.44/M
Output cost$0.87/M
Context1M
+ PROS
  • +Top open-weights coding score at release
  • +Codeforces 3206 — elite competitive coding
  • +MIT license, 1M context, $0.87/M output
– CONS
  • –Independent evals land below self-reported numbers
  • –Text-only; surge pricing during Beijing hours
VIEW FULL REPORT →
2026-07-1623D AGO · RELEASE
Kimi K3
MOONSHOTPREMIUM

The closest Chinese challenger to the frontier — #4 of all models on AA Intelligence Index, ahead of Opus 4.8. At 2.8T parameters, it's the largest open-weight release in history.

SWE
70
TERM
88
MMLU
92
VIS
85
MATH
94
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
  • +#4 overall on aggregate intelligence
  • +Largest open-weight model ever (2.8T, weights July 26)
  • +Terminal-Bench 2.0 88.3 (Moonshot-reported)
– CONS
  • –Always-on thinking burns output tokens; slow
  • –Signups paused July 19 over GPU capacity
VIEW FULL REPORT →
2026-07-0930D AGO · RELEASE
GPT-5.6 Luna
OPENAIBUDGET

The budget disruptor of 2026. After the ~80% July 30 price cut, Luna posts scores within 2–4 points of Sol at 1/25th the output cost. Just don't trust it past 512K tokens.

SWE
63
TERM
85
MMLU
90
VIS
80
MATH
91
Input cost$0.20/M
Output cost$1.20/M
Context1M
+ PROS
  • +GPQA Diamond 92.3% at $0.20/M input
  • +1.05M context at commodity pricing
  • +Within ~2–4 points of Sol on most benchmarks
– CONS
  • –Long-context recall collapses past 512K (41.3%)
  • –Text and image input only
VIEW FULL REPORT →
2026-07-0831D AGO · RELEASE
Grok 4.5
XAIBALANCED

xAI's coding-first model on the 1.5T V9 base. Wins SWE Marathon (29% pass@1, ahead of Opus 4.8 and Fable 5) and solves SWE tasks with ~4.2x fewer tokens than Opus 4.8.

SWE
65
TERM
78
MMLU
88
VIS
78
MATH
89
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
  • +Top SWE Marathon score for long sessions
  • +~60–80% cheaper per solved task than rivals
  • +Exceptional token efficiency
– CONS
  • –Raw ceiling trails the frontier (64.7% SWE-bench Pro)
  • –500K context; price doubles at ≥200K prompt
VIEW FULL REPORT →
2026-06-3039D AGO · RELEASE
Claude Sonnet 5
ANTHROPICPREMIUM

The most agentic Sonnet yet and the new claude.ai default — 72.7% SWE-bench Verified (vs 62.3% for Sonnet 4.6) at the same price, with a $2/$10 promo through August 31.

SWE
66
TERM
80
MMLU
91
VIS
86
MATH
92
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
  • +Big agentic jump over Sonnet 4.6 at the same price
  • +78.5% OSWorld-Verified computer use
  • +Promo pricing $2/$10 through Aug 31, 2026
– CONS
  • –20+ points behind Opus 5 on SWE-bench Verified
  • –New tokenizer raises effective cost ~1.0–1.35x
VIEW FULL REPORT →
2026-06-2643D AGO · RELEASE
GPT-5.6 Sol
OPENAIPREMIUM

OpenAI's flagship — Terminal-Bench 2.1 leader at 88.8% (91.9% in ultra mode with sub-agents) and the first frontier model to clear a US government review. Opus 5 still owns repo-level coding.

SWE
65
TERM
89
MMLU
93
VIS
90
MATH
95
Input cost$5.00/M
Output cost$30.00/M
Context1M
+ PROS
  • +Terminal-Bench 2.1 leader (88.8%, 91.9% ultra)
  • +GPQA Diamond 94.6%, BrowseComp 90.4%
  • +Ultra mode spawns sub-agents for long workflows
– CONS
  • –SWE-bench Pro 64.6% vs Opus 5's 79.2%
  • –Long-context surcharge $10/$45 above 272K
VIEW FULL REPORT →
2026-06-2643D AGO · RELEASE
GPT-5.6 Terra
OPENAIBALANCED

The sensible OpenAI default: within 1–4 points of Sol on core benchmarks at 60% lower output cost after the July 30 price cut. Makes GPT-5.5's rate card look obsolete.

SWE
63
TERM
87
MMLU
92
VIS
88
MATH
93
Input cost$2.00/M
Output cost$12.00/M
Context1M
+ PROS
  • +~97% of Sol's benchmark line at $2/$12
  • +Long-context recall nearly matches Sol
  • +Cut ~20% on July 30, 2026
– CONS
  • –Big gap to Sol on computer use
  • –Well behind Opus 5 on repo-level coding
VIEW FULL REPORT →
2026-06-1653D AGO · RELEASE
GLM-5.2
Z.AIOPEN-WEIGHTS

The top open-weights coding model of mid-2026 — beats GPT-5.5 on SWE-bench Pro at roughly a sixth of the cost, MIT-licensed, with two reasoning-effort modes.

SWE
62
TERM
74
MMLU
85
VIS
60
MATH
86
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
  • +SWE-bench Pro 62.1 — ahead of GPT-5.5
  • +MIT license; third-party hosts from ~$0.75/M
  • +Coding plan from ~$12.60/mo effective
– CONS
  • –AA Index 51 vs Opus 5's 61 — clear frontier gap
  • –No launch-day benchmarks from Z.ai itself
VIEW FULL REPORT →
2026-06-1257D AGO · RELEASE
Kimi K2.7 Code
MOONSHOTBUDGET

Open-weight 1T MoE coding specialist with only 32B active params — fast, cheap to serve, and ~30% more token-efficient than its predecessor.

SWE
60
TERM
75
MMLU
82
VIS
55
MATH
83
Input cost$0.95/M
Output cost$4.00/M
Context256K
+ PROS
  • ++21.8% over K2.6 on coding evals with fewer thinking tokens
  • +Modified MIT license, weights on Hugging Face
  • +$0.95/M input at coding-specialist quality
– CONS
  • –256K context limits big-repo agent work
  • –General reasoning lags the frontier
VIEW FULL REPORT →
MODEL OF THE YEAR
2026-06-0960D AGO · RELEASE
Claude Fable 5
ANTHROPICPREMIUM

The new global #1. 80.3% SWE-Bench Pro is an 11-point leap over Opus 4.8 (69.2%) — the biggest single-release jump of 2026. 1932 GDPval-AA, 1M context, native parallel subagents. Costs 2x Opus 4.8 ($10/$50), so reserve it for the hardest agentic and engineering work.

SWE
80
TERM
86
MMLU
94
VIS
90
MATH
97
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
  • +80.3% SWE-Bench Pro — new #1, +11 pts over Opus 4.8
  • +1932 GDPval-AA, ahead of Opus 4.8 (1890)
  • +Mythos-class capability, generally available
  • +1M context + native parallel subagents
– CONS
  • –$10/$50 — double Opus 4.8's price
  • –Deliberate pace; not for latency-sensitive apps
VIEW FULL REPORT →
2026-06-0960D AGO · RESTRICTED
Claude Mythos 5
ANTHROPICFRONTIERNOT AVAILABLE

The same model as Fable 5 with safeguards lifted in high-risk areas — restricted to vetted partners for advanced cybersecurity and research. For everyone else, Fable 5 is identical at $10/$50 with standard safety controls.

SWE
80
TERM
86
MMLU
94
VIS
90
MATH
97
Input cost—
Output cost—
Context1M
+ PROS
  • +Tied with Fable 5 as the highest public coding score (80.3% SWE-Bench Pro)
  • +Safeguards lifted for advanced security and research
  • +Same 1M context + 1932 GDPval-AA as Fable 5
– CONS
  • –Not generally available — vetted partners only
  • –Most teams should use Fable 5 instead
VIEW FULL REPORT →
2026-05-2773D AGO · RELEASE
Claude Opus 4.8
ANTHROPICPREMIUM

The best-value premium model. 69.2% SWE-Bench Pro and 1890 Elo at $5/$25 — but Claude Fable 5 (80.3%) now leads the frontier at 2x the price, so Opus 4.8 is the smarter default for most premium work.

SWE
69
TERM
83
MMLU
93
VIS
99
MATH
96
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +69.2% SWE-Bench Pro at $5/$25 — best value at the top
  • +1890 Arena Elo (67% win rate vs GPT-5.5)
  • +Native parallel subagents built in
– CONS
  • –Superseded by Fable 5 (80.3%) on raw coding
  • –Deliberate speed — not for latency-sensitive apps
VIEW FULL REPORT →
2026-05-2080D AGO · RELEASE
Qwen 3.7 Max
ALIBABABALANCED

Alibaba's first closed-weight flagship — an agent-first model built for ~35-hour autonomous runs. Cracked the global top 5 on agentic evals at launch; superseded by Qwen 3.8 Max in August.

SWE
61
TERM
70
MMLU
89
VIS
82
MATH
90
Input cost$2.50/M
Output cost$7.50/M
Context1M
+ PROS
  • +GPQA Diamond 92.4 — near the top of the field
  • +Built for marathon agent sessions
  • +90% cached-input discount
– CONS
  • –API-only, no open weights — a break from Qwen tradition
  • –Qwen 3.8 Max is stronger and cheaper
VIEW FULL REPORT →
2026-05-1981D AGO · RELEASE
Gemini 3.5 Flash
GOOGLEBALANCED

Google's I/O headliner: a Flash-tier model that beat Gemini 3.1 Pro on agentic benchmarks at ~278 tokens/sec. Superseded by 3.6 Flash two months later.

SWE
55
TERM
76
MMLU
88
VIS
87
MATH
89
Input cost$1.50/M
Output cost$9.00/M
Context1M
+ PROS
  • +Beat 3.1 Pro on agentic/coding despite Flash tier
  • +~4x faster than comparable frontier models
  • +Default model in the Gemini app
– CONS
  • –3.6 Flash beats it on every tested benchmark, cheaper
  • –No 3.5 Pro sibling ever shipped
VIEW FULL REPORT →
2026-05-1288D AGO · RELEASE
Mistral Medium 3.1
MISTRALBALANCED

Europe's strongest open release of 2026. A clean middle option for teams that need a non-US model.

SWE
49
TERM
62
MMLU
86
VIS
74
MATH
82
Input cost$1.00/M
Output cost$4.50/M
Context256K
+ PROS
  • +EU-hosted option
  • +Apache 2.0 license
  • +Good speed
– CONS
  • –Below frontier on coding
  • –Smaller ecosystem
VIEW FULL REPORT →
2026-04-30100D AGO · RELEASE
GPT-5.5
OPENAIPREMIUM

Best for agentic, computer-use, and Codex workflows. The right pick if your stack is already OpenAI-native.

SWE
60
TERM
83
MMLU
91
VIS
86
MATH
93
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +Top Terminal-Bench at 82.7%
  • +Best computer-use ability
  • +1M context
– CONS
  • –Vision lags Opus 4.7
  • –More expensive than 5.4 for marginal gains
VIEW FULL REPORT →
2026-04-29101D AGO · RELEASE
Mistral Medium 3.5
MISTRALBALANCED

77.6% SWE-bench Verified from a 128B dense open-weight model with vision — within ~2 points of Sonnet 4.6 at half the price, and actually self-hostable.

SWE
58
TERM
68
MMLU
87
VIS
80
MATH
84
Input cost$1.50/M
Output cost$7.50/M
Context256K
+ PROS
  • +Strongest dense open-weights coding score at release
  • +Single-checkpoint text + vision
  • +Modified MIT license, EU-hosted option
– CONS
  • –256K context lags the 1M frontier norm
  • –Sparse benchmark disclosure at launch
VIEW FULL REPORT →
2026-04-16114D AGO · RELEASE
Claude Opus 4.7
ANTHROPICPREMIUM

Was #1 on SWE-Bench Pro at 64.3% — now superseded by Opus 4.8 (69.2%) at the same price. Vision accuracy 98.5%, strong agentic recall. Still fully supported.

SWE
64
TERM
78
MMLU
92
VIS
99
MATH
94
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
  • +SWE-Bench Pro 64.3% — still top-tier
  • +Vision accuracy 98.5%
  • +1M context with sharp recall
– CONS
  • –Opus 4.8 is strictly better at the same price
  • –New tokenizer can raise effective cost by 35%
VIEW FULL REPORT →
2026-04-11119D AGO · DISCLOSED
Claude Mythos
ANTHROPICINTERNALNOT AVAILABLE

Anthropic's most powerful internal model. Found thousands of zero-days autonomously. Not released publicly.

SWE
87
TERM
93
MMLU
96
VIS
99
MATH
98
Input cost—
Output cost—
Context—
+ PROS
  • +Frontier of frontier
  • +Autonomous capabilities reported
– CONS
  • –Not available — disclosed only
VIEW FULL REPORT →
2026-04-08122D AGO · RELEASE
Llama 5
METAOPEN-WEIGHTS

600B open weights, a 5M-token context window, and Recursive Self-Improvement that re-checks its own reasoning mid-task. No major host has published verified per-token pricing yet — it joins our ranked catalog when one does.

SWE
62
TERM
70
MMLU
90
VIS
78
MATH
88
Input cost—
Output cost—
Context5M
+ PROS
  • +5M context — largest of any 2026 frontier release
  • +Open weights under the Meta Open License
  • +Self-corrects reasoning without external feedback
– CONS
  • –No verified hosted API pricing yet
  • –Needs 8x H100 minimum to self-host
VIEW FULL REPORT →
2026-04-08122D AGO · RELEASE
Muse Spark
METABALANCED

Meta's first closed frontier model — GPT-5.5-tier intelligence with the broadest multimodal input (text, image, video, audio, PDF) at $1.25/$4.25. The surprise value play of the year.

SWE
60
TERM
80
MMLU
89
VIS
88
MATH
87
Input cost$1.25/M
Output cost$4.25/M
Context1M
+ PROS
  • +#5 overall on GDPval-AA v2, ahead of Opus 4.8 on agentic tasks
  • +Video, audio, and PDF input — broadest of any model
  • +~$0.40/task — undercuts OpenAI and Anthropic
– CONS
  • –AA Intelligence Index 54 trails Opus 5 (61) and Sol (59)
  • –Public API only since July — young ecosystem
VIEW FULL REPORT →
2026-03-22139D AGO · RELEASE
GPT-5.4
OPENAIPREMIUM

OpenAI's best price/quality. Pair with Claude Opus for hybrid stacks — they're complementary, not competitive.

SWE
58
TERM
80
MMLU
90
VIS
84
MATH
92
Input cost$2.50/M
Output cost$15.00/M
Context272K
+ PROS
  • +7× cheaper than Opus 4.7
  • +Top-tier reasoning
  • +Mature tools/agents
– CONS
  • –Context capped at 272K
  • –Vision lags Gemini
VIEW FULL REPORT →
2026-03-04157D AGO · RELEASE
Gemini 3.1 Pro
GOOGLEPREMIUM

Research workhorse. 2M context, native multimodality, and the best-priced premium model in the directory.

SWE
53
TERM
67
MMLU
90
VIS
90
MATH
90
Input cost$1.25/M
Output cost$10.00/M
Context2M
+ PROS
  • +2M context
  • +Best research score
  • +Best price-per-quality at premium tier
– CONS
  • –Lags top tier on raw coding
  • –Stuck inside Google's tooling
VIEW FULL REPORT →
2026-02-18171D AGO · RELEASE
Grok 4
XAIBALANCED

Strong coding value at 2M context. Underrated at this price tier. The contrarian voice helps in research.

SWE
56
TERM
70
MMLU
86
VIS
72
MATH
87
Input cost$2.00/M
Output cost$10.00/M
Context2M
+ PROS
  • +2M context at $2/M input
  • +Strong reasoning
  • +Real-time X data integration
– CONS
  • –Writing voice is uneven
  • –Smaller ecosystem
VIEW FULL REPORT →
2026-02-04185D AGO · RELEASE
Llama 4 Scout
METABUDGET

10M context is the headline. Useful for indexing entire codebases but accuracy degrades past 1M.

SWE
41
TERM
56
MMLU
82
VIS
70
MATH
78
Input cost$0.30/M
Output cost$1.20/M
Context10M
+ PROS
  • +10M context window — by far the largest
  • +Open weights
  • +Cheap
– CONS
  • –Long-context accuracy thins out past 1M
  • –Below frontier on reasoning
VIEW FULL REPORT →
2026-02-04185D AGO · RELEASE
Llama 4 Maverick
METABUDGET

Biggest open-weight leap of 2026. Competitive with GPT-5.4 on general tasks at a quarter of the price.

SWE
47
TERM
60
MMLU
85
VIS
76
MATH
84
Input cost$0.60/M
Output cost$2.40/M
Context256K
+ PROS
  • +Open weights at near-frontier quality
  • +Fast
  • +Strong math
– CONS
  • –256K context lags Scout
  • –No native multimodality
VIEW FULL REPORT →
2026-01-22198D AGO · RELEASE
DeepSeek V4
DEEPSEEKOPEN-WEIGHTS

Open-weights, $0.27/M input, beats GPT-4o on coding. Quietly the most disruptive release of January.

SWE
53
TERM
64
MMLU
84
VIS
60
MATH
88
Input cost$0.27/M
Output cost$1.10/M
Context128K
+ PROS
  • +Cheapest serious code model
  • +Open weights — self-hostable
  • +Strong math
– CONS
  • –Data residency questions for some teams
  • –Vision is weak
VIEW FULL REPORT →
2026-01-08212D AGO · RELEASE
Claude Haiku 4.5
ANTHROPICBUDGET

Fast, cheap, surprisingly capable. The cheapest model in the lineup that you can actually ship behind a feature flag.

SWE
28
TERM
41
MMLU
79
VIS
68
MATH
72
Input cost$0.80/M
Output cost$4.00/M
Context200K
+ PROS
  • +96-score speed — fastest in directory
  • +Cheapest serious model at $0.80/M input
  • +Vision matches mid-tier from 2025
– CONS
  • –SWE-Bench Pro under 30%
  • –Context capped at 200K
VIEW FULL REPORT →
FAQ / END

Frequently actually asked.

6 ENTRIES
+

Keep going. More guides.

RELATED
DIRECTORY
All AI models

Every model currently tracked, with live pricing.

COMPARE
Opus 5 vs GPT-5.6 Sol

The 2026 flagship head-to-head: coding, agentic, and price.

GUIDE
Best AI for coding

The 2026 winner by use case and budget.

PRICING
Price history

Track every API price move since launch.

DATA
Benchmark scores

Raw numbers across every benchmark we cite.

GUIDE
Best cheap AI

Free + $20 chatbots ranked by Value Index.