UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsGuidesEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

Benchmarks

AI Model Benchmark Scores

Compare MMLU, HumanEval, SWE-bench, GPQA, and MATH scores across all major AI models. Click any column to sort. Click a benchmark name to learn what it measures.

Scores are reported values from provider papers and public leaderboards · Updated as new results are published

See the dedicated SWE-bench leaderboard with pricing →

New to SWE-bench? What counts as a good score · Verified vs full vs Lite

ModelDetails
GPT-5.6 Sol
OpenAI
——
96.2%
94.6%
——Full review →
Claude Opus 5
Anthropic
——
96%
———Full review →
Claude Mythos 5
Anthropic
94.2%
97.1%
95.5%
83.5%
97.4%
1,932
Full review →
Claude Fable 5
Anthropic
94.2%
97.1%
95%
83.5%
97.4%
1,932
Full review →
Claude Opus 4.8
Anthropic
93%
95.5%
88.6%
79.2%
96%
1,890
Full review →
Claude Opus 4.7
Anthropic
92%
93%
87.6%
76%
94%
1,800
Full review →
Grok 4.5
xAI
——
86.6%
———Full review →
Claude Sonnet 5
Anthropic
——
85.2%
———Full review →
Claude Opus 4.6
Anthropic
90.4%
92%
80.8%
74.9%
89.3%
1,360
Full review →
Gemini 3.7 Flash
Google
——
80.8%
———Full review →
Gemini 3.1 Pro
Google
90%
92%
80.6%
84%
91.6%
1,380
Full review →
Qwen 3.7 Max
Alibaba
——
80.4%
———Full review →
GPT-5.2
OpenAI
——
80%
———Full review →
Claude Sonnet 4.6
Anthropic
88.3%
90.1%
79.6%
68%
85.1%
1,340
Full review →
Mistral Medium 3.5
Mistral
——
77.6%
———Full review →
Muse Spark
Meta
——
77.4%
———Full review →
Muse Glimmer 30B
Meta
——
76%
———Full review →
GPT-5.4
OpenAI
91%
91.5%
74.9%
75.4%
91%
1,355
Full review →
Grok 4
xAI
87.5%
88%
54%
72%
87%
1,305
Full review →
DeepSeek R1
DeepSeek
90.8%
92%
49.2%
71.5%
97.3%
1,320
Full review →
GPT-4o
OpenAI
88.7%
90.2%
46%
53.6%
76.6%
1,295
Full review →
Claude Haiku 4
Anthropic
80%
84%
43%
41.5%
71%
1,210
Full review →
DeepSeek V3
DeepSeek
88.5%
90.2%
42%
59.1%
90.2%
1,305
Full review →
Gemini 3.1 Flash
Google
84%
86.5%
35%
51%
78.4%
1,265
Full review →
Llama 4 Maverick
Meta
85.5%
87.5%
32%
52%
80.5%
1,250
Full review →
Mistral Large 2
Mistral
84%
92%
28%
49.6%
72%
1,225
Full review →
GPT-4o Mini
OpenAI
82%
87.2%
23.6%
40.2%
70.2%
1,235
Full review →
GPT-6 Astra
OpenAI
———
96%
——Full review →
GLM-5.3
Z.ai
———
91.7%
——Full review →
Grok 4.6
xAI
——————Full review →
GLM-5.3 Flash
Z.ai
——————Full review →
Qwen 3.8 Flash
Alibaba
——————Full review →

What do these benchmarks measure?

MMLU

General knowledge and reasoning across 57 academic subjects

HumanEval

Python code generation — pass@1 accuracy on 164 problems

SWE-bench

Real-world GitHub issues resolved autonomously

GPQA

Graduate-level biology, chemistry, physics questions

MATH

Competition mathematics — algebra, geometry, calculus

Arena Elo

Human preference Elo score from Chatbot Arena head-to-head battles

A note on benchmarks

Benchmarks measure specific, testable capabilities — not overall "intelligence" or real-world usefulness. A model that tops SWE-bench may still frustrate developers with its API latency or context handling. Use these scores as one signal, not the final word. Our model reviews combine benchmarks with practical verdict assessments for a fuller picture.

Compare models side by side →View price history →