UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

HomeBenchmarksSWE-bench Leaderboard

Leaderboard

SWE-bench Leaderboard

Every major AI model ranked by SWE-bench — the benchmark that measures whether a model can resolve real GitHub issues, not toy problems. Pricing and context columns are included because a leaderboard you can't act on is just trivia.

Scores from provider publications and public leaderboards · Pricing verified daily

Current leader
GPT-5.6 Sol96.2%OpenAI · $2/1M input · 1.1M tokens context

Best value in the top 10: Gemini 3.7 Flash — 80.8% at $0.75/1M input tokens.

GPT-5.6 Sol vs Claude Opus 5: full comparison →Cheapest above 70%: Muse Glimmer 30B ($0.35/1M) →What's a good score? →
#ModelProviderSWE-benchInput $/1MContext
1GPT-5.6 SolOpenAI96.2%$21.1M tokens
2Claude Opus 5Anthropic96%$51M tokens
3Claude Mythos 5Anthropic95.5%$101M tokens
4Claude Fable 5Anthropic95%$101M tokens
5Claude Opus 4.8Anthropic88.6%$51M tokens
6Claude Opus 4.7Anthropic87.6%$51M tokens
7Grok 4.5xAI86.6%$2500k tokens
8Claude Sonnet 5Anthropic85.2%$21M tokens
9Claude Opus 4.6Anthropic80.8%$51M tokens
10Gemini 3.7 FlashGoogle80.8%$0.751.0M tokens
11Gemini 3.1 ProGoogle80.6%$22M tokens
12Qwen 3.7 MaxAlibaba80.4%$2.51M tokens
13GPT-5.2OpenAI80%$1.75200k tokens
14Claude Sonnet 4.6Anthropic79.6%$31M tokens
15Mistral Medium 3.5Mistral77.6%$1.5256k tokens
16Muse SparkMeta77.4%$1.251.0M tokens
17Muse Glimmer 30BMeta76%$0.35131k tokens
18GPT-5.4OpenAI74.9%$2.5272k tokens
19Grok 4xAI54%$22M tokens
20DeepSeek R1DeepSeek49.2%$0.55128k tokens
21GPT-4oOpenAI46%$2.5128k tokens
22Claude Haiku 4Anthropic43%$0.8200k tokens
23DeepSeek V3DeepSeek42%$0.27128k tokens
24Gemini 3.1 FlashGoogle35%$0.51M tokens
25Llama 4 MaverickMeta32%$0.6256k tokens
26Mistral Large 2Mistral28%$3128k tokens
27GPT-4o MiniOpenAI23.6%$0.15128k tokens
All benchmark scores →Best AI for coding →GPT-5.6 Sol alternatives →API cost calculator →What is a good SWE-bench score? →Verified vs full SWE-bench →

FAQ

What is SWE-bench?

SWE-bench is a benchmark that tests whether an AI model can resolve real GitHub issues from real open-source repositories — write a patch, pass the tests. Unlike puzzle-style benchmarks, it measures the messy, multi-file work software engineers actually do, which is why it has become the standard for comparing coding models.

Which AI model has the highest SWE-bench score?

GPT-5.6 Sol (OpenAI) currently leads at 96.2%, ahead of Claude Opus 5 at 96%.

What is a good SWE-bench score?

Anything above 70% is frontier-class in 2026 — the model can resolve most real GitHub issues autonomously. The current leader, GPT-5.6 Sol, is at 96.2%. Two years ago the best models scored under 20%, which is how fast this benchmark moves.

Does the highest SWE-bench score mean the best coding model for me?

Not always. Score-per-dollar matters for daily work: Gemini 3.7 Flash delivers 80.8% at $0.75/1M input tokens, which is the best value in the top 10. Reserve the outright leader for the hardest tasks and route volume work to the value pick.

How often is this leaderboard updated?

Scores are updated whenever providers publish new results, and the pricing columns update automatically from our daily-verified pricing data.

Newsletter

Get notified when the SWE-bench leader changes

We track new benchmark publications. When a model takes the top spot, you'll know first.

No spam. Useful updates only. Affiliate disclosures always clearly labeled.