Every major model that shipped in 2026 — ranked, benchmarked, and dated. Scrub the timeline to see how the field reshaped itself month by month.
New global #1 on SWE-Bench Pro at 80.3% — an 11-point leap over Opus 4.8. Mythos-class reasoning with native parallel subagents. The frontier you can actually use.
2M context, $2/M input, strong coding. The best ratio on the chart by a clear margin.
Frontier-adjacent quality, self-hostable, no API costs. The release that mattered most for builders.
Each dot is one model. Up means better quality, right means more expensive. The top-left edge is the Pareto frontier — every dot under it is strictly dominated.
The new global #1. 80.3% SWE-Bench Pro is an 11-point leap over Opus 4.8 (69.2%) — the biggest single-release jump of 2026. 1932 GDPval-AA, 1M context, native parallel subagents. Costs 2x Opus 4.8 ($10/$50), so reserve it for the hardest agentic and engineering work.
The same model as Fable 5 with safeguards lifted in high-risk areas — restricted to vetted partners for advanced cybersecurity and research. For everyone else, Fable 5 is identical at $10/$50 with standard safety controls.
The best-value premium model. 69.2% SWE-Bench Pro and 1890 Elo at $5/$25 — but Claude Fable 5 (80.3%) now leads the frontier at 2x the price, so Opus 4.8 is the smarter default for most premium work.
Europe's strongest open release of 2026. A clean middle option for teams that need a non-US model.
Best for agentic, computer-use, and Codex workflows. The right pick if your stack is already OpenAI-native.
Was #1 on SWE-Bench Pro at 64.3% — now superseded by Opus 4.8 (69.2%) at the same price. Vision accuracy 98.5%, strong agentic recall. Still fully supported.
Anthropic's most powerful internal model. Found thousands of zero-days autonomously. Not released publicly.
OpenAI's best price/quality. Pair with Claude Opus for hybrid stacks — they're complementary, not competitive.
Research workhorse. 2M context, native multimodality, and the best-priced premium model in the directory.
Strong coding value at 2M context. Underrated at this price tier. The contrarian voice helps in research.
10M context is the headline. Useful for indexing entire codebases but accuracy degrades past 1M.
Biggest open-weight leap of 2026. Competitive with GPT-5.4 on general tasks at a quarter of the price.
Open-weights, $0.27/M input, beats GPT-4o on coding. Quietly the most disruptive release of January.
Fast, cheap, surprisingly capable. The cheapest model in the lineup that you can actually ship behind a feature flag.
Every model currently tracked, with live pricing.
Head-to-head on coding, vision, agentic, and price.
The 2026 winner by use case and budget.
Track every API price move since launch.
Raw numbers across every benchmark we cite.
Free + $20 chatbots ranked by Value Index.