2026-08-035D AGO · RELEASE
ALIBABABALANCED
Alibaba's 2.4T MoE flagship — a legitimate SWE-bench Pro upset over GPT-5.6 Sol (67.7 vs 64.6) and #2 globally for vision. The strongest Chinese multimodal model yet.
Input cost$2.00/M
Output cost$6.00/M
Context1M
+ PROS
- +Beats GPT-5.6 Sol on SWE-bench Pro
- +#2 globally on vision, behind only a Fable 5 variant
- +First Alibaba open-weights release at this scale
– CONS
- –No independent benchmarks at GA yet
- –Trails Fable 5 badly on the coding ceiling
VIEW FULL REPORT →2026-07-318D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
The best cheap agent engine of 2026: Terminal-Bench 82.7 at $0.14/M input, with MIT weights self-hostable in ~110 GB. Nothing touches its agentic capability per dollar.
Input cost$0.14/M
Output cost$0.28/M
Context1M
+ PROS
- +Terminal-Bench 82.7 rivals models 30x its price
- +Beats the V4-Pro preview on all nine agent benchmarks
- +2,500 concurrent requests, 1M context
– CONS
- –Text-only
- –Several headline numbers from DeepSeek's own evals
VIEW FULL REPORT →2026-07-2415D AGO · RELEASE
ANTHROPICPREMIUM
The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +SWE-bench Verified leader (~96%) as of August 2026
- +Terminal-Bench 2.1 89.1% — edges GPT-5.6 Sol
- +Within 0.5% of Fable 5 on CursorBench at half the cost
– CONS
- –Fable 5 still holds the absolute frontier ceiling
- –Fast mode doubles pricing to $10/$50
VIEW FULL REPORT →2026-07-2118D AGO · RELEASE
GOOGLEBALANCED
The best Google model for agents: beats 3.5 Flash on every tested benchmark with ~17% fewer output tokens and cheaper output. The default Gemini pick until Gemini 4 lands.
Input cost$1.50/M
Output cost$7.50/M
Context1M
+ PROS
- +OSWorld-Verified 83.0% — best computer use in class
- +~17% fewer output tokens + lower price compound
- +~280–304 tokens/sec
– CONS
- –Point release, not a generational leap
- –No Pro tier in this generation
VIEW FULL REPORT →2026-07-2118D AGO · RELEASE
GOOGLEBUDGET
The fastest model in Google's lineup at 350 tokens/sec, with real agentic chops at $0.30/M. GPT-5.6 Luna beats it on price-per-benchmark; Flash-Lite answers with speed and full multimodal input.
Input cost$0.30/M
Output cost$2.50/M
Context1M
+ PROS
- +350 tokens/sec — fastest in the 3.5 line
- +Huge jump over 3.1 Flash-Lite
- +1M context with video/audio/PDF input
– CONS
- –Trails full Flash on hard agentic work
- –Luna wins pure price-per-benchmark math
VIEW FULL REPORT →2026-07-2118D AGO · RESTRICTED
GOOGLERESTRICTEDNOT AVAILABLE
A cybersecurity fine-tune of 3.5 Flash that found 55 confirmed V8 issues vs 47 for the base model. Gated to governments and trusted partners via the CodeMender pilot — no public access or pricing.
Input cost—
Output cost—
Context1M
+ PROS
- +Beat base 3.5 Flash and Opus 4.6 at vulnerability discovery
- +Found RCE bugs in Google's own APIs within 2 hours
– CONS
- –Not publicly available — pilot only
- –No published pricing
VIEW FULL REPORT →2026-07-2019D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
The 1.6T MoE open-weights flagship: 80.6% SWE-bench Verified (self-reported) at an order of magnitude below closed-frontier pricing. Independent agentic scores land lower — but the value is still absurd.
Input cost$0.44/M
Output cost$0.87/M
Context1M
+ PROS
- +Top open-weights coding score at release
- +Codeforces 3206 — elite competitive coding
- +MIT license, 1M context, $0.87/M output
– CONS
- –Independent evals land below self-reported numbers
- –Text-only; surge pricing during Beijing hours
VIEW FULL REPORT →2026-07-1623D AGO · RELEASE
MOONSHOTPREMIUM
The closest Chinese challenger to the frontier — #4 of all models on AA Intelligence Index, ahead of Opus 4.8. At 2.8T parameters, it's the largest open-weight release in history.
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
- +#4 overall on aggregate intelligence
- +Largest open-weight model ever (2.8T, weights July 26)
- +Terminal-Bench 2.0 88.3 (Moonshot-reported)
– CONS
- –Always-on thinking burns output tokens; slow
- –Signups paused July 19 over GPU capacity
VIEW FULL REPORT →2026-07-0930D AGO · RELEASE
OPENAIBUDGET
The budget disruptor of 2026. After the ~80% July 30 price cut, Luna posts scores within 2–4 points of Sol at 1/25th the output cost. Just don't trust it past 512K tokens.
Input cost$0.20/M
Output cost$1.20/M
Context1M
+ PROS
- +GPQA Diamond 92.3% at $0.20/M input
- +1.05M context at commodity pricing
- +Within ~2–4 points of Sol on most benchmarks
– CONS
- –Long-context recall collapses past 512K (41.3%)
- –Text and image input only
VIEW FULL REPORT →2026-07-0831D AGO · RELEASE
XAIBALANCED
xAI's coding-first model on the 1.5T V9 base. Wins SWE Marathon (29% pass@1, ahead of Opus 4.8 and Fable 5) and solves SWE tasks with ~4.2x fewer tokens than Opus 4.8.
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
- +Top SWE Marathon score for long sessions
- +~60–80% cheaper per solved task than rivals
- +Exceptional token efficiency
– CONS
- –Raw ceiling trails the frontier (64.7% SWE-bench Pro)
- –500K context; price doubles at ≥200K prompt
VIEW FULL REPORT →2026-06-3039D AGO · RELEASE
ANTHROPICPREMIUM
The most agentic Sonnet yet and the new claude.ai default — 72.7% SWE-bench Verified (vs 62.3% for Sonnet 4.6) at the same price, with a $2/$10 promo through August 31.
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
- +Big agentic jump over Sonnet 4.6 at the same price
- +78.5% OSWorld-Verified computer use
- +Promo pricing $2/$10 through Aug 31, 2026
– CONS
- –20+ points behind Opus 5 on SWE-bench Verified
- –New tokenizer raises effective cost ~1.0–1.35x
VIEW FULL REPORT →2026-06-2643D AGO · RELEASE
OPENAIPREMIUM
OpenAI's flagship — Terminal-Bench 2.1 leader at 88.8% (91.9% in ultra mode with sub-agents) and the first frontier model to clear a US government review. Opus 5 still owns repo-level coding.
Input cost$5.00/M
Output cost$30.00/M
Context1M
+ PROS
- +Terminal-Bench 2.1 leader (88.8%, 91.9% ultra)
- +GPQA Diamond 94.6%, BrowseComp 90.4%
- +Ultra mode spawns sub-agents for long workflows
– CONS
- –SWE-bench Pro 64.6% vs Opus 5's 79.2%
- –Long-context surcharge $10/$45 above 272K
VIEW FULL REPORT →2026-06-2643D AGO · RELEASE
OPENAIBALANCED
The sensible OpenAI default: within 1–4 points of Sol on core benchmarks at 60% lower output cost after the July 30 price cut. Makes GPT-5.5's rate card look obsolete.
Input cost$2.00/M
Output cost$12.00/M
Context1M
+ PROS
- +~97% of Sol's benchmark line at $2/$12
- +Long-context recall nearly matches Sol
- +Cut ~20% on July 30, 2026
– CONS
- –Big gap to Sol on computer use
- –Well behind Opus 5 on repo-level coding
VIEW FULL REPORT →2026-06-1653D AGO · RELEASE
Z.AIOPEN-WEIGHTS
The top open-weights coding model of mid-2026 — beats GPT-5.5 on SWE-bench Pro at roughly a sixth of the cost, MIT-licensed, with two reasoning-effort modes.
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
- +SWE-bench Pro 62.1 — ahead of GPT-5.5
- +MIT license; third-party hosts from ~$0.75/M
- +Coding plan from ~$12.60/mo effective
– CONS
- –AA Index 51 vs Opus 5's 61 — clear frontier gap
- –No launch-day benchmarks from Z.ai itself
VIEW FULL REPORT →2026-06-1257D AGO · RELEASE
MOONSHOTBUDGET
Open-weight 1T MoE coding specialist with only 32B active params — fast, cheap to serve, and ~30% more token-efficient than its predecessor.
Input cost$0.95/M
Output cost$4.00/M
Context256K
+ PROS
- ++21.8% over K2.6 on coding evals with fewer thinking tokens
- +Modified MIT license, weights on Hugging Face
- +$0.95/M input at coding-specialist quality
– CONS
- –256K context limits big-repo agent work
- –General reasoning lags the frontier
VIEW FULL REPORT →MODEL OF THE YEAR2026-06-0960D AGO · RELEASE
ANTHROPICPREMIUM
The new global #1. 80.3% SWE-Bench Pro is an 11-point leap over Opus 4.8 (69.2%) — the biggest single-release jump of 2026. 1932 GDPval-AA, 1M context, native parallel subagents. Costs 2x Opus 4.8 ($10/$50), so reserve it for the hardest agentic and engineering work.
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
- +80.3% SWE-Bench Pro — new #1, +11 pts over Opus 4.8
- +1932 GDPval-AA, ahead of Opus 4.8 (1890)
- +Mythos-class capability, generally available
- +1M context + native parallel subagents
– CONS
- –$10/$50 — double Opus 4.8's price
- –Deliberate pace; not for latency-sensitive apps
VIEW FULL REPORT → 2026-06-0960D AGO · RESTRICTED
ANTHROPICFRONTIERNOT AVAILABLE
The same model as Fable 5 with safeguards lifted in high-risk areas — restricted to vetted partners for advanced cybersecurity and research. For everyone else, Fable 5 is identical at $10/$50 with standard safety controls.
Input cost—
Output cost—
Context1M
+ PROS
- +Tied with Fable 5 as the highest public coding score (80.3% SWE-Bench Pro)
- +Safeguards lifted for advanced security and research
- +Same 1M context + 1932 GDPval-AA as Fable 5
– CONS
- –Not generally available — vetted partners only
- –Most teams should use Fable 5 instead
VIEW FULL REPORT →2026-05-2773D AGO · RELEASE
ANTHROPICPREMIUM
The best-value premium model. 69.2% SWE-Bench Pro and 1890 Elo at $5/$25 — but Claude Fable 5 (80.3%) now leads the frontier at 2x the price, so Opus 4.8 is the smarter default for most premium work.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +69.2% SWE-Bench Pro at $5/$25 — best value at the top
- +1890 Arena Elo (67% win rate vs GPT-5.5)
- +Native parallel subagents built in
– CONS
- –Superseded by Fable 5 (80.3%) on raw coding
- –Deliberate speed — not for latency-sensitive apps
VIEW FULL REPORT →2026-05-2080D AGO · RELEASE
ALIBABABALANCED
Alibaba's first closed-weight flagship — an agent-first model built for ~35-hour autonomous runs. Cracked the global top 5 on agentic evals at launch; superseded by Qwen 3.8 Max in August.
Input cost$2.50/M
Output cost$7.50/M
Context1M
+ PROS
- +GPQA Diamond 92.4 — near the top of the field
- +Built for marathon agent sessions
- +90% cached-input discount
– CONS
- –API-only, no open weights — a break from Qwen tradition
- –Qwen 3.8 Max is stronger and cheaper
VIEW FULL REPORT →2026-05-1981D AGO · RELEASE
GOOGLEBALANCED
Google's I/O headliner: a Flash-tier model that beat Gemini 3.1 Pro on agentic benchmarks at ~278 tokens/sec. Superseded by 3.6 Flash two months later.
Input cost$1.50/M
Output cost$9.00/M
Context1M
+ PROS
- +Beat 3.1 Pro on agentic/coding despite Flash tier
- +~4x faster than comparable frontier models
- +Default model in the Gemini app
– CONS
- –3.6 Flash beats it on every tested benchmark, cheaper
- –No 3.5 Pro sibling ever shipped
VIEW FULL REPORT →2026-05-1288D AGO · RELEASE
MISTRALBALANCED
Europe's strongest open release of 2026. A clean middle option for teams that need a non-US model.
Input cost$1.00/M
Output cost$4.50/M
Context256K
+ PROS
- +EU-hosted option
- +Apache 2.0 license
- +Good speed
– CONS
- –Below frontier on coding
- –Smaller ecosystem
VIEW FULL REPORT →2026-04-30100D AGO · RELEASE
OPENAIPREMIUM
Best for agentic, computer-use, and Codex workflows. The right pick if your stack is already OpenAI-native.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +Top Terminal-Bench at 82.7%
- +Best computer-use ability
- +1M context
– CONS
- –Vision lags Opus 4.7
- –More expensive than 5.4 for marginal gains
VIEW FULL REPORT →2026-04-29101D AGO · RELEASE
MISTRALBALANCED
77.6% SWE-bench Verified from a 128B dense open-weight model with vision — within ~2 points of Sonnet 4.6 at half the price, and actually self-hostable.
Input cost$1.50/M
Output cost$7.50/M
Context256K
+ PROS
- +Strongest dense open-weights coding score at release
- +Single-checkpoint text + vision
- +Modified MIT license, EU-hosted option
– CONS
- –256K context lags the 1M frontier norm
- –Sparse benchmark disclosure at launch
VIEW FULL REPORT →2026-04-16114D AGO · RELEASE
ANTHROPICPREMIUM
Was #1 on SWE-Bench Pro at 64.3% — now superseded by Opus 4.8 (69.2%) at the same price. Vision accuracy 98.5%, strong agentic recall. Still fully supported.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +SWE-Bench Pro 64.3% — still top-tier
- +Vision accuracy 98.5%
- +1M context with sharp recall
– CONS
- –Opus 4.8 is strictly better at the same price
- –New tokenizer can raise effective cost by 35%
VIEW FULL REPORT →2026-04-11119D AGO · DISCLOSED
ANTHROPICINTERNALNOT AVAILABLE
Anthropic's most powerful internal model. Found thousands of zero-days autonomously. Not released publicly.
Input cost—
Output cost—
Context—
+ PROS
- +Frontier of frontier
- +Autonomous capabilities reported
– CONS
- –Not available — disclosed only
VIEW FULL REPORT →2026-04-08122D AGO · RELEASE
METAOPEN-WEIGHTS
600B open weights, a 5M-token context window, and Recursive Self-Improvement that re-checks its own reasoning mid-task. No major host has published verified per-token pricing yet — it joins our ranked catalog when one does.
Input cost—
Output cost—
Context5M
+ PROS
- +5M context — largest of any 2026 frontier release
- +Open weights under the Meta Open License
- +Self-corrects reasoning without external feedback
– CONS
- –No verified hosted API pricing yet
- –Needs 8x H100 minimum to self-host
VIEW FULL REPORT →2026-04-08122D AGO · RELEASE
METABALANCED
Meta's first closed frontier model — GPT-5.5-tier intelligence with the broadest multimodal input (text, image, video, audio, PDF) at $1.25/$4.25. The surprise value play of the year.
Input cost$1.25/M
Output cost$4.25/M
Context1M
+ PROS
- +#5 overall on GDPval-AA v2, ahead of Opus 4.8 on agentic tasks
- +Video, audio, and PDF input — broadest of any model
- +~$0.40/task — undercuts OpenAI and Anthropic
– CONS
- –AA Intelligence Index 54 trails Opus 5 (61) and Sol (59)
- –Public API only since July — young ecosystem
VIEW FULL REPORT →2026-03-22139D AGO · RELEASE
OPENAIPREMIUM
OpenAI's best price/quality. Pair with Claude Opus for hybrid stacks — they're complementary, not competitive.
Input cost$2.50/M
Output cost$15.00/M
Context272K
+ PROS
- +7× cheaper than Opus 4.7
- +Top-tier reasoning
- +Mature tools/agents
– CONS
- –Context capped at 272K
- –Vision lags Gemini
VIEW FULL REPORT →2026-03-04157D AGO · RELEASE
GOOGLEPREMIUM
Research workhorse. 2M context, native multimodality, and the best-priced premium model in the directory.
Input cost$1.25/M
Output cost$10.00/M
Context2M
+ PROS
- +2M context
- +Best research score
- +Best price-per-quality at premium tier
– CONS
- –Lags top tier on raw coding
- –Stuck inside Google's tooling
VIEW FULL REPORT →2026-02-18171D AGO · RELEASE
XAIBALANCED
Strong coding value at 2M context. Underrated at this price tier. The contrarian voice helps in research.
Input cost$2.00/M
Output cost$10.00/M
Context2M
+ PROS
- +2M context at $2/M input
- +Strong reasoning
- +Real-time X data integration
– CONS
- –Writing voice is uneven
- –Smaller ecosystem
VIEW FULL REPORT →2026-02-04185D AGO · RELEASE
METABUDGET
10M context is the headline. Useful for indexing entire codebases but accuracy degrades past 1M.
Input cost$0.30/M
Output cost$1.20/M
Context10M
+ PROS
- +10M context window — by far the largest
- +Open weights
- +Cheap
– CONS
- –Long-context accuracy thins out past 1M
- –Below frontier on reasoning
VIEW FULL REPORT →2026-02-04185D AGO · RELEASE
METABUDGET
Biggest open-weight leap of 2026. Competitive with GPT-5.4 on general tasks at a quarter of the price.
Input cost$0.60/M
Output cost$2.40/M
Context256K
+ PROS
- +Open weights at near-frontier quality
- +Fast
- +Strong math
– CONS
- –256K context lags Scout
- –No native multimodality
VIEW FULL REPORT →2026-01-22198D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
Open-weights, $0.27/M input, beats GPT-4o on coding. Quietly the most disruptive release of January.
Input cost$0.27/M
Output cost$1.10/M
Context128K
+ PROS
- +Cheapest serious code model
- +Open weights — self-hostable
- +Strong math
– CONS
- –Data residency questions for some teams
- –Vision is weak
VIEW FULL REPORT →2026-01-08212D AGO · RELEASE
ANTHROPICBUDGET
Fast, cheap, surprisingly capable. The cheapest model in the lineup that you can actually ship behind a feature flag.
Input cost$0.80/M
Output cost$4.00/M
Context200K
+ PROS
- +96-score speed — fastest in directory
- +Cheapest serious model at $0.80/M input
- +Vision matches mid-tier from 2025
– CONS
- –SWE-Bench Pro under 30%
- –Context capped at 200K
VIEW FULL REPORT →