MODEL OF THE YEAR2026-09-039D AGO · RELEASE
OPENAIPREMIUM
OpenAI's first GPT-6 model and its answer to Fable 5.1, two days later and at the same $10/$50. The new ceiling on computer use (72.6% OSWorld 2.0, ~47% less time per task than Sol), agentic coding (57.9% Terminal-Bench 4.0 vs Fable 5.1's 55.8%) and frontier math (97.6% FrontierMath Tier 4). Fable 5.1 keeps Humanity's Last Exam (65.0% vs 57.2%). No SWE-bench figure published.
Input cost$10.00/M
Output cost$50.00/M
Context1.05M
+ PROS
- +72.6% OSWorld 2.0 — the computer-use ceiling, in ~47% less time than Sol
- +57.9% Terminal-Bench 4.0, ahead of Fable 5.1 (55.8%) and Opus 5 (52.3%)
- +97.6% FrontierMath Tier 4 and 96.0% GPQA Diamond
- +1.05M context, 128K output, April 30 2026 cutoff
– CONS
- –$10/$50 — five times GPT-5.6 Sol
- –Trails Fable 5.1 on Humanity's Last Exam and the AA Intelligence Index
- –No published SWE-bench score; Enterprise access off by default
VIEW FULL REPORT → MODEL OF THE YEAR2026-09-0111D AGO · RELEASE
ANTHROPICPREMIUM
The new frontier ceiling, and the rare upgrade that costs less than what it replaces. 52.6% on Terminal-Bench-Science 0.1 against Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0 against Opus 5's 52.3%. Base pricing is unchanged at $10/$50, but cache reads fell 75% to $0.25/1M — about 25% cheaper on typical workloads and up to 45% on agentic ones. Anthropic published no SWE-bench figure at launch.
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
- +52.6% Terminal-Bench-Science 0.1 — 2.1x Fable 5's 24.7%
- +55.8% Terminal-Bench 4.0, ahead of Opus 5 (52.3%)
- +Cache reads cut 75% to $0.25/1M — cheaper than Fable 5 to run
- +Vals AI ranks it #1 of 51 overall and #1 on LiveCodeBench (90.5%)
– CONS
- –$10/$50 base — still double Claude Opus 5
- –No published SWE-bench Verified or Pro score at launch
VIEW FULL REPORT → 2026-09-0111D AGO · RESTRICTED
ANTHROPICFRONTIERNOT AVAILABLE
Fable 5.1 with the safeguards lifted — the highest agentic-coding score Anthropic has published (60.9% Terminal-Bench 4.0), restricted to vetted US cybersecurity and life-sciences organisations. Identical price, context and cache discount to Fable 5.1.
Input cost—
Output cost—
Context1M
+ PROS
- +60.9% Terminal-Bench 4.0 — Anthropic's highest published figure
- +Same $10/$50 pricing and $0.25/1M cache reads as Fable 5.1
- +Safeguards lifted for advanced security and life-sciences research
– CONS
- –Not generally available — vetted trusted access only
- –No advantage over Fable 5.1 for ordinary work
VIEW FULL REPORT →2026-08-2617D AGO · RELEASE
Z.AIOPEN-WEIGHTS
The first natively multimodal GLM-5 — a 320B-A18B MoE with vision and video built in, MIT weights, a 1M context, and a fifteen-cent price. The cheapest serious multimodal model on the market.
Input cost$0.15/M
Output cost$0.50/M
Context1M
+ PROS
- +$0.15/$0.50 — halved again through Sep 9 launch promo
- +Native vision + video, not a bolt-on
- +MIT weights, 18B active params, 1M context
– CONS
- –No published SWE-bench score
- –Volume model, not a ceiling model — frontier reasoning is elsewhere
VIEW FULL REPORT →2026-08-2617D AGO · RELEASE
ALIBABABUDGET
Alibaba's Qwen4 architecture preview: a 125B MoE with just 6B active per token that posts SWE-bench Pro 62.5 and GPQA Diamond 91.7 — roughly GLM-5.2-class coding at a tenth of the running cost.
Input cost$0.16/M
Output cost$0.47/M
Context991K
+ PROS
- +SWE-bench Pro 62.5 at $0.16/M input
- +6B active params — cheap to host, very fast
- +Open-weight Qwen3.8-Flash-Next variant previews the Qwen4 design
– CONS
- –Architecture preview, not the flagship — that's still Qwen 3.8 Max
- –Self-reported model-card numbers, no independent evals yet
VIEW FULL REPORT →2026-08-1429D AGO · RELEASE
Z.AIOPEN-WEIGHTS
The clear GLM-5.2 upgrade at the exact same price — huge agentic gains (DeepSWE 46.2 → 66.9, Terminal-Bench 3.0 4.6 → 28.3) plus a first: 84.5% on CyberGym, narrowly ahead of Claude Mythos 5's 83.8% on security.
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
- +Beats Claude Mythos 5 on CyberGym (84.5 vs 83.8)
- +Same $1.40/$4.40 price as GLM-5.2
- +Terminal-Bench 2.1 88.2, GPQA Diamond 91.7
– CONS
- –No published SWE-bench Verified score
- –Access initially routed through the GLM Coding Plan ($18/mo)
VIEW FULL REPORT →2026-08-1330D AGO · RELEASE
GOOGLEBALANCED
Google's best coding score per dollar — 80.8% SWE-bench Verified and GPQA Diamond 94.8% at half the price of 3.6 Flash, shipped three weeks after it. The catch: the price doubles to $1.50/$7.50 on January 1, 2027.
Input cost$0.75/M
Output cost$3.75/M
Context1M
+ PROS
- +80.8% SWE-bench Verified at $0.75/M input
- +GPQA Diamond 94.8% — expert reasoning near the frontier
- +Big SWE jumps over 3.6 Flash: DeepSWE 49.0 → 65.3
– CONS
- –Introductory price expires Dec 31, 2026 — then doubles
- –No SWE-bench Pro or Terminal-Bench figures published yet
- –Makes Google's own 3.6 Flash look overpriced 3 weeks after launch
VIEW FULL REPORT →2026-08-1231D AGO · RELEASE
XAIBALANCED
xAI's long-horizon agent play: finishes agent tasks in roughly half the turns of rivals, so real cost per completed task lands well under the sticker price. AA Intelligence Index 61 — 4th overall.
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
- +~2x turn efficiency on long agent runs
- +Terminal-Bench 2.1 88.4 — ahead of GLM-5.3
- +Same $2/$6 price as Grok 4.5
– CONS
- –No published SWE-bench score of any kind
- –Price doubles to $4/$12 for prompts ≥200K
- –GPT-5.6 Sol Max beats it on DeepSWE and Terminal-Bench 3.0
VIEW FULL REPORT →2026-08-0934D AGO · RELEASE
METAOPEN-WEIGHTS
Meta's return to genuine open source — the first model from Meta Superintelligence Labs, a 30B dense Apache 2.0 release built for always-on agents. 76.0% SWE-bench Verified from a model that runs on a 24GB GPU.
Input cost$0.35/M
Output cost$1.50/M
Context131K
+ PROS
- +Apache 2.0 — no commercial restrictions at all
- +76.0% SWE-bench Verified at 30B dense
- +Runs locally: 24GB GPU with quantization
– CONS
- –Qwen3.6-27B beats it on several agent tests in Meta's own table
- –131K context is small next to the 1M field
- –No first-party API — third-party hosting or self-host only
VIEW FULL REPORT →2026-08-0340D AGO · RELEASE
ALIBABABALANCED
Alibaba's 2.4T MoE flagship — a legitimate SWE-bench Pro upset over GPT-5.6 Sol (67.7 vs 64.6) and #2 globally for vision. The strongest Chinese multimodal model yet.
Input cost$2.00/M
Output cost$6.00/M
Context1M
+ PROS
- +Beats GPT-5.6 Sol on SWE-bench Pro
- +#2 globally on vision, behind only a Fable 5 variant
- +First Alibaba open-weights release at this scale
– CONS
- –No independent benchmarks at GA yet
- –Trails Fable 5 badly on the coding ceiling
VIEW FULL REPORT →2026-07-3143D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
The best cheap agent engine of 2026: Terminal-Bench 82.7 at $0.14/M input, with MIT weights self-hostable in ~110 GB. Nothing touches its agentic capability per dollar.
Input cost$0.14/M
Output cost$0.28/M
Context1M
+ PROS
- +Terminal-Bench 82.7 rivals models 30x its price
- +Beats the V4-Pro preview on all nine agent benchmarks
- +2,500 concurrent requests, 1M context
– CONS
- –Text-only
- –Several headline numbers from DeepSeek's own evals
VIEW FULL REPORT →2026-07-2450D AGO · RELEASE
ANTHROPICPREMIUM
The new premium coding default. Takes the SWE-bench Verified lead (~96%) at Opus 4.8's exact $5/$25 price, and gets within half a point of Fable 5 on agentic coding at half the cost.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +SWE-bench Verified leader (~96%) as of August 2026
- +Terminal-Bench 2.1 89.1% — edges GPT-5.6 Sol
- +Within 0.5% of Fable 5 on CursorBench at half the cost
– CONS
- –Fable 5 still holds the absolute frontier ceiling
- –Fast mode doubles pricing to $10/$50
VIEW FULL REPORT →2026-07-2153D AGO · RELEASE
GOOGLEBALANCED
The best Google model for agents: beats 3.5 Flash on every tested benchmark with ~17% fewer output tokens and cheaper output. The default Gemini pick until Gemini 4 lands.
Input cost$1.50/M
Output cost$7.50/M
Context1M
+ PROS
- +OSWorld-Verified 83.0% — best computer use in class
- +~17% fewer output tokens + lower price compound
- +~280–304 tokens/sec
– CONS
- –Point release, not a generational leap
- –No Pro tier in this generation
VIEW FULL REPORT →2026-07-2153D AGO · RELEASE
GOOGLEBUDGET
The fastest model in Google's lineup at 350 tokens/sec, with real agentic chops at $0.30/M. GPT-5.6 Luna beats it on price-per-benchmark; Flash-Lite answers with speed and full multimodal input.
Input cost$0.30/M
Output cost$2.50/M
Context1M
+ PROS
- +350 tokens/sec — fastest in the 3.5 line
- +Huge jump over 3.1 Flash-Lite
- +1M context with video/audio/PDF input
– CONS
- –Trails full Flash on hard agentic work
- –Luna wins pure price-per-benchmark math
VIEW FULL REPORT →2026-07-2153D AGO · RESTRICTED
GOOGLERESTRICTEDNOT AVAILABLE
A cybersecurity fine-tune of 3.5 Flash that found 55 confirmed V8 issues vs 47 for the base model. Gated to governments and trusted partners via the CodeMender pilot — no public access or pricing.
Input cost—
Output cost—
Context1M
+ PROS
- +Beat base 3.5 Flash and Opus 4.6 at vulnerability discovery
- +Found RCE bugs in Google's own APIs within 2 hours
– CONS
- –Not publicly available — pilot only
- –No published pricing
VIEW FULL REPORT →2026-07-2054D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
The 1.6T MoE open-weights flagship: 80.6% SWE-bench Verified (self-reported) at an order of magnitude below closed-frontier pricing. Independent agentic scores land lower — but the value is still absurd.
Input cost$0.44/M
Output cost$0.87/M
Context1M
+ PROS
- +Top open-weights coding score at release
- +Codeforces 3206 — elite competitive coding
- +MIT license, 1M context, $0.87/M output
– CONS
- –Independent evals land below self-reported numbers
- –Text-only; surge pricing during Beijing hours
VIEW FULL REPORT →2026-07-1658D AGO · RELEASE
MOONSHOTPREMIUM
The closest Chinese challenger to the frontier — #4 of all models on AA Intelligence Index, ahead of Opus 4.8. At 2.8T parameters, it's the largest open-weight release in history.
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
- +#4 overall on aggregate intelligence
- +Largest open-weight model ever (2.8T, weights July 26)
- +Terminal-Bench 2.0 88.3 (Moonshot-reported)
– CONS
- –Always-on thinking burns output tokens; slow
- –Signups paused July 19 over GPU capacity
VIEW FULL REPORT →2026-07-0965D AGO · RELEASE
OPENAIBUDGET
The budget disruptor of 2026. After the ~80% July 30 price cut, Luna posts scores within 2–4 points of Sol at 1/25th the output cost. Just don't trust it past 512K tokens.
Input cost$0.20/M
Output cost$1.20/M
Context1M
+ PROS
- +GPQA Diamond 92.3% at $0.20/M input
- +1.05M context at commodity pricing
- +Within ~2–4 points of Sol on most benchmarks
– CONS
- –Long-context recall collapses past 512K (41.3%)
- –Text and image input only
VIEW FULL REPORT →2026-07-0866D AGO · RELEASE
XAIBALANCED
xAI's coding-first model on the 1.5T V9 base. Wins SWE Marathon (29% pass@1, ahead of Opus 4.8 and Fable 5) and solves SWE tasks with ~4.2x fewer tokens than Opus 4.8.
Input cost$2.00/M
Output cost$6.00/M
Context500K
+ PROS
- +Top SWE Marathon score for long sessions
- +~60–80% cheaper per solved task than rivals
- +Exceptional token efficiency
– CONS
- –Raw ceiling trails the frontier (64.7% SWE-bench Pro)
- –500K context; price doubles at ≥200K prompt
VIEW FULL REPORT →2026-06-3074D AGO · RELEASE
ANTHROPICPREMIUM
The most agentic Sonnet yet and the new claude.ai default — 72.7% SWE-bench Verified (vs 62.3% for Sonnet 4.6) at the same price, with a $2/$10 promo through August 31.
Input cost$3.00/M
Output cost$15.00/M
Context1M
+ PROS
- +Big agentic jump over Sonnet 4.6 at the same price
- +78.5% OSWorld-Verified computer use
- +Promo pricing $2/$10 through Aug 31, 2026
– CONS
- –20+ points behind Opus 5 on SWE-bench Verified
- –New tokenizer raises effective cost ~1.0–1.35x
VIEW FULL REPORT →2026-06-2678D AGO · RELEASE
OPENAIPREMIUM
OpenAI's flagship — Terminal-Bench 2.1 leader at 88.8% (91.9% in ultra mode with sub-agents) and the first frontier model to clear a US government review. Opus 5 still owns repo-level coding.
Input cost$5.00/M
Output cost$30.00/M
Context1M
+ PROS
- +Terminal-Bench 2.1 leader (88.8%, 91.9% ultra)
- +GPQA Diamond 94.6%, BrowseComp 90.4%
- +Ultra mode spawns sub-agents for long workflows
– CONS
- –SWE-bench Pro 64.6% vs Opus 5's 79.2%
- –Long-context surcharge $10/$45 above 272K
VIEW FULL REPORT →2026-06-2678D AGO · RELEASE
OPENAIBALANCED
The sensible OpenAI default: within 1–4 points of Sol on core benchmarks at 60% lower output cost after the July 30 price cut. Makes GPT-5.5's rate card look obsolete.
Input cost$2.00/M
Output cost$12.00/M
Context1M
+ PROS
- +~97% of Sol's benchmark line at $2/$12
- +Long-context recall nearly matches Sol
- +Cut ~20% on July 30, 2026
– CONS
- –Big gap to Sol on computer use
- –Well behind Opus 5 on repo-level coding
VIEW FULL REPORT →2026-06-1688D AGO · RELEASE
Z.AIOPEN-WEIGHTS
The top open-weights coding model of mid-2026 — beats GPT-5.5 on SWE-bench Pro at roughly a sixth of the cost, MIT-licensed, with two reasoning-effort modes.
Input cost$1.40/M
Output cost$4.40/M
Context1M
+ PROS
- +SWE-bench Pro 62.1 — ahead of GPT-5.5
- +MIT license; third-party hosts from ~$0.75/M
- +Coding plan from ~$12.60/mo effective
– CONS
- –AA Index 51 vs Opus 5's 61 — clear frontier gap
- –No launch-day benchmarks from Z.ai itself
VIEW FULL REPORT →2026-06-1292D AGO · RELEASE
MOONSHOTBUDGET
Open-weight 1T MoE coding specialist with only 32B active params — fast, cheap to serve, and ~30% more token-efficient than its predecessor.
Input cost$0.95/M
Output cost$4.00/M
Context256K
+ PROS
- ++21.8% over K2.6 on coding evals with fewer thinking tokens
- +Modified MIT license, weights on Hugging Face
- +$0.95/M input at coding-specialist quality
– CONS
- –256K context limits big-repo agent work
- –General reasoning lags the frontier
VIEW FULL REPORT →MODEL OF THE YEAR2026-06-0995D AGO · RELEASE
ANTHROPICPREMIUM
The new global #1. 80.3% SWE-Bench Pro is an 11-point leap over Opus 4.8 (69.2%) — the biggest single-release jump of 2026. 1932 GDPval-AA, 1M context, native parallel subagents. Costs 2x Opus 4.8 ($10/$50), so reserve it for the hardest agentic and engineering work.
Input cost$10.00/M
Output cost$50.00/M
Context1M
+ PROS
- +80.3% SWE-Bench Pro — new #1, +11 pts over Opus 4.8
- +1932 GDPval-AA, ahead of Opus 4.8 (1890)
- +Mythos-class capability, generally available
- +1M context + native parallel subagents
– CONS
- –$10/$50 — double Opus 4.8's price
- –Deliberate pace; not for latency-sensitive apps
VIEW FULL REPORT → 2026-06-0995D AGO · RESTRICTED
ANTHROPICFRONTIERNOT AVAILABLE
The same model as Fable 5 with safeguards lifted in high-risk areas — restricted to vetted partners for advanced cybersecurity and research. For everyone else, Fable 5 is identical at $10/$50 with standard safety controls.
Input cost—
Output cost—
Context1M
+ PROS
- +Tied with Fable 5 as the highest public coding score (80.3% SWE-Bench Pro)
- +Safeguards lifted for advanced security and research
- +Same 1M context + 1932 GDPval-AA as Fable 5
– CONS
- –Not generally available — vetted partners only
- –Most teams should use Fable 5 instead
VIEW FULL REPORT →2026-05-27108D AGO · RELEASE
ANTHROPICPREMIUM
The best-value premium model. 69.2% SWE-Bench Pro and 1890 Elo at $5/$25 — but Claude Fable 5 (80.3%) now leads the frontier at 2x the price, so Opus 4.8 is the smarter default for most premium work.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +69.2% SWE-Bench Pro at $5/$25 — best value at the top
- +1890 Arena Elo (67% win rate vs GPT-5.5)
- +Native parallel subagents built in
– CONS
- –Superseded by Fable 5 (80.3%) on raw coding
- –Deliberate speed — not for latency-sensitive apps
VIEW FULL REPORT →2026-05-20115D AGO · RELEASE
ALIBABABALANCED
Alibaba's first closed-weight flagship — an agent-first model built for ~35-hour autonomous runs. Cracked the global top 5 on agentic evals at launch; superseded by Qwen 3.8 Max in August.
Input cost$2.50/M
Output cost$7.50/M
Context1M
+ PROS
- +GPQA Diamond 92.4 — near the top of the field
- +Built for marathon agent sessions
- +90% cached-input discount
– CONS
- –API-only, no open weights — a break from Qwen tradition
- –Qwen 3.8 Max is stronger and cheaper
VIEW FULL REPORT →2026-05-19116D AGO · RELEASE
GOOGLEBALANCED
Google's I/O headliner: a Flash-tier model that beat Gemini 3.1 Pro on agentic benchmarks at ~278 tokens/sec. Superseded by 3.6 Flash two months later.
Input cost$1.50/M
Output cost$9.00/M
Context1M
+ PROS
- +Beat 3.1 Pro on agentic/coding despite Flash tier
- +~4x faster than comparable frontier models
- +Default model in the Gemini app
– CONS
- –3.6 Flash beats it on every tested benchmark, cheaper
- –No 3.5 Pro sibling ever shipped
VIEW FULL REPORT →2026-05-12123D AGO · RELEASE
MISTRALBALANCED
Europe's strongest open release of 2026. A clean middle option for teams that need a non-US model.
Input cost$1.00/M
Output cost$4.50/M
Context256K
+ PROS
- +EU-hosted option
- +Apache 2.0 license
- +Good speed
– CONS
- –Below frontier on coding
- –Smaller ecosystem
VIEW FULL REPORT →2026-04-30135D AGO · RELEASE
OPENAIPREMIUM
Best for agentic, computer-use, and Codex workflows. The right pick if your stack is already OpenAI-native.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +Top Terminal-Bench at 82.7%
- +Best computer-use ability
- +1M context
– CONS
- –Vision lags Opus 4.7
- –More expensive than 5.4 for marginal gains
VIEW FULL REPORT →2026-04-29136D AGO · RELEASE
MISTRALBALANCED
77.6% SWE-bench Verified from a 128B dense open-weight model with vision — within ~2 points of Sonnet 4.6 at half the price, and actually self-hostable.
Input cost$1.50/M
Output cost$7.50/M
Context256K
+ PROS
- +Strongest dense open-weights coding score at release
- +Single-checkpoint text + vision
- +Modified MIT license, EU-hosted option
– CONS
- –256K context lags the 1M frontier norm
- –Sparse benchmark disclosure at launch
VIEW FULL REPORT →2026-04-16149D AGO · RELEASE
ANTHROPICPREMIUM
Was #1 on SWE-Bench Pro at 64.3% — now superseded by Opus 4.8 (69.2%) at the same price. Vision accuracy 98.5%, strong agentic recall. Still fully supported.
Input cost$5.00/M
Output cost$25.00/M
Context1M
+ PROS
- +SWE-Bench Pro 64.3% — still top-tier
- +Vision accuracy 98.5%
- +1M context with sharp recall
– CONS
- –Opus 4.8 is strictly better at the same price
- –New tokenizer can raise effective cost by 35%
VIEW FULL REPORT →2026-04-11154D AGO · DISCLOSED
ANTHROPICINTERNALNOT AVAILABLE
Anthropic's most powerful internal model. Found thousands of zero-days autonomously. Not released publicly.
Input cost—
Output cost—
Context—
+ PROS
- +Frontier of frontier
- +Autonomous capabilities reported
– CONS
- –Not available — disclosed only
VIEW FULL REPORT →2026-04-08157D AGO · RELEASE
METAOPEN-WEIGHTS
600B open weights, a 5M-token context window, and Recursive Self-Improvement that re-checks its own reasoning mid-task. No major host has published verified per-token pricing yet — it joins our ranked catalog when one does.
Input cost—
Output cost—
Context5M
+ PROS
- +5M context — largest of any 2026 frontier release
- +Open weights under the Meta Open License
- +Self-corrects reasoning without external feedback
– CONS
- –No verified hosted API pricing yet
- –Needs 8x H100 minimum to self-host
VIEW FULL REPORT →2026-04-08157D AGO · RELEASE
METABALANCED
Meta's first closed frontier model — GPT-5.5-tier intelligence with the broadest multimodal input (text, image, video, audio, PDF) at $1.25/$4.25. The surprise value play of the year.
Input cost$1.25/M
Output cost$4.25/M
Context1M
+ PROS
- +#5 overall on GDPval-AA v2, ahead of Opus 4.8 on agentic tasks
- +Video, audio, and PDF input — broadest of any model
- +~$0.40/task — undercuts OpenAI and Anthropic
– CONS
- –AA Intelligence Index 54 trails Opus 5 (61) and Sol (59)
- –Public API only since July — young ecosystem
VIEW FULL REPORT →2026-03-22174D AGO · RELEASE
OPENAIPREMIUM
OpenAI's best price/quality. Pair with Claude Opus for hybrid stacks — they're complementary, not competitive.
Input cost$2.50/M
Output cost$15.00/M
Context272K
+ PROS
- +7× cheaper than Opus 4.7
- +Top-tier reasoning
- +Mature tools/agents
– CONS
- –Context capped at 272K
- –Vision lags Gemini
VIEW FULL REPORT →2026-03-04192D AGO · RELEASE
GOOGLEPREMIUM
Research workhorse. 2M context, native multimodality, and the best-priced premium model in the directory.
Input cost$1.25/M
Output cost$10.00/M
Context2M
+ PROS
- +2M context
- +Best research score
- +Best price-per-quality at premium tier
– CONS
- –Lags top tier on raw coding
- –Stuck inside Google's tooling
VIEW FULL REPORT →2026-02-18206D AGO · RELEASE
XAIBALANCED
Strong coding value at 2M context. Underrated at this price tier. The contrarian voice helps in research.
Input cost$2.00/M
Output cost$10.00/M
Context2M
+ PROS
- +2M context at $2/M input
- +Strong reasoning
- +Real-time X data integration
– CONS
- –Writing voice is uneven
- –Smaller ecosystem
VIEW FULL REPORT →2026-02-04220D AGO · RELEASE
METABUDGET
10M context is the headline. Useful for indexing entire codebases but accuracy degrades past 1M.
Input cost$0.30/M
Output cost$1.20/M
Context10M
+ PROS
- +10M context window — by far the largest
- +Open weights
- +Cheap
– CONS
- –Long-context accuracy thins out past 1M
- –Below frontier on reasoning
VIEW FULL REPORT →2026-02-04220D AGO · RELEASE
METABUDGET
Biggest open-weight leap of 2026. Competitive with GPT-5.4 on general tasks at a quarter of the price.
Input cost$0.60/M
Output cost$2.40/M
Context256K
+ PROS
- +Open weights at near-frontier quality
- +Fast
- +Strong math
– CONS
- –256K context lags Scout
- –No native multimodality
VIEW FULL REPORT →2026-01-22233D AGO · RELEASE
DEEPSEEKOPEN-WEIGHTS
Open-weights, $0.27/M input, beats GPT-4o on coding. Quietly the most disruptive release of January.
Input cost$0.27/M
Output cost$1.10/M
Context128K
+ PROS
- +Cheapest serious code model
- +Open weights — self-hostable
- +Strong math
– CONS
- –Data residency questions for some teams
- –Vision is weak
VIEW FULL REPORT →2026-01-08247D AGO · RELEASE
ANTHROPICBUDGET
Fast, cheap, surprisingly capable. The cheapest model in the lineup that you can actually ship behind a feature flag.
Input cost$0.80/M
Output cost$4.00/M
Context200K
+ PROS
- +96-score speed — fastest in directory
- +Cheapest serious model at $0.80/M input
- +Vision matches mid-tier from 2025
– CONS
- –SWE-Bench Pro under 30%
- –Context capped at 200K
VIEW FULL REPORT →