UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

Home/AI Models Under $1 per Million Tokens
Best under $1/1MPrice Filter

AI Models Under $1 per Million Tokens

19 models in this directory cost $1 or less per million input tokens. Gemini 3.6 Flash is the most capable of them ($0.75/1M), Mistral Small 3.1 is the absolute cheapest at $0.1/1M, and GPT-5.6 Luna gives you the largest context window (1.05M tokens) at this price level.

Last verified Sep 3, 2026/Model data modified Sep 3, 2026
Rankings refresh dailyScored on 6 criteriaNo paid rankings
GoogleBalanced
Input cost
$0.75/1M
Context
1.0M tokens
Speed
Fast

Clear recommendation block

The safest this comparison default, the cheaper option worth trying first, and the specialist pick — before you read the detail below.

Best overall model

Gemini 3.6 Flash

View
Why this recommendation

Gemini 3.6 Flash is the strongest answer here for this comparison — pick it when quality of output matters more than the $0.75/1M/1M input you pay for it.

GoogleBalanced
Best for
Cost-efficient long-horizon agents and computer use
Price
$0.75/1M
Context
1.0M tokens
Best value model

DeepSeek V4-Pro

View
Why this recommendation

DeepSeek V4-Pro handles the same job for about 71% less per token. Start here and only move up if the output is not good enough.

DeepSeekBudget
Best for
Frontier-level coding and reasoning on a budget
Price
$0.43/1M
Context
1M tokens
Best for speed

Qwen 3.8 Flash

View
Why this recommendation

Qwen 3.8 Flash is the fastest of these for this comparison — worth it when latency is what the reader notices, not the last few points of reasoning depth.

AlibabaBudget
Best for
Cheap high-throughput coding and reasoning
Price
$0.16/1M
Context
991k tokens

Why this page recommends it

Gemini 3.6 Flash is the most capable model under $1/1M — $0.75/1M input, $3.75/1M output, 1.048576M context.

Mistral Small 3.1 is the absolute cheapest at $0.1/1M input — 100x cheaper than Claude Fable 5.1.

DeepSeek V4-Pro is the strongest budget coding pick (coding score 93/100).

Decision notes

Choose Gemini 3.6 Flash as your budget default — the best capability-per-dollar in this price band.

Choose Mistral Small 3.1 for very high-volume tasks like classification, tagging, and extraction.

Route hard tasks to a premium model and keep everything else here — a two-tier setup usually cuts spend 60–80%.

Interactive decision lab

Test the recommendation against your priority

Switch the scoring lens to see whether the this comparison answer changes when cost, speed, or long-document depth leads the decision.

#1Gemini 3.6 Flash88 pts
#2Gemini 3.7 Flash85 pts
#3DeepSeek V4-Pro83 pts
#4GPT-5.6 Luna82 pts
#5GLM-5.3 Flash82 pts
Quality first

Gemini 3.6 Flash

Google / Balanced / Sep 3, 2026

88

Best Gemini for agents — efficiency king with native computer use.

Ranks models by the broadest mix of coding, writing, research, and long-context usefulness.

Cost
$0.75/1M
$3.75/1M out
Speed
Fast
4/5 score
Context
1.0M tokens
input window
View model
Data-backed recommendation
Avoid this pick if

You need raw frontier reasoning ceiling — Claude Opus 5 and GPT-5.6 Sol lead the hardest tasks.

Recommended comparisons

Where the this comparison recommendation shifts once you weigh price or latency differently.

GoogleBalancedBest under $1/1M

Gemini 3.6 Flash

Best Gemini for agents — efficiency king with native computer use.

Best use case
Cost-efficient long-horizon agents and computer use
Input
$0.75/1M
Pricing
Balanced
Speed
Fast
Context
1.0M tokens
AgenticComputer useEfficient
DeepSeekBudgetOption 2

DeepSeek V4-Pro

Best open-weights flagship — near-frontier coding at a tenth of the price.

Best use case
Frontier-level coding and reasoning on a budget
Input
$0.43/1M
Pricing
Budget
Speed
Balanced
Context
1M tokens
Open weightsCodingReasoning
OpenAIBudgetOption 3

GPT-5.6 Luna

Best budget model from a frontier lab — near-frontier scores at commodity price.

Best use case
Cheap high-throughput summarization, drafting, and routine agent steps
Input
$0.20/1M
Pricing
Budget
Speed
Fast
Context
1.1M tokens
BudgetFastHigh volume
GoogleBalancedOption 4

Gemini 3.7 Flash

80.8% SWE-bench Verified at introductory Flash pricing.

Best use case
High-volume coding and long-context work at introductory Flash pricing
Input
$0.75/1M
Pricing
Balanced
Speed
Fast
Context
1.0M tokens
Coding1M contextFast
DeepSeekBudgetOption 5

DeepSeek V3

GPT-4o-class coding quality at under $0.30/1M — the best value in the directory.

Best use case
Coding, reasoning, and general tasks at extreme cost efficiency
Input
$0.27/1M
Pricing
Budget
Speed
Fast
Context
128k tokens
Open sourceBudgetCoding
DeepSeekBudgetOption 6

DeepSeek V4-Flash

Best agentic capability per dollar in the directory.

Best use case
High-volume agentic coding and tool-use pipelines
Input
$0.14/1M
Pricing
Budget
Speed
Fast
Context
1M tokens
Open weightsBudgetAgentic
AlibabaBudgetOption 7

Qwen 3.8 Flash

SWE-bench Pro 62.5 at sixteen cents per million input.

Best use case
Cheap high-throughput coding and reasoning
Input
$0.16/1M
Pricing
Budget
Speed
Very fast
Context
991k tokens
Open weightsBudgetCoding
Z.aiBudgetOption 8

GLM-5.3 Flash

Native vision and video, MIT weights, fifteen cents per million.

Best use case
Cheap multimodal work at scale on MIT-licensed weights
Input
$0.15/1M
Pricing
Budget
Speed
Fast
Context
1M tokens
Open weightsMultimodalBudget

Side-by-side specs

List prices and published scores — the numbers this page's pick is built from.

ModelInputOutputEst. monthContextSpeedCodingWritingResearch
Gemini 3.6 FlashGoogle$0.75/1M$3.75/1M$151.0M tokensFast928691
DeepSeek V4-ProDeepSeek$0.43/1M$0.87/1M$6.091M tokensBalanced938085
GPT-5.6 LunaOpenAI$0.20/1M$1.20/1M$4.401.1M tokensFast888584
Gemini 3.7 FlashGoogle$0.75/1M$3.75/1M$151.0M tokensFast898285

Scores out of 100 — how we evaluate models. “Est. month” is 10M in / 2M out at list price: a ceiling, no discounts.

The case for each model

Why each one is on the shortlist for this comparison, what it is genuinely good at, and where we would steer you away from it.

Gemini 3.6 Flash

Best under $1/1MGoogle

The default answer for this comparison — 92/100 on the coding axis, and the model we would start with unless the price below rules it out.

Google's efficiency-focused successor to 3.5 Flash — higher scores on every benchmark Google tested, ~17% fewer output tokens, and cheaper output pricing.

Input
$0.75/1M
Output
$3.75/1M
Context
1.0M tokens
Speed
Fast

What people actually use it for

  • Long-horizon engineering agents — DeepSWE 49% with up to 65% token reduction on long tasks
  • Native computer-use automation (83.0% OSWorld-Verified)
  • High-throughput multimodal work with video, audio, and PDF ingestion

Where it wins

  • Beats 3.5 Flash across the board: DeepSWE 49% vs 37%, MLE Bench 63.9% vs 49.7%, OSWorld-Verified 83.0% vs 78.4%
  • ~17% fewer output tokens plus $7.50/1M output — compounds into materially cheaper agent runs
  • ~280–304 tokens/sec with computer use built in as a native tool

Where it falls down

  • Point release, not a generational leap — Gemini 4 is teased but unreleased
  • Google's own lineup tops out at Flash tier for this generation; no 3.5/3.6 Pro exists

Skip it if

You need raw frontier reasoning ceiling — Claude Opus 5 and GPT-5.6 Sol lead the hardest tasks.

Our verdict

The best Google model for agents right now. Cheaper, faster, and stronger than 3.5 Flash with the best OSWorld computer-use score in its class. The default Gemini pick until Gemini 4 lands.

Full pricing, benchmark table and release notes on the Gemini 3.6 Flash page.

DeepSeek V4-Pro

DeepSeek

The value option for this comparison: about 71% less per token than Gemini 3.6 Flash, at 93/100 on coding. Worth starting here and moving up only if the output disappoints.

DeepSeek's 1.6T-parameter (49B active) MoE flagship with hybrid sparse attention — near-frontier coding and reasoning at roughly a tenth of closed-rival pricing, MIT-licensed open weights.

Input
$0.43/1M
Output
$0.87/1M
Context
1M tokens
Speed
Balanced

What people actually use it for

  • Repository-level coding — 80.6% SWE-bench Verified (self-reported), the top open-weights score at release
  • Competitive-programming-grade reasoning (Codeforces rating 3206)
  • Self-hosted frontier capability under an MIT license

Where it wins

  • 80.6% SWE-bench Verified (self-reported) — reported as tied with Gemini 3.1 Pro
  • 93.5% LiveCodeBench and Codeforces 3206 — elite competitive-coding results
  • 1M context with 384K max output at $0.87/1M output — an order of magnitude cheaper than closed frontier models

Where it falls down

  • Independent harnesses report much lower agentic scores than the self-reported numbers; trails GPT-5.6 and Opus-class on hard agentic evals
  • Peak-hour surge pricing doubles rates, a price increase is announced, and it's text-only (no vision)

Skip it if

You need vision input, verified agentic performance, or predictable pricing (surge pricing and an announced increase loom).

Our verdict

The open-weights frontier flagship of 2026. Self-reported numbers flatter it and independent agentic scores land lower, but even discounted it's the most capability per dollar in the directory's upper tier — with MIT-licensed weights.

Full pricing, benchmark table and release notes on the DeepSeek V4-Pro page.

GPT-5.6 Luna

OpenAI

Also worth a look for this comparison, at 88/100 on the coding axis.

Input
$0.20/1M
Output
$1.20/1M
Context
1.1M tokens
Speed
Fast

Best budget model from a frontier lab — near-frontier scores at commodity price. Full GPT-5.6 Luna review →

Gemini 3.7 Flash

Google

Rounds out the shortlist for this comparison at 89/100 on coding.

Input
$0.75/1M
Output
$3.75/1M
Context
1.0M tokens
Speed
Fast

80.8% SWE-bench Verified at introductory Flash pricing. Full Gemini 3.7 Flash review →

Explore related decisions

Price Filter
AI Models Under 50¢ per Million TokensEvery AI model with input pricing under 50¢ per million tokens, ranked by real…Read guide
Guide
Best Cheap AIThe cheapest AI models ranked by real value: GPT-4o Mini at $0.15/1M, Gemini Flash…Read guide
Guide
Best Cheap AI API in 2026The cheapest AI APIs ranked by actual value — DeepSeek V3 at $0.07/1M, Gemini…Read guide
Tool
AI API cost calculatorModel your monthly spend from real token prices — input and output sides both…Read guide
Pricing
AI API pricing comparisonInput and output cost per million tokens for every model, updated when providers change…Read guide
Google
Gemini 3.6 FlashBest Gemini for agents — efficiency king with native computer use.Read guide

Quick links

Browse all modelsCompare pricingView Gemini 3.6 FlashView DeepSeek V4-ProView GPT-5.6 Luna

How we evaluate AI models

UseRightAI recommendations are based on practical decision factors people actually feel in day-to-day use.

Newsletter

Get updates when ai models under $1 per million tokens changes

We email when the this comparison pick changes, when one of these models moves on price, or when something new displaces the current leader.

No spam. Useful updates only. Affiliate disclosures always clearly labeled.

FAQ

What is the best AI model under $1 per million tokens?

Gemini 3.6 Flash — $0.75/1M input tokens with the highest capability average in this price band. Best Gemini for agents — efficiency king with native computer use.

What is the cheapest AI model overall?

Mistral Small 3.1 at $0.1/1M input and $0.3/1M output. It handles ultra-high-volume classification, summarisation, and lightweight vision tasks well despite the price.

Which model under $1/1M is best for coding?

DeepSeek V4-Pro, with a coding score of 93/100 at $0.435/1M input.

What is the catch with cheap AI models?

Budget models trail flagships on hard reasoning, nuanced writing, and complex multi-step coding. Claude Fable 5.1, the current capability leader, scores 99/100 on average vs 90/100 for Gemini 3.6 Flash — use cheap models for volume, not for your hardest work.

Do output tokens cost more than input tokens?

Yes — usually 3–5x more. Gemini 3.6 Flash charges $0.75/1M input but $3.75/1M output, so long responses drive the real bill. Our API cost calculator models both sides.