UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

HomeModelsMax output tokens

Output limits

AI Max Output Tokens

The longest single response each of 38 tracked models can return, largest first, with the equivalent word count and what a full-length answer costs at list price.

Limits from the AI Gateway catalog on September 3, 2026; prices are the live list price. Words assume 0.75 words per token.

Largest max output
1M tokens
GLM-5.3
≥ 128K per response
26 models
Book-length drafts
≥ 64K per response
32 models
Long reports and code
Cheapest ≥ 64K
$0.28/1M
DeepSeek V4-Flash

Max output by model

ModelProviderMax output~WordsContextOutput $/MFull response
GLM-5.3Z.ai1M tokens~750,0001M tokens$4.40$4.40
Grok 4.5xAI500k tokens~375,000500k tokens$6.00$3.00
Grok 4.6xAI500k tokens~375,000500k tokens$6.00$3.00
DeepSeek V4-FlashDeepSeek384k tokens~288,0001M tokens$0.28$0.11
DeepSeek V4-ProDeepSeek384k tokens~288,0001M tokens$0.87$0.33
Mistral Medium 3.5Mistral256k tokens~192,000256k tokens$7.50$1.92
Muse Glimmer 30BMeta131k tokens~98,304131k tokens$1.50$0.20
Kimi K3Moonshot131k tokens~98,3041M tokens$15.00$1.97
GLM-5.3 FlashZ.ai131k tokens~98,2501M tokens$0.50$0.066
Qwen 3.8 FlashAlibaba128k tokens~96,000991k tokens$0.47$0.060
GPT-5.6 LunaOpenAI128k tokens~96,0001.1M tokens$1.20$0.15
GLM-5.2Z.ai128k tokens~96,0001M tokens$4.40$0.56
Qwen 3.8 MaxAlibaba128k tokens~96,0001M tokens$6.00$0.77
Claude Sonnet 5Anthropic128k tokens~96,0001M tokens$10.00$1.28
GPT-5.6 SolOpenAI128k tokens~96,0001.1M tokens$10.00$1.28
GPT-5.6 TerraOpenAI128k tokens~96,0001.1M tokens$12.00$1.54
GPT-5.2OpenAI128k tokens~96,000400k tokens$14.00$1.79
GPT-5.4OpenAI128k tokens~96,0001.1M tokens$15.00$1.92
Claude Sonnet 4.6Anthropic128k tokens~96,0001M tokens$15.00$1.92
Claude Opus 4.6Anthropic128k tokens~96,0001M tokens$25.00$3.20
Claude Opus 4.8Anthropic128k tokens~96,0001M tokens$25.00$3.20
Claude Opus 4.7Anthropic128k tokens~96,0001M tokens$25.00$3.20
Claude Opus 5Anthropic128k tokens~96,0001M tokens$25.00$3.20
GPT-5.5OpenAI128k tokens~96,0001M tokens$30.00$3.84
Claude Fable 5.1Anthropic128k tokens~96,0001M tokens$50.00$6.40
Claude Fable 5Anthropic128k tokens~96,0001M tokens$50.00$6.40
Gemini 3.7 FlashGoogle66k tokens~49,1521M tokens$3.75$0.25
Gemini 3.5 Flash-LiteGoogle65k tokens~48,7501M tokens$2.50$0.16
Gemini 3.6 FlashGoogle64k tokens~48,0001M tokens$3.75$0.24
Qwen 3.7 MaxAlibaba64k tokens~48,000991k tokens$7.50$0.48
Gemini 3.5 FlashGoogle64k tokens~48,0001M tokens$9.00$0.58
Gemini 3.1 ProGoogle64k tokens~48,0001M tokens$12.00$0.77
Kimi K2.7 CodeMoonshot33k tokens~24,576256k tokens$4.00$0.13
GPT-4o MiniOpenAI16k tokens~12,288128k tokens$0.60$0.010
GPT-4oOpenAI16k tokens~12,288128k tokens$10.00$0.16
Llama 4 ScoutMeta8k tokens~6,144128k tokens$1.20$0.010
Llama 4 MaverickMeta8k tokens~6,144128k tokens$1.60$0.013
DeepSeek R1DeepSeek8k tokens~6,144128k tokens$2.19$0.018

Full response = max output tokens × output price, a ceiling at list rates. Models without a published output limit are not listed.

Frequently asked questions

What does max output tokens mean?

It is the longest single response a model will return in one request, measured in tokens (about 0.75 English words each). It is separate from the context window: the context window is how much the model can read, including your prompt; the max output is how much it can write back at once. If you need more than the limit, you continue in a follow-up turn.

Which AI model has the largest max output?

GLM-5.3 (Z.ai) has the largest maximum output of the 38 models tracked here: 1M tokens per response, roughly 750,000 words. 26 models can return 128K tokens or more in one response.

What is the cheapest model for long outputs?

Among models that can return at least 64K tokens in one response, DeepSeek V4-Flash has the lowest output price at $0.28/1M tokens. A maximum-length 384k tokens response costs about $0.11 at list price. Long outputs are billed on output tokens, which are usually several times the input rate.

How much does a maximum-length response cost?

Output tokens × output price. For GLM-5.3, a full 1M tokens response costs about $4.40 at $4.4/1M. The table above shows the same figure for every model; it is a ceiling, since most responses stop well short of the limit.

Why does a model stop before reaching its max output?

Models end a response when they judge it complete, when your request sets a lower limit (max_tokens), or when the context window fills up because a long prompt plus a long answer must fit together. The published maximum is what the API allows, not what the model usually produces. If you need a long document, ask for it in sections.

Which tracked models have the smallest max output?

DeepSeek R1 has the smallest published maximum on this page at 8k tokens per response. Smaller limits are common on older or speed-focused models and are fine for chat, classification, and short drafts.

Related comparisons

Context window comparisonKnowledge cutoff datesAI cost per taskBest long-context AIAPI pricingAll AI models