UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

HomeModelsGPT-5.6 Luna
OpenAIBudget

GPT-5.6 Luna

Best budget model from a frontier lab — near-frontier scores at commodity price.

88
Coding
85
Writing
84
Research
80
Images
95
Value
58
Long Context
Use this when

Cheap high-throughput summarization, drafting, and routine agent steps

Skip this if

Your workload actually uses the long context window — recall drops to 41% past 512K tokens.

Pricing
$0.20/1M in
$1.20/1M out
↑100%since Aug 2026
Context
1.1M tokens
Speed
Fast

GPT-5.6 Lunaspecs & pricing

Verified Sep 4, 2026 against the AI Gateway catalog
Input price
$0.20 / 1M tokens
Output price
$1.20 / 1M tokens
Cached input(prompt-cache read)
$0.020 / 1M tokens
Cache write
$0.25 / 1M tokens
Context window
1.1M tokens
Max output
128k tokens
Knowledge cutoff
Feb 16, 2026
Released
Jul 9, 2026
Input modalities
Text, Image, PDF
Output modalities
Text
Reasoning mode
Yes
Tool use
Yes
Gateway model ID
openai/gpt-5.6-luna

Compare every model's knowledge cutoff, max output, and context window.

Fully public July 9, 2026; price cut ~80% to $0.20/$1.20 on July 30, 2026 (launched at $1/$6). Many third-party pages still show the old price.

How to access
Subscription
ChatGPT Plus — $20/mo
API
$0.2/1M input tokens
Subscription = chat interface. API = build with it. Compare all subscription plans · ChatGPT Plus usage limits
Switch to instead if...
Best overall
GPT-6 Astra
Cheaper option
GPT-5.1-Codex-Max
Faster option
GPT-3.5 Turbo

Strengths

Punches far above its price: GPQA Diamond 92.3%, SWE-bench Pro 62.7%, Terminal-Bench 2.1 84.7%

$0.20/$1.20 per 1M after the July 30, 2026 price cut — dramatically cheaper per token than Gemini 3.6 Flash

Full 1.05M-token context at budget pricing — larger than most rival small models

Weaknesses

Long-context recall collapses at scale: 41.3% on 512K–1M token tasks vs Terra's 72.5%

Text and image input only — no video, audio, or native PDF ingestion like Gemini 3.6 Flash

Real-world use cases

What people actually use GPT-5.6 Luna for.

High-volume summarization and drafting at $0.20/1M input

Routine steps in agent pipelines where Sol/Terra would be overkill

Budget coding assistance — 62.7% SWE-bench Pro within ~2 points of Sol at 1/25th the output cost

How GPT-5.6 Luna compares

The nearest models people weigh against it, and what actually separates them.

vs GPT-3.5 Turbo — Against GPT-3.5 Turbo (OpenAI), GPT-5.6 Luna runs about 30% cheaper per token, takes 64.1x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.

vs GPT-3.5 Turbo (older v0613) — Against GPT-3.5 Turbo (older v0613) (OpenAI), GPT-5.6 Luna runs about 53% cheaper per token, takes 256.4x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.

vs GPT-3.5 Turbo Instruct — Against GPT-3.5 Turbo Instruct (OpenAI), GPT-5.6 Luna runs about 60% cheaper per token, takes 256.4x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.

Price History

GPT-5.6 Luna pricing over time

↑100% since Aug 7

$0.216$0.185$0.154$0.123$0.092Aug 7Aug 14Aug 22Sep 5Sep 13Sep 20

38 data points · tracked daily since Aug 7, 2026

Ready to try it?

Start using GPT-5.6 Luna

Cheap high-throughput summarization, drafting, and routine agent steps. Start free — no card required.

Try GPT-5.6 Luna freeCompare alternatives

Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.

Compare alternatives

Similar models worth checking before you commit.

All GPT-5.6 Luna alternatives →
OpenAIBudget

GPT-3.5 Turbo

GPT-3.5 Turbo is OpenAI's legacy fast and affordable chat model, optimized for dialogue and straightforward text tasks at low cost. It was the backbone of early ChatGPT and remains a go-to for high-volume, cost-sensitive deployments.

Verdict
A once-dominant budget model now outclassed by cheaper, smarter alternatives like GPT-4o mini.
Quality score
35%
Pricing
$0.50/1M in
$1.50/1M out
Speed
Very fast
5/5 speed
Context
16k tokens
GPT-3.5 Turbo is still available via OpenAI API and supports fine-tuning, which keeps it relevant for teams with existing trained models. However, OpenAI has deprioritized its development in favor of the GPT-4o family. Not multimodal — text only.
BudgetLegacyFastHigh-volumeChatbot
Best for
High-volume, low-complexity tasks like chatbots, classification, summarization, and simple Q&A where cost matters more than cutting-edge quality.
View model
OpenAIBalanced

GPT-3.5 Turbo (older v0613)

An older versioned snapshot of GPT-3.5 Turbo (v0613), OpenAI's once-dominant mid-tier language model optimized for fast chat completions and instruction following. This specific checkpoint is frozen in time, predating later capability improvements introduced in subsequent GPT-3.5 Turbo updates.

Verdict
A once-useful workhorse now completely overshadowed by cheaper, more capable successors.
Quality score
31%
Pricing
$1.00/1M in
$2.00/1M out
Speed
Very fast
5/5 speed
Context
4k tokens
This is a pinned legacy snapshot (v0613) and may eventually be deprecated by OpenAI. The 4,095-token context window is its most significant practical limitation. OpenAI's own GPT-4o mini offers drastically more context and better quality at a comparable price — strongly consider migrating.
LegacyBudgetFastShort ContextOpenAI
Best for
High-volume, cost-sensitive text tasks like classification, summarization, and simple Q&A where bleeding-edge quality is not required.
View model
OpenAIBalanced

GPT-3.5 Turbo Instruct

GPT-3.5 Turbo Instruct is a legacy completion-style model from OpenAI, designed for instruction-following tasks using the older text completion API rather than the chat API. It excels at structured text generation, fill-in-the-middle tasks, and traditional NLP workflows that predate the chat paradigm.

Verdict
A legacy model only worth using if your pipeline depends on the text completion API.
Quality score
30%
Pricing
$1.50/1M in
$2.00/1M out
Speed
Very fast
5/5 speed
Context
4k tokens
Uses the legacy /v1/completions endpoint, not /v1/chat/completions. The 4,095-token context window is a hard constraint that makes it unsuitable for most modern tasks. OpenAI has not deprecated it, but it receives no capability updates.
LegacyCompletion APILow LatencyNarrow TasksOld Gen
Best for
Legacy completion API workflows, structured text generation, and simple instruction-following tasks where the chat format is not required.
View model

GPT-5.6 Luna head-to-head

All GPT-5.6 Luna alternatives →GPT-5.6 Luna vs Gemini 3.5 Flash-Lite →GPT-5.6 Luna vs DeepSeek V4-Flash →GLM-5.3 Flash vs GPT-5.6 Luna →Qwen 3.8 Flash vs GPT-5.6 Luna →View benchmark scores →

Change history

Pricing moves, ranking shifts, and capability updates.

PricingSep 1, 2026

GPT-5.6 Luna — output price cut

GPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).

View model
PricingSep 1, 2026

GPT-5.6 Luna — input price cut

GPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).

View model
PricingAug 23, 2026

GPT-5.6 Luna — output price cut

GPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).

View model
PricingAug 23, 2026

GPT-5.6 Luna — input price cut

GPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).

View model
PricingAug 7, 2026

GPT-5.6 Luna — output price cut

GPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).

View model
PricingAug 7, 2026

GPT-5.6 Luna — input price cut

GPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).

View model

FAQ

How much does GPT-5.6 Luna cost?

GPT-5.6 Luna costs $0.2 per million input tokens and $1.2 per million output tokens on the API, with cached input at $0.02 per million. A month of 10M input and 2M output tokens runs about $4.40 at list price, before any batch or caching discounts.

What is the context window of GPT-5.6 Luna?

GPT-5.6 Luna has a 1.1M tokens context window, with up to 128k tokens of output per response. That is the total of prompt plus response the model can hold in one request.

What is the knowledge cutoff of GPT-5.6 Luna?

GPT-5.6 Luna's training data runs through February 16, 2026, and the model was released on July 9, 2026. For anything after that date it needs web search or documents in the prompt.

What is GPT-5.6 Luna best for?

GPT-5.6 Luna is best for cheap high-throughput summarization, drafting, and routine agent steps. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.

When should I avoid GPT-5.6 Luna?

Your workload actually uses the long context window — recall drops to 41% past 512K tokens.

What is a cheaper alternative to GPT-5.6 Luna?

GPT-5.1-Codex-Max (OpenAI) at $1.25/1M/1M input against GPT-5.6 Luna's $0.20/1M/1M. The strongest choice for serious software engineering work, provided you can absorb the output-side pricing. Compare it first if GPT-5.6 Luna's pricing is the thing stopping you.

What is a faster alternative to GPT-5.6 Luna?

GPT-3.5 Turbo — very fast against GPT-5.6 Luna's fast, with 16k tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.

Newsletter

Get notified when GPT-5.6 Luna pricing changes

We track pricing daily. When this model drops or spikes, you'll know first.

No spam. Useful updates only. Affiliate disclosures always clearly labeled.

User reviews

No reviews yet — be the first.