GPT-3.5 Turbo
GPT-3.5 Turbo is OpenAI's legacy fast and affordable chat model, optimized for dialogue and straightforward text tasks at low cost. It was the backbone of early ChatGPT and remains a go-to for high-volume, cost-sensitive deployments.
Best budget model from a frontier lab — near-frontier scores at commodity price.
Cheap high-throughput summarization, drafting, and routine agent steps
Your workload actually uses the long context window — recall drops to 41% past 512K tokens.
Compare every model's knowledge cutoff, max output, and context window.
Fully public July 9, 2026; price cut ~80% to $0.20/$1.20 on July 30, 2026 (launched at $1/$6). Many third-party pages still show the old price.
Punches far above its price: GPQA Diamond 92.3%, SWE-bench Pro 62.7%, Terminal-Bench 2.1 84.7%
$0.20/$1.20 per 1M after the July 30, 2026 price cut — dramatically cheaper per token than Gemini 3.6 Flash
Full 1.05M-token context at budget pricing — larger than most rival small models
Long-context recall collapses at scale: 41.3% on 512K–1M token tasks vs Terra's 72.5%
Text and image input only — no video, audio, or native PDF ingestion like Gemini 3.6 Flash
What people actually use GPT-5.6 Luna for.
High-volume summarization and drafting at $0.20/1M input
Routine steps in agent pipelines where Sol/Terra would be overkill
Budget coding assistance — 62.7% SWE-bench Pro within ~2 points of Sol at 1/25th the output cost
The nearest models people weigh against it, and what actually separates them.
vs GPT-3.5 Turbo — Against GPT-3.5 Turbo (OpenAI), GPT-5.6 Luna runs about 30% cheaper per token, takes 64.1x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.
vs GPT-3.5 Turbo (older v0613) — Against GPT-3.5 Turbo (older v0613) (OpenAI), GPT-5.6 Luna runs about 53% cheaper per token, takes 256.4x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.
vs GPT-3.5 Turbo Instruct — Against GPT-3.5 Turbo Instruct (OpenAI), GPT-5.6 Luna runs about 60% cheaper per token, takes 256.4x the context and answers slower. Which one wins depends on whether context depth or latency is your constraint.
Price History
↑100% since Aug 7
38 data points · tracked daily since Aug 7, 2026
Cheap high-throughput summarization, drafting, and routine agent steps. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
GPT-3.5 Turbo is OpenAI's legacy fast and affordable chat model, optimized for dialogue and straightforward text tasks at low cost. It was the backbone of early ChatGPT and remains a go-to for high-volume, cost-sensitive deployments.
An older versioned snapshot of GPT-3.5 Turbo (v0613), OpenAI's once-dominant mid-tier language model optimized for fast chat completions and instruction following. This specific checkpoint is frozen in time, predating later capability improvements introduced in subsequent GPT-3.5 Turbo updates.
GPT-3.5 Turbo Instruct is a legacy completion-style model from OpenAI, designed for instruction-following tasks using the older text completion API rather than the chat API. It excels at structured text generation, fill-in-the-middle tasks, and traditional NLP workflows that predate the chat paradigm.
Pricing moves, ranking shifts, and capability updates.
GPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna output pricing changed from $1.20/1M to $0.60/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna input pricing changed from $0.20/1M to $0.10/1M (↓ cheaper, 50% cut).
View modelGPT-5.6 Luna costs $0.2 per million input tokens and $1.2 per million output tokens on the API, with cached input at $0.02 per million. A month of 10M input and 2M output tokens runs about $4.40 at list price, before any batch or caching discounts.
GPT-5.6 Luna has a 1.1M tokens context window, with up to 128k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
GPT-5.6 Luna's training data runs through February 16, 2026, and the model was released on July 9, 2026. For anything after that date it needs web search or documents in the prompt.
GPT-5.6 Luna is best for cheap high-throughput summarization, drafting, and routine agent steps. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.
Your workload actually uses the long context window — recall drops to 41% past 512K tokens.
GPT-5.1-Codex-Max (OpenAI) at $1.25/1M/1M input against GPT-5.6 Luna's $0.20/1M/1M. The strongest choice for serious software engineering work, provided you can absorb the output-side pricing. Compare it first if GPT-5.6 Luna's pricing is the thing stopping you.
GPT-3.5 Turbo — very fast against GPT-5.6 Luna's fast, with 16k tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.