A once-dominant budget model now outclassed by cheaper, smarter alternatives like GPT-4o mini.
42
Coding
50
Writing
35
Research
0
Images
78
Value
20
Long Context
Use this when
High-volume, low-complexity tasks like chatbots, classification, summarization, and simple Q&A where cost matters more than cutting-edge quality.
Skip this if
You need strong reasoning, long document handling, code generation beyond simple snippets, or high accuracy on factual tasks — modern budget alternatives like GPT-4o mini outperform it at similar cost.
Pricing
$0.50/1M in
$1.50/1M out
→0%since May 2026
Context
16k tokens
Speed
Very fast
GPT-3.5 Turbospecs & pricing
Verified Sep 4, 2026 against the AI Gateway catalog
GPT-3.5 Turbo is still available via OpenAI API and supports fine-tuning, which keeps it relevant for teams with existing trained models. However, OpenAI has deprioritized its development in favor of the GPT-4o family. Not multimodal — text only.
Extremely low cost at $0.50/$1.50 per million tokens, undercutting most modern competitors
Very fast inference speed, suitable for real-time applications and large batch jobs
Reliable for simple instruction-following, FAQ bots, and form-filling tasks
Massive community adoption means abundant fine-tuning resources and documentation
Weaknesses
Significantly lags behind GPT-4o, Claude Sonnet 4.6, and Gemini 3.1 Pro on complex reasoning, nuanced writing, and multi-step tasks
16K context window is limiting compared to modern models offering 128K–1M tokens
Prone to hallucinations and instruction drift on longer or more complex prompts
Real-world use cases
What people actually use GPT-3.5 Turbo for.
Building a customer support chatbot that handles common FAQs at scale
Batch-classifying thousands of product reviews into sentiment categories
Generating short templated email replies or form auto-fills
How GPT-3.5 Turbo compares
The nearest models people weigh against it, and what actually separates them.
vs GPT-3.5 Turbo (older v0613) — Against GPT-3.5 Turbo (older v0613) (OpenAI), GPT-3.5 Turbo runs about 33% cheaper per token and takes 4x the context. Take GPT-3.5 Turbo unless you specifically need what GPT-3.5 Turbo (older v0613) does better.
vs GPT-3.5 Turbo Instruct — Against GPT-3.5 Turbo Instruct (OpenAI), GPT-3.5 Turbo runs about 43% cheaper per token and takes 4x the context. Take GPT-3.5 Turbo unless you specifically need what GPT-3.5 Turbo Instruct does better.
vs GPT-4.1 Mini — Against GPT-4.1 Mini (OpenAI), GPT-3.5 Turbo lands within a few percent on price and gives up 63.9x on context. Which one wins depends on whether context depth or latency is your constraint.
Price History
GPT-3.5 Turbo pricing over time
→0% since May 30
90 data points · tracked daily since May 30, 2026
Ready to try it?
Start using GPT-3.5 Turbo
High-volume, low-complexity tasks like chatbots, classification, summarization, and simple Q&A where cost matters more than cutting-edge quality.. Start free — no card required.
An older versioned snapshot of GPT-3.5 Turbo (v0613), OpenAI's once-dominant mid-tier language model optimized for fast chat completions and instruction following. This specific checkpoint is frozen in time, predating later capability improvements introduced in subsequent GPT-3.5 Turbo updates.
Verdict
A once-useful workhorse now completely overshadowed by cheaper, more capable successors.
Quality score
31%
Pricing
$1.00/1M in
$2.00/1M out
Speed
Very fast
5/5 speed
Context
4k tokens
This is a pinned legacy snapshot (v0613) and may eventually be deprecated by OpenAI. The 4,095-token context window is its most significant practical limitation. OpenAI's own GPT-4o mini offers drastically more context and better quality at a comparable price — strongly consider migrating.
LegacyBudgetFastShort ContextOpenAI
Best for
High-volume, cost-sensitive text tasks like classification, summarization, and simple Q&A where bleeding-edge quality is not required.
GPT-3.5 Turbo Instruct is a legacy completion-style model from OpenAI, designed for instruction-following tasks using the older text completion API rather than the chat API. It excels at structured text generation, fill-in-the-middle tasks, and traditional NLP workflows that predate the chat paradigm.
Verdict
A legacy model only worth using if your pipeline depends on the text completion API.
Quality score
30%
Pricing
$1.50/1M in
$2.00/1M out
Speed
Very fast
5/5 speed
Context
4k tokens
Uses the legacy /v1/completions endpoint, not /v1/chat/completions. The 4,095-token context window is a hard constraint that makes it unsuitable for most modern tasks. OpenAI has not deprecated it, but it receives no capability updates.
LegacyCompletion APILow LatencyNarrow TasksOld Gen
Best for
Legacy completion API workflows, structured text generation, and simple instruction-following tasks where the chat format is not required.
GPT-4.1 Mini is OpenAI's cost-optimized small model from the GPT-4.1 family, designed to deliver strong instruction-following and coding performance at a fraction of flagship pricing. It targets high-volume, latency-sensitive applications where cost efficiency matters more than peak capability.
Verdict
The go-to budget workhorse for high-volume OpenAI API users who need GPT-4.1 quality at GPT-3.5 prices.
Quality score
65%
Pricing
$0.40/1M in
$1.60/1M out
Speed
Very fast
5/5 speed
Context
1.0M tokens
Pricing shown is $0.40 input / $1.60 output per 1M tokens. Cached input tokens are significantly cheaper. The 1M token context window is a standout feature at this price tier — few competitors match it. Supersedes GPT-4o as the recommended default for cost-conscious applications.
BudgetFastLong ContextOpenAIProduction
Best for
High-volume production workloads that need reliable GPT-4-class instruction following without flagship pricing.
GPT-3.5 Turbo costs $0.5 per million input tokens and $1.5 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $8.00 at list price, before any batch or caching discounts.
What is the context window of GPT-3.5 Turbo?
GPT-3.5 Turbo has a 16k tokens context window, with up to 4k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
What is the knowledge cutoff of GPT-3.5 Turbo?
GPT-3.5 Turbo's training data runs through September 2021, and the model was released on March 1, 2023. For anything after that date it needs web search or documents in the prompt.
What is GPT-3.5 Turbo best for?
GPT-3.5 Turbo is best for high-volume, low-complexity tasks like chatbots, classification, summarization, and simple q&a where cost matters more than cutting-edge quality.. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and very fast speed.
When should I avoid GPT-3.5 Turbo?
You need strong reasoning, long document handling, code generation beyond simple snippets, or high accuracy on factual tasks — modern budget alternatives like GPT-4o mini outperform it at similar cost.
What is a cheaper alternative to GPT-3.5 Turbo?
Claude Opus 4.5 (Anthropic) at $5.00/1M/1M input against GPT-3.5 Turbo's $0.50/1M/1M. Anthropic's most capable model delivers best-in-class reasoning and writing quality, but the steep output cost demands genuinely complex use cases to justify it. Compare it first if GPT-3.5 Turbo's pricing is the thing stopping you.
What is a faster alternative to GPT-3.5 Turbo?
GPT-3.5 Turbo (older v0613) — very fast against GPT-3.5 Turbo's very fast, with 4k tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
Get notified when GPT-3.5 Turbo pricing changes
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.