GPT-3.5 Turbo
GPT-3.5 Turbo is OpenAI's legacy fast and affordable chat model, optimized for dialogue and straightforward text tasks at low cost. It was the backbone of early ChatGPT and remains a go-to for high-volume, cost-sensitive deployments.
A legacy model only worth using if your pipeline depends on the text completion API.
Legacy completion API workflows, structured text generation, and simple instruction-following tasks where the chat format is not required.
Avoid if you need multi-turn conversations, long document processing, strong reasoning, or are building any new application — modern alternatives offer far better quality at comparable cost.
Uses the legacy /v1/completions endpoint, not /v1/chat/completions. The 4,095-token context window is a hard constraint that makes it unsuitable for most modern tasks. OpenAI has not deprecated it, but it receives no capability updates.
Completion API support makes it uniquely suitable for legacy integrations and fine-tuning pipelines
Low latency for short, structured outputs like classifications or templated text
Cheaper than GPT-4o Mini for simple, high-volume completion tasks
Reliable instruction-following for well-defined, narrow tasks
Tiny 4,095-token context window severely limits document processing and long conversations
Significantly behind modern models like GPT-4o Mini, Claude Haiku 3.5, and Gemini Flash on reasoning and nuanced writing
No native chat format support; unsuitable for multi-turn conversation applications
What people actually use GPT-3.5 Turbo Instruct for.
Filling in structured templates like cover letter boilerplates or form responses
Running text classification or extraction in a high-volume batch pipeline via the completion API
Migrating or maintaining legacy OpenAI integrations built before the chat API era
The nearest models people weigh against it, and what actually separates them.
vs GPT-3.5 Turbo — Against GPT-3.5 Turbo (OpenAI), GPT-3.5 Turbo Instruct costs about 43% more per token and gives up 4x on context. GPT-3.5 Turbo is the one to check first if the price difference matters more than the ceiling.
vs GPT-3.5 Turbo (older v0613) — Against GPT-3.5 Turbo (older v0613) (OpenAI), GPT-3.5 Turbo Instruct costs about 14% more per token. GPT-3.5 Turbo (older v0613) is the one to check first if the price difference matters more than the ceiling.
vs GPT-4.1 Mini — Against GPT-4.1 Mini (OpenAI), GPT-3.5 Turbo Instruct costs about 43% more per token and gives up 255.8x on context. GPT-4.1 Mini is the one to check first if the price difference matters more than the ceiling.
Price History
→0% since May 30
90 data points · tracked daily since May 30, 2026
Legacy completion API workflows, structured text generation, and simple instruction-following tasks where the chat format is not required.. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
GPT-3.5 Turbo is OpenAI's legacy fast and affordable chat model, optimized for dialogue and straightforward text tasks at low cost. It was the backbone of early ChatGPT and remains a go-to for high-volume, cost-sensitive deployments.
An older versioned snapshot of GPT-3.5 Turbo (v0613), OpenAI's once-dominant mid-tier language model optimized for fast chat completions and instruction following. This specific checkpoint is frozen in time, predating later capability improvements introduced in subsequent GPT-3.5 Turbo updates.
GPT-4.1 Mini is OpenAI's cost-optimized small model from the GPT-4.1 family, designed to deliver strong instruction-following and coding performance at a fraction of flagship pricing. It targets high-volume, latency-sensitive applications where cost efficiency matters more than peak capability.
Pricing moves, ranking shifts, and capability updates.
OpenAI: GPT-3.5 Turbo Instruct (OpenAI) is now indexed. A legacy model only worth using if your pipeline depends on the text completion API.
View modelGPT-3.5 Turbo Instruct costs $1.5 per million input tokens and $2 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $19.00 at list price, before any batch or caching discounts.
GPT-3.5 Turbo Instruct is best for legacy completion api workflows, structured text generation, and simple instruction-following tasks where the chat format is not required.. It is a strong fit when that workflow matters more than the tradeoffs around balanced pricing and very fast speed.
Avoid if you need multi-turn conversations, long document processing, strong reasoning, or are building any new application — modern alternatives offer far better quality at comparable cost.
GPT-4.1 Mini (OpenAI) at $0.40/1M/1M input against GPT-3.5 Turbo Instruct's $1.50/1M/1M — roughly 43% less per token all in. The go-to budget workhorse for high-volume OpenAI API users who need GPT-4.1 quality at GPT-3.5 prices. Compare it first if GPT-3.5 Turbo Instruct's pricing is the thing stopping you.
GPT-3.5 Turbo — very fast against GPT-3.5 Turbo Instruct's very fast, with 16k tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.