Claude Sonnet 4.6 and Gemini 3.1 Pro are both premium models at similar prices but with different strengths. Claude leads on coding (79.6% SWE-bench, default in Cursor/Windsurf) and writing quality. Gemini 3.1 Pro wins on research depth, context window (2M vs 1M tokens), and is 33% cheaper ($2 vs $3/1M input). For developers and content teams, Claude is the stronger daily driver. For research, document analysis, and cost-sensitive workloads, Gemini is the smarter pick.
AnthropicPremium
Claude Sonnet 4.6
Best daily driver for coding and writing — the model most developers actually reach for.
VS
GooglePremium
Gemini 3.1 Pro
Best for research and deep document analysis — 2M context at the best premium price.
At a glance
Claude Sonnet 4.6
Gemini 3.1 Pro
Input cost / 1M tokens
$$3.00/1M
$$2.00/1M
Output cost / 1M tokens
$$15.00/1M
$$12.00/1M
Context window
1M tokens
2M tokens
Speed
Balanced
Balanced
Price tier
Premium
Premium
Benchmarks
SWE-bench (coding)
79.6%
80.6%
Arena Elo
1,340
1,380
MMLU
88.3%
90%
How they compare
Which model wins for each use case — and why.
CodingClaude Sonnet 4.6 wins
Claude Sonnet 4.6 scores 79.6% on SWE-bench and is the default in Cursor and Windsurf — the leading AI code editors. Gemini handles code but trails on benchmarks.
WritingClaude Sonnet 4.6 wins
Claude Sonnet 4.6 leads the writing category and consistently produces cleaner, more natural prose with better tone control than Gemini 3.1 Pro.
ResearchGemini 3.1 Pro wins
Gemini 3.1 Pro leads ARC-AGI-2 at 77.1% and has a 2M token context window for processing large research documents and corpora in a single pass.
Context WindowGemini 3.1 Pro wins
Gemini 3.1 Pro supports 2M tokens vs Claude Sonnet 4.6's 1M. For very large inputs, Gemini has twice the capacity.
PriceGemini 3.1 Pro wins
Gemini 3.1 Pro costs $2/1M input vs Claude Sonnet 4.6's $3/1M — 33% cheaper. Output is also cheaper at $12 vs $15/1M.
Which should you pick?
Pick Claude Sonnet 4.6 if…
Coding is your primary use case — Claude leads on SWE-bench and AI code editors
Writing quality, tone, and content polish matter
You use Cursor or Windsurf, which default to Claude
You want a strong all-rounder for daily developer and content work
Against Gemini 3.1 Pro it costs about 22% more per token.
The default model powering Cursor and Windsurf. 79.6% SWE-bench, 1M context window, and best-in-tier writing quality — all at $3/1M input.
Input
$3.00/1M
Output
$15.00/1M
Context
1M tokens
Speed
Balanced
What people actually use it for
Daily coding in Cursor — debugging, refactoring, and feature implementation
Drafting polished client reports, strategy memos, and long-form editorial content
Answering research questions with up to 1M tokens of document context
Where it wins
Strong coding quality with 1M context at $3/1M input
Default model in Cursor and Windsurf, the two most popular AI coding editors
Best writing quality in its price tier — tone, long-form clarity, editorial polish
Where it falls down
Claude Opus 4.7 has a higher current premium coding ceiling
GPT-5.5 or GPT-5.4 are better picks when OpenAI computer-use workflows are the priority
Skip it if
You specifically need desktop-control capabilities (GPT-5.5/GPT-5.4) or the absolute highest coding ceiling (Opus 4.7).
Our verdict
The best all-around model for most developers and writers. Strong SWE-bench, excellent writing, 1M context — all at $3/1M input. Hard to beat as a daily driver.
Against Claude Sonnet 4.6 it costs about 22% less per token and takes 2x the context.
Google's flagship with the largest context window of any frontier model at 2M tokens, Deep Think reasoning, and the best price-to-performance among premium models.
Input
$2.00/1M
Output
$12.00/1M
Context
2M tokens
Speed
Balanced
What people actually use it for
Analyzing entire contracts, codebases, or research corpora in a single 2M-token prompt
Due diligence synthesis across large sets of financial documents or legal agreements
Multi-step reasoning across dense technical specifications with Deep Think mode
Where it wins
2M token context window — the largest of any frontier model
Leads ARC-AGI-2 reasoning benchmark at 77.1%
Best price-to-performance among premium models at $2/$12 per 1M tokens
Where it falls down
Slower than Flash for everyday lightweight tasks
Claude Sonnet 4.6 is better for writing quality
Skip it if
Your primary use case is writing quality or agentic coding — Claude wins both.
Our verdict
The best research and long-context model available. Handles entire codebases, legal documents, and large datasets in a single pass — at a lower price than GPT-5.4 or Claude Sonnet 4.6.
Claude Sonnet 4.6 is better for coding and writing. Gemini 3.1 Pro is better for research, large-context work, and cost efficiency. For everyday developer use, Claude wins. For large-document analysis, Gemini wins.
Which is cheaper?
Gemini 3.1 Pro costs $2/$12 per 1M tokens vs Claude Sonnet 4.6's $3/$15. Gemini is 33% cheaper on input and 20% cheaper on output.
Which has a bigger context window?
Gemini 3.1 Pro has a 2M token context window vs Claude Sonnet 4.6's 1M — twice as large. For very long documents or codebases, Gemini has the capacity edge.