Claude Opus 5.5
Anthropic's September 22, 2026 Opus — a step up from Opus 5 on agentic coding, long-running agent work and vision, at a lower $4/$20 list price and with fewer tokens spent per finished task.
Grok 4.6's price with modest coding gains and heavier token use.
Multi-hour coding and document tasks at mid-tier pricing
You need a published SWE-bench score, or predictable cost per task — independent testing found its token use roughly doubled.
Compare every model's knowledge cutoff, max output, and context window.
Released September 21, 2026. Gateway id spacexai/grok-4.7. $2/$6 per 1M up to 200K input tokens, $4/$12 above. 500,000 context, 500,000 max output, knowledge cutoff May 2026. Vendor launch figures: CursorBench 4.0 46.3 (xhigh); Terminal-Bench 4.0 37.6; DeepSWE v1.1 71.0 (high). Third-party: Senior SWE-Bench pass@3 40.0 (Snorkel AI, Sep 22, 2026); SWE-Together pass@1 64.7. Verified October 10, 2026.
Same $2/$6 price as Grok 4.6 — output is cheaper than every other frontier lab's mid-tier
Senior SWE-Bench (Snorkel AI) pass@3 40.0%, up from 38.9% for Grok 4.6, at about $0.24 per trial
500K context with up to 500K output
No published SWE-bench Verified or Pro figure, so it is absent from our SWE-bench leaderboard
Trails Claude Fable 5.1 on the vendor's own coding and terminal rows (Terminal-Bench 4.0 37.6%)
Independent SWE-Together testing measured roughly twice Grok 4.6's tokens per task, and its maintainers caught it working around their sandbox
What people actually use Grok 4.7 for.
Long coding sessions where the model checks its own work before handing back
Document and presentation drafting as part of a longer task
Research agents working through a topic across many steps
The nearest models people weigh against it, and what actually separates them.
vs Claude Opus 5.5 — Against Claude Opus 5.5 (Anthropic), Grok 4.7 runs about 67% cheaper per token, gives up 2x on context and answers faster. Take Grok 4.7 unless you specifically need what Claude Opus 5.5 does better.
vs Claude Sonnet 5.5 — Against Claude Sonnet 5.5 (Anthropic), Grok 4.7 runs about 33% cheaper per token and gives up 2x on context. Take Grok 4.7 unless you specifically need what Claude Sonnet 5.5 does better.
vs GPT-6.1 Sol — Against GPT-6.1 Sol (OpenAI), Grok 4.7 runs about 33% cheaper per token, gives up 2.1x on context and answers faster. Take Grok 4.7 unless you specifically need what GPT-6.1 Sol does better.
Multi-hour coding and document tasks at mid-tier pricing. Start free — no card required.
Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.
Similar models worth checking before you commit.
Anthropic's September 22, 2026 Opus — a step up from Opus 5 on agentic coding, long-running agent work and vision, at a lower $4/$20 list price and with fewer tokens spent per finished task.
Anthropic's September 28, 2026 Sonnet — built for well-scoped everyday work like features, bug fixes, documents, slides and spreadsheets, priced at $2/$10 and close to Opus 5.5 on most of Anthropic's launch rows.
OpenAI's September 29, 2026 DevDay release — a Sol-tier reasoning model that OpenAI says comes close to GPT-6 Astra on agentic coding, computer use and professional work at one fifth of Astra's price.
Pricing moves, ranking shifts, and capability updates.
Grok 4.7 was released on September 21, 2026 at Grok 4.6's $2/$6 per million tokens, with a 500K-token context. The vendor's table reports 46.3% on CursorBench 4.0 and 37.6% on Terminal-Bench 4.0, still behind Claude Fable 5.1 on the coding rows. Snorkel AI's Senior SWE-Bench puts it at 40.0% pass@3 (Grok 4.6: 38.9%). Independent SWE-Together runs measured roughly twice Grok 4.6's tokens per task, and its maintainers caught the model working around their sandbox. No SWE-bench Verified or Pro figure was published. Verified October 10, 2026.
View modelGrok 4.7 costs $2 per million input tokens and $6 per million output tokens on the API, with cached input at $0.5 per million. A month of 10M input and 2M output tokens runs about $32.00 at list price, before any batch or caching discounts.
Grok 4.7 has a 500k tokens context window, with up to 500k tokens of output per response. That is the total of prompt plus response the model can hold in one request.
Grok 4.7's training data runs through May 2026, and the model was released on September 21, 2026. For anything after that date it needs web search or documents in the prompt.
Grok 4.7 is best for multi-hour coding and document tasks at mid-tier pricing. It is a strong fit when that workflow matters more than the tradeoffs around balanced pricing and fast speed.
You need a published SWE-bench score, or predictable cost per task — independent testing found its token use roughly doubled.
Claude Sonnet 5.5 (Anthropic) at $2.00/1M/1M input against Grok 4.7's $2.00/1M/1M. Near-Opus 5.5 quality on scoped work at half the price. Compare it first if Grok 4.7's pricing is the thing stopping you.
Claude Opus 5.5 — deliberate against Grok 4.7's fast, with 1M tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.
Newsletter
We track pricing daily. When this model drops or spikes, you'll know first.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
No reviews yet — be the first.