Kimi K3
Kimi K3 is the safest overall answer here when you want the strongest default instead of the lowest list price.
- Best for
- Frontier-level reasoning and agentic coding
- Price
- $3.00/1M
- Context
- 1M tokens
Kimi K3 is Moonshot's best model for long-context work — it scores 93/100 vs 70/100 for Kimi K2.7 Code, at $3/1M input tokens. Across all providers, Claude Fable 5 still leads long-context work at 99/100 — worth considering if you're not committed to Moonshot.
The shortest way to see the safest default, the lower-cost option, and the specialist pick before you read deeper.
Kimi K3 is the safest overall answer here when you want the strongest default instead of the lowest list price.
Meta: Llama 3.1 8B Instruct is the lower-cost option to start with when you still need useful output at scale.
Kimi K2.7 Code is the better pick when large documents, transcripts, or knowledge-heavy work lead the decision.
Kimi K3 leads Moonshot's lineup for long-context work at 93/100 ($3/1M input, 1M context).
Kimi K2.7 Code is the value pick at $0.95/1M input with a long-context work score of 70/100.
Claude Fable 5 (Anthropic) is the overall long-context work leader at 99/100 if provider choice is open.
Choose Kimi K3 when long-context work quality is the priority and you're staying on Moonshot.
Choose Kimi K2.7 Code when token volume matters more than peak quality.
Teams open to other providers should also evaluate Claude Fable 5 before committing.
Switch the scoring lens to see whether the top answer changes when you care more about cost, speed, or long-document work.
Moonshot / Premium / Aug 6, 2026
Closest Chinese challenger to the frontier — #4 overall on intelligence.
Ranks models by the broadest mix of coding, writing, research, and long-context usefulness.
You need fast responses or predictable output costs — always-on thinking burns tokens.
The fastest way to see where the recommendation shifts when your priority changes.
Closest Chinese challenger to the frontier — #4 overall on intelligence.
Value coding specialist — 1T MoE agentic coder at budget prices.
Every figure below is the provider's list price or a published capability score — the same numbers the recommendation on this page is built from.
| Model | Input | Output | Est. month | Context | Speed | Coding | Writing | Research |
|---|---|---|---|---|---|---|---|---|
| Kimi K3Moonshot | $3.00/1M | $15.00/1M | $60 | 1M tokens | Deliberate | 96 | 90 | 93 |
| Kimi K2.7 CodeMoonshot | $0.95/1M | $4.00/1M | $18 | 256k tokens | Fast | 88 | 68 | 70 |
Capability scores are out of 100 and reflect our own weighting of published benchmarks and production signals — see how we evaluate models. “Est. month” assumes 10M input and 2M output tokens at list price, with no batch or caching discounts applied, so treat it as a ceiling.
What each one is genuinely good at, where it falls down, and the situations we would steer you away from it — not just the headline score.
Moonshot's 2.8-trillion-parameter multimodal reasoning flagship with always-on thinking — the largest open-weight model ever released and the closest Chinese challenger to the Western frontier.
You need fast responses or predictable output costs — always-on thinking burns tokens.
The first Chinese model to genuinely crowd the Western frontier — #4 on aggregate intelligence ahead of Opus 4.8. The always-on thinking makes it slow and output-heavy, so cost per task runs above the sticker price. A serious Opus-class alternative if latency isn't critical.
Released July 16, 2026; open weights July 26. Cache-hit input $0.30/1M. Subscriptions: Adagio (free) to Vivace $199/mo; full 1M context only on Allegro ($99) and up. New signups paused July 19 near GPU capacity, reopening in batches.
An open-weight 1T-parameter MoE (32B active) coding specialist tuned for long-horizon agentic software engineering with markedly better token efficiency than its predecessor.
Your agent needs big-repo context (256K cap) or frontier general reasoning.
The value pick among coding specialists. K3 superseded it at the frontier a month later, but for pure coding-agent volume at a quarter of K3's input price, K2.7 Code remains the smarter buy.
Model ID kimi-k2.7-code; weights on Hugging Face June 12, 2026. Kimi Code membership from $19/mo.
UseRightAI recommendations are based on practical decision factors people actually feel in day-to-day use.
Newsletter
Useful if you care about ranking shifts, pricing changes, or a better recommendation appearing in this decision path.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
Kimi K3 — it scores 93/100 on long-context work in this directory, ahead of Kimi K2.7 Code at 70/100. Closest Chinese challenger to the frontier — #4 overall on intelligence.
Not overall. Claude Fable 5 (Anthropic) leads the directory for long-context work at 99/100 vs Kimi K3's 93/100. Kimi K3 is the best pick if you're staying within Moonshot's ecosystem.
Kimi K2.7 Code at $0.95/1M input tokens (long-context work score: 70/100). Use it for volume work and reserve Kimi K3 for the tasks where quality matters most.
$3/1M input tokens and $15/1M output tokens via the API, or through Kimi Moderato at $19/mo for chat use. Context window: 1M tokens. On a moderate month — 10M input and 2M output tokens — that works out to about $60.00, against $17.50 for Kimi K2.7 Code.
Most expensive Chinese-lab model ever ($3/$15) with always-on thinking driving high output-token burn and slow responses. 2.8T size makes self-hosting impractical despite open weights; consumer signups were paused July 19 over GPU capacity. Concretely, avoid it if you need fast responses or predictable output costs — always-on thinking burns tokens. If none of that is negotiable, Claude Fable 5 (Anthropic) is the cross-provider leader at 99/100.
hardest reasoning tasks — #4 of all models on AA Intelligence Index v4.1 (57.1), ahead of Claude Opus 4.8, agentic coding at 81.2 FrontierSWE and 88.3 Terminal-Bench 2.0 (Moonshot-reported), and 1M-context research synthesis with always-on extended thinking. Its 1M-token context window is the practical limit on how much you can hand it in one go.
Kimi K3 scores 93/100 on long-context work against 70/100 for Kimi K2.7 Code, at 3x the input price. That premium is worth it on work where a wrong answer costs real time or money, and hard to justify on high-volume, low-stakes calls. Most teams run both and route by task rather than picking one.