The problem with a single number
Every model on this site carries a capability score out of 100. It's tempting to read that as an objective measurement, the way a sprinter's 100-metre time is objective. It isn't. The underlying benchmarks — SWE-bench, MMLU, GPQA, ARC-AGI-2, and a handful of others — are real, and the scores on them are usually accurate to what a provider published. But a benchmark is a fixed test administered once, and a recommendation is a bet on how a model will behave across thousands of different real prompts, at a price that has to still make sense in three months.
Those are different questions, and treating the benchmark answer as the recommendation answer is how a lot of AI comparison content goes wrong. We try not to do that, which means the score you see is a weighted judgment call, not a benchmark passthrough — and we say so on the editorial policy page rather than pretending it's more scientific than it is.
What actually goes into the number
Published benchmark results are the largest single input, because they're the closest thing to a controlled comparison the industry has. But three things sit alongside them, and any one can override a benchmark lead:
- Verified, current API pricing — not a launch-day price that changed six weeks later. A model that costs twice as much for a marginal benchmark gain rarely earns the top recommendation for cost-sensitive use cases.
- Context window and max output, because a model that wins a coding benchmark but can't hold an entire repository in context loses to one that can, for real refactoring work.
- Whether the model is still there. A provider retiring a model — or an entire model family — with no warning is not a hypothetical for us; it has happened to infrastructure we run ourselves, and it changes how we weigh a model that has only been generally available for a few weeks against one with a longer track record from the same provider.
A worked example, honestly
Take coding, the category readers ask about most. GPT-6 Astra publishes a strong agentic-coding number (57.9% on Terminal-Bench 4.0) and no SWE-bench Verified or Pro figure at all — OpenAI simply didn't publish one. Claude Opus 5 has a published SWE-bench Pro figure around 79%, the single most-cited coding benchmark in the industry, but trails Astra on the agentic test OpenAI chose to highlight. Neither model 'wins' cleanly, and a site that only surfaces the benchmark each provider chose to publish would give you two contradictory answers depending on which press release you read.
Our approach is to prefer the benchmark with the widest, most independently-reproduced adoption for a given task — SWE-bench for coding specifically — while noting in the comparison itself where a model wins on a benchmark the other one didn't bother to report. That's a judgment call, not arithmetic, and it's the reason two people reading the same two spec sheets could reasonably disagree with our verdict. We'd rather show that disagreement than hide it behind a single misleadingly precise score.
The time we got it wrong
In late August 2026, our SWE-bench leaderboard — one of the more heavily trafficked pages on the site — was still naming Claude Fable 5.1 the leader at 93.4%, after GPT-5.6 Sol had already published a 96.2% figure weeks earlier. The underlying data hadn't been touched since the previous refresh, and nothing in our pipeline flagged that a new model launch had quietly made a live page wrong. It wasn't a formula problem. It was a boring data-freshness problem, and it sat there wrong for longer than it should have.
The fix wasn't clever: we backfilled seventeen rows to thirty-one, corrected five stale scores, and — more importantly — made re-verifying benchmark data part of what happens every time a new model is added to the catalog, rather than a separate task someone might forget to do. If you ever see a page on this site that looks stale relative to a model you know changed, that's the failure mode to assume, and the correction address on our editorial policy page goes straight to someone who will fix it the same way.
What we deliberately don't score
We don't score 'vibes' — writing style, personality, how a model 'feels' to talk to — because it's real but not comparable across evaluators, and turning it into a number would manufacture false precision. Where it matters (creative writing, conversational tone) we say so in prose instead of a score. We also don't let a provider's marketing materials substitute for an independent benchmark; where a provider publishes a comparison table showing its own model winning, we cross-check it against the rival's own published numbers before repeating it, and we say when we can't verify a claim independently.