UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsGuidesEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

Home/Guides/How We Actually Score AI Models (And Where the Benchmarks Lie)

Editorial · Methodology

How We Actually Score AI Models (And Where the Benchmarks Lie)

A benchmark leaderboard can tell you a model wins. It can't tell you whether the provider will still sell it to you next quarter, or whether the win matters for what you're actually building. Here's the judgment layer we add on top, and the mistake that taught us to add it.

Published September 15, 2026 7 min readBy the UseRightAI editorial team
Key takeaways
  • A published benchmark score is one input to our recommendation, not the recommendation itself — pricing, context window, speed, and provider stability all get weighed against it.
  • Deprecation risk is the input most comparison sites skip entirely, because it isn't in any benchmark table. We add it after watching an entire provider's model catalog disappear (see the field notes below).
  • We shipped a wrong leaderboard leader for a stretch in August 2026 because a benchmark row didn't get re-verified when a new model launched. The fix wasn't a smarter formula, it was a boring re-verification pass — which is now routine.

The problem with a single number

Every model on this site carries a capability score out of 100. It's tempting to read that as an objective measurement, the way a sprinter's 100-metre time is objective. It isn't. The underlying benchmarks — SWE-bench, MMLU, GPQA, ARC-AGI-2, and a handful of others — are real, and the scores on them are usually accurate to what a provider published. But a benchmark is a fixed test administered once, and a recommendation is a bet on how a model will behave across thousands of different real prompts, at a price that has to still make sense in three months.

Those are different questions, and treating the benchmark answer as the recommendation answer is how a lot of AI comparison content goes wrong. We try not to do that, which means the score you see is a weighted judgment call, not a benchmark passthrough — and we say so on the editorial policy page rather than pretending it's more scientific than it is.

What actually goes into the number

Published benchmark results are the largest single input, because they're the closest thing to a controlled comparison the industry has. But three things sit alongside them, and any one can override a benchmark lead:

  • Verified, current API pricing — not a launch-day price that changed six weeks later. A model that costs twice as much for a marginal benchmark gain rarely earns the top recommendation for cost-sensitive use cases.
  • Context window and max output, because a model that wins a coding benchmark but can't hold an entire repository in context loses to one that can, for real refactoring work.
  • Whether the model is still there. A provider retiring a model — or an entire model family — with no warning is not a hypothetical for us; it has happened to infrastructure we run ourselves, and it changes how we weigh a model that has only been generally available for a few weeks against one with a longer track record from the same provider.

A worked example, honestly

Take coding, the category readers ask about most. GPT-6 Astra publishes a strong agentic-coding number (57.9% on Terminal-Bench 4.0) and no SWE-bench Verified or Pro figure at all — OpenAI simply didn't publish one. Claude Opus 5 has a published SWE-bench Pro figure around 79%, the single most-cited coding benchmark in the industry, but trails Astra on the agentic test OpenAI chose to highlight. Neither model 'wins' cleanly, and a site that only surfaces the benchmark each provider chose to publish would give you two contradictory answers depending on which press release you read.

Our approach is to prefer the benchmark with the widest, most independently-reproduced adoption for a given task — SWE-bench for coding specifically — while noting in the comparison itself where a model wins on a benchmark the other one didn't bother to report. That's a judgment call, not arithmetic, and it's the reason two people reading the same two spec sheets could reasonably disagree with our verdict. We'd rather show that disagreement than hide it behind a single misleadingly precise score.

The time we got it wrong

In late August 2026, our SWE-bench leaderboard — one of the more heavily trafficked pages on the site — was still naming Claude Fable 5.1 the leader at 93.4%, after GPT-5.6 Sol had already published a 96.2% figure weeks earlier. The underlying data hadn't been touched since the previous refresh, and nothing in our pipeline flagged that a new model launch had quietly made a live page wrong. It wasn't a formula problem. It was a boring data-freshness problem, and it sat there wrong for longer than it should have.

The fix wasn't clever: we backfilled seventeen rows to thirty-one, corrected five stale scores, and — more importantly — made re-verifying benchmark data part of what happens every time a new model is added to the catalog, rather than a separate task someone might forget to do. If you ever see a page on this site that looks stale relative to a model you know changed, that's the failure mode to assume, and the correction address on our editorial policy page goes straight to someone who will fix it the same way.

What we deliberately don't score

We don't score 'vibes' — writing style, personality, how a model 'feels' to talk to — because it's real but not comparable across evaluators, and turning it into a number would manufacture false precision. Where it matters (creative writing, conversational tone) we say so in prose instead of a score. We also don't let a provider's marketing materials substitute for an independent benchmark; where a provider publishes a comparison table showing its own model winning, we cross-check it against the rival's own published numbers before repeating it, and we say when we can't verify a claim independently.

Keep exploring

Full editorial policy SWE-bench leaderboard All benchmark scores
Written and reviewed by the UseRightAI editorial team — not generated on demand for this page. See our editorial policy for how we source and review everything we publish, or send a correction.