UseRightAI
HomeModelsAsk AIComparePricingWhat's New
UseRightAICut through AI hype. Pick what works.

Independent AI model tracker. Live pricing, real benchmarks, zero vendor bias.

X (Twitter)LinkedInUpdatesContact

Compare

Opus 4.8 vs Opus 4.7Fable 5 vs Opus 4.8New AI Models 2026ChatGPT vs ClaudeGPT-4o vs Claude SonnetClaude vs GeminiDeepSeek vs ChatGPTMistral vs ClaudeGemini Flash vs GPT-4o MiniLlama vs ChatGPTAll comparisons →Build your own →

Best For

CodingWritingDevelopersProduct ManagersDesignersSalesBest Cheap AIBest Free AI

Pricing & Data

API Token PricingCost per TaskPrice HistoryBenchmark ScoresPrivacy & SafetySubscription PlansPlan Usage LimitsCost CalculatorWhich AI is Cheapest?Cheapest AI APIs

Company

About UseRightAIContactWhat ChangedAll ModelsEditorial PolicyDisclosuresPrivacy PolicyTerms of Service

© 2026 UseRightAI. Independent · Free forever · Not affiliated with any AI provider.

Affiliate links are clearly labeled. See disclosures.

HomeModelsVoxtral Small 24B 2507
MistralBudget

Voxtral Small 24B 2507

A purpose-built budget audio model that excels at voice tasks but stumbles on context length and general-purpose depth.

52
Coding
63
Writing
58
Research
0
Images
88
Value
28
Long Context
Use this when

Transcribing, analyzing, and responding to audio input cost-effectively without needing a separate speech-to-text pipeline.

Skip this if

You need long-context document analysis, image understanding, or top-tier reasoning performance, as the 32K window and 24B scale will bottleneck complex tasks.

Pricing
$0.10/1M in
$0.30/1M out
→0%since May 2026
Context
32k tokens
Speed
Fast

Voxtral Small is audio-in capable but does not support image input. The 32K context window is notably short for a 2025 model. Pricing is via Mistral's API; availability through third-party providers may vary. Check whether your use case requires audio input — the text-only version of Mistral Small 3.1 may be more appropriate for pure text workloads.

How to access
API
$0.09999999999999999/1M input tokens
Subscription = chat interface. API = build with it. Compare all subscription plans
Switch to instead if...
Best overall
GPT-6 Astra
Cheaper option
Mistral Small 3.1
Faster option
Mistral Medium 3.1

Strengths

Native audio input support — understands spoken language directly without requiring external STT preprocessing

Extremely low cost at $0.10/$0.30 per 1M tokens, undercutting GPT-4o Audio and Gemini 1.5 Flash significantly

Solid multilingual audio handling, reflecting Mistral's strong European language coverage

Compact 24B size keeps inference fast despite multimodal capabilities

Weaknesses

32K context window is restrictive compared to competitors — Gemini 3.1 Pro offers 1M+ tokens, limiting long audio or document tasks

No image understanding despite the multimodal framing, narrowing its real-world versatility

Reasoning and complex coding benchmarks lag behind GPT-4o and Claude Sonnet 4.6 at their respective tiers

Real-world use cases

What people actually use Voxtral Small 24B 2507 for.

Transcribing and summarizing customer support call recordings in multiple languages

Building a budget voice assistant that processes spoken queries and returns structured text responses

Extracting action items from meeting audio files without a separate STT service

How Voxtral Small 24B 2507 compares

The nearest models people weigh against it, and what actually separates them.

vs Mistral Medium 3.1 — Against Mistral Medium 3.1 (Mistral), Voxtral Small 24B 2507 runs about 83% cheaper per token and gives up 4.1x on context. Take Voxtral Small 24B 2507 unless you specifically need what Mistral Medium 3.1 does better.

vs Llama 3.2 11B Vision Instruct — Against Llama 3.2 11B Vision Instruct (Meta), Voxtral Small 24B 2507 runs about 42% cheaper per token and gives up 4.1x on context. Take Voxtral Small 24B 2507 unless you specifically need what Llama 3.2 11B Vision Instruct does better.

vs Mistral Large 3 2512 — Against Mistral Large 3 2512 (Mistral), Voxtral Small 24B 2507 runs about 80% cheaper per token, gives up 8.2x on context and answers faster. Take Voxtral Small 24B 2507 unless you specifically need what Mistral Large 3 2512 does better.

Price History

Voxtral Small 24B 2507 pricing over time

→0% since May 31

$0.108$0.104$0.100$0.096$0.092May 31Jun 18Jul 10Jul 27Aug 14Sep 8

90 data points · tracked daily since May 31, 2026

Ready to try it?

Start using Voxtral Small 24B 2507

Transcribing, analyzing, and responding to audio input cost-effectively without needing a separate speech-to-text pipeline.. Start free — no card required.

Try Voxtral Small 24B 2507 freeCompare alternatives

Recommendations are made independently based on real-world use and public benchmarks. See our disclosures for details.

Compare alternatives

Similar models worth checking before you commit.

All Voxtral Small 24B 2507 alternatives →
MistralBudget

Mistral Medium 3.1

Mistral Medium 3.1 is a multimodal mid-tier model from Mistral that supersedes Mistral Large 2, offering vision capabilities alongside strong text performance at a significantly reduced price point. It targets the sweet spot between budget models and expensive flagships, with a 128K context window and competitive multilingual support.

Verdict
The best Mistral model for budget-conscious builders who still need multimodal capability and solid multilingual output.
Quality score
70%
Pricing
$0.40/1M in
$2.00/1M out
Speed
Fast
4/5 speed
Context
131k tokens
Officially supersedes Mistral Large 2, representing a generational shift in Mistral's lineup toward multimodal capability at lower cost tiers. Available via Mistral API and select cloud providers. No function calling limitations noted at this tier.
BudgetMultimodalMultilingualMid-tierVision
Best for
Cost-sensitive teams needing solid coding, instruction-following, and basic vision tasks without paying flagship prices.
View model
MetaBudget

Llama 3.2 11B Vision Instruct

Llama 3.2 11B Vision Instruct is Meta's open-weight multimodal model capable of understanding both text and images at an extremely low price point. It handles image captioning, visual question answering, and document analysis alongside standard text tasks.

Verdict
The go-to vision model when budget is the top constraint and good-enough accuracy is acceptable.
Quality score
57%
Pricing
$0.34/1M in
$0.34/1M out
Speed
Fast
4/5 speed
Context
131k tokens
Available via multiple inference providers including Together AI, Fireworks, and OpenRouter. As an open-weight model, it can also be self-hosted for even lower marginal costs at scale. Part of Meta's Llama 3.2 family which also includes a 90B vision variant for heavier workloads.
Open-weightVisionBudgetMultimodalMeta
Best for
Budget-conscious developers who need basic vision capabilities without paying premium multimodal prices.
View model
MistralBudget

Mistral Large 3 2512

Mistral Large 3 2512 is Mistral's flagship dense model updated in December 2025, offering strong multilingual reasoning and coding capabilities at a significantly reduced price point compared to its predecessor. It targets enterprise workloads that need high-quality outputs without paying top-tier frontier model prices.

Verdict
The best price-per-quality ratio in the non-mini flagship tier, especially for multilingual and long-context enterprise tasks.
Quality score
69%
Pricing
$0.50/1M in
$1.50/1M out
Speed
Balanced
3/5 speed
Context
262k tokens
Pricing of $0.50 input / $1.50 output per 1M tokens places it firmly in the budget-flagship category. Available via Mistral API (La Plateforme) and major cloud providers. December 2025 update ('2512') improves instruction following over the earlier 2407 release.
Budget flagshipMultilingualLong contextEnterpriseCode
Best for
Multilingual enterprise tasks, code generation, and long-document analysis where cost efficiency matters more than absolute state-of-the-art performance.
View model

Change history

Pricing moves, ranking shifts, and capability updates.

New ModelMar 27, 2026

Mistral: Voxtral Small 24B 2507 — added to UseRightAI

Mistral: Voxtral Small 24B 2507 (Mistral) is now indexed. It supersedes Mistral Small 3.1. A purpose-built budget audio model that excels at voice tasks but stumbles on context length and general-purpose depth.

View model

FAQ

How much does Voxtral Small 24B 2507 cost?

Voxtral Small 24B 2507 costs $0.09999999999999999 per million input tokens and $0.3 per million output tokens on the API. A month of 10M input and 2M output tokens runs about $1.60 at list price, before any batch or caching discounts.

What is Voxtral Small 24B 2507 best for?

Voxtral Small 24B 2507 is best for transcribing, analyzing, and responding to audio input cost-effectively without needing a separate speech-to-text pipeline.. It is a strong fit when that workflow matters more than the tradeoffs around budget pricing and fast speed.

When should I avoid Voxtral Small 24B 2507?

You need long-context document analysis, image understanding, or top-tier reasoning performance, as the 32K window and 24B scale will bottleneck complex tasks.

What is a cheaper alternative to Voxtral Small 24B 2507?

Mistral Small 3.1 (Mistral) at $0.10/1M/1M input against Voxtral Small 24B 2507's $0.10/1M/1M. Ultra-cheap multimodal model for massive-volume, low-complexity pipelines. Compare it first if Voxtral Small 24B 2507's pricing is the thing stopping you.

What is a faster alternative to Voxtral Small 24B 2507?

Mistral Medium 3.1 — fast against Voxtral Small 24B 2507's fast, with 131k tokens of context. Worth the swap when response time is what your users notice rather than the last few points of reasoning depth.

Newsletter

Get notified when Voxtral Small 24B 2507 pricing changes

We track pricing daily. When this model drops or spikes, you'll know first.

No spam. Useful updates only. Affiliate disclosures always clearly labeled.

User reviews

No reviews yet — be the first.