Muse Glimmer 30B
Muse Glimmer 30B is the safest overall answer here when you want the strongest default instead of the lowest list price.
- Best for
- Local and self-hosted agents that run continuously
- Price
- $0.35/1M
- Context
- 131k tokens
Muse Glimmer 30B wins on coding (80 vs 54) and writing quality and price ($0.35 vs $0.5/1M input). Llama 4 Scout wins on context window (512K vs 131K). For most workflows, Muse Glimmer 30B is the stronger default — apache 2.0 agent model that runs on a 24gb gpu.
The shortest way to see the safest default, the lower-cost option, and the specialist pick before you read deeper.
Muse Glimmer 30B is the safest overall answer here when you want the strongest default instead of the lowest list price.
Mistral: Mistral Nemo is the lower-cost option to start with when you still need useful output at scale.
Llama 4 Scout is the better pick when response speed matters more than maximum reasoning depth.
Muse Glimmer 30B leads on coding with a score of 80 vs 54 for Llama 4 Scout.
Llama 4 Scout has the larger context window: 512K vs 131K for Muse Glimmer 30B.
Muse Glimmer 30B is cheaper at $0.35/1M input tokens vs $0.5/1M for Llama 4 Scout.
Muse Glimmer 30B is the safer default: it is built for local and self-hosted agents that run continuously, which covers most of what people bring to this comparison.
Switch to Llama 4 Scout when your work is mostly affordable self-hosted long-context workflows and analysis pipelines; on that narrower brief it is the better tool.
Both models serve different primary workflows — Muse Glimmer 30B for local and self-hosted agents that run continuously, Llama 4 Scout for affordable self-hosted long-context workflows and analysis pipelines — so running each where it has a clear edge often beats forcing one to do both.
Switch the scoring lens to see whether the top answer changes when you care more about cost, speed, or long-document work.
Meta / Budget / Aug 27, 2026
Apache 2.0 agent model that runs on a 24GB GPU.
Ranks models by the broadest mix of coding, writing, research, and long-context usefulness.
You need a large context window or the strongest agent scores at this size — check Qwen's 27B first.
The fastest way to see where the recommendation shifts when your priority changes.
Apache 2.0 agent model that runs on a 24GB GPU.
Best open-weight long-context option for self-hosted pipelines.
Every figure below is the provider's list price or a published capability score — the same numbers the recommendation on this page is built from.
| Model | Input | Output | Est. month | Context | Speed | Coding | Writing | Research |
|---|---|---|---|---|---|---|---|---|
| Muse Glimmer 30BMeta | $0.35/1M | $1.50/1M | $6.50 | 131k tokens | Fast | 80 | 74 | 76 |
| Llama 4 ScoutMeta | $0.50/1M | $1.20/1M | $7.40 | 512k tokens | Fast | 54 | 60 | 78 |
Capability scores are out of 100 and reflect our own weighting of published benchmarks and production signals — see how we evaluate models. “Est. month” assumes 10M input and 2M output tokens at list price, with no batch or caching discounts applied, so treat it as a ceiling.
What each one is genuinely good at, where it falls down, and the situations we would steer you away from it — not just the headline score.
Meta's return to genuine open source — a 30B dense model under Apache 2.0, built for always-on agents rather than chat, and the first release from Meta Superintelligence Labs.
You need a large context window or the strongest agent scores at this size — check Qwen's 27B first.
The best Apache 2.0 agent model you can run on consumer hardware right now. Pick it when licence freedom and local execution matter more than the last few benchmark points — otherwise Qwen3.6-27B edges it on agent tasks.
Released August 9, 2026 — the first model from Meta Superintelligence Labs and Meta's return to a genuinely permissive licence. No Meta API price; hosted rates from third parties such as Together and OpenRouter land around $0.35/$1.50. Local hardware targets: 24GB for K-Quant-17GB, 32GB for K-Quant-Dynamic, 64GB for full precision.
Long-window open-weight model that handles large document sets at a low price point.
You want a hosted solution — Gemini 3.1 Flash gives more context for roughly the same cost.
A compelling pick for self-hosted long-context pipelines — but Gemini 3.1 Flash now offers 1M context hosted at a similar price.
Worth considering for internal search, analysis, and review workflows where data sovereignty matters.
UseRightAI recommendations are based on practical decision factors people actually feel in day-to-day use.
Newsletter
Useful if you care about ranking shifts, pricing changes, or a better recommendation appearing in this decision path.
No spam. Useful updates only. Affiliate disclosures always clearly labeled.
Muse Glimmer 30B wins on more of the categories we score — coding, budget, reasoning — so it is the better default of the two. Llama 4 Scout is the better pick when your work is mostly affordable self-hosted long-context workflows and analysis pipelines. Neither is universally "better": Muse Glimmer 30B is aimed at local and self-hosted agents that run continuously, Llama 4 Scout at affordable self-hosted long-context workflows and analysis pipelines.
Muse Glimmer 30B is cheaper at $0.35/1M input and $1.5/1M output. Llama 4 Scout costs $0.5/1M input and $1.2/1M output.
Llama 4 Scout has the larger context window at 512K tokens vs Muse Glimmer 30B's 131K. For large document analysis, Llama 4 Scout is the stronger pick.
Muse Glimmer 30B is better for coding with a score of 80 vs Llama 4 Scout's 54 (out of 100). Claude Fable 5 is the overall coding leader in this directory at 100/100.
Both Muse Glimmer 30B and Llama 4 Scout have similar speed profiles — rated fast. Neither will be the bottleneck if latency is your deciding factor.
Qwen3.6-27B beats it on several practical agent and multimodal tests in Meta's own comparison table. Meta publishes no first-party API price — you pay a third-party host or run it yourself. 131K context is small next to the 1M-token field. Avoid it if you need a large context window or the strongest agent scores at this size — check Qwen's 27B first. That is the main case for looking at Llama 4 Scout instead.
Less polished than hosted frontier models on nuanced tasks. Gemini 3.1 Flash now offers 1M context at only $0.50/1M — bigger and hosted. Avoid it if you want a hosted solution — Gemini 3.1 Flash gives more context for roughly the same cost. Against Muse Glimmer 30B specifically, the gap shows up most on coding (80 vs 54).
Take a moderate workload of 10M input and 2M output tokens a month. Muse Glimmer 30B runs $6.50 (at $0.35/1M in and $1.5/1M out); Llama 4 Scout runs $7.40 (at $0.5/1M in and $1.2/1M out). The gap is small enough that price should not decide this one. Output tokens dominate the bill on both, so prompt length matters far less than response length.
Yes, and for most teams that beats picking one. A common split is Muse Glimmer 30B for local and self-hosted agents that run continuously, with Llama 4 Scout handling affordable self-hosted long-context workflows and analysis pipelines. Since Muse Glimmer 30B is both the stronger and the cheaper option here, a split mainly makes sense if Llama 4 Scout covers a capability you specifically need.