Llama 4 Scout
Language Available Full comparison ↗ meta/llama-4-scout · by Meta
· mixture-of-experts
Pricing — 1 offering(s)
Input tokens
- $0.10 / 1M tokens (input) Historical 2025-04-05 → 2026-07-21
- $0.18 / 1M tokens (input) Current 2026-07-22 → present
Output tokens
- $0.30 / 1M tokens (output) Historical 2025-04-05 → 2026-07-21
- $0.59 / 1M tokens (output) Current 2026-07-22 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Llama 4 Scout fits into a cost-aware routing setup
See how →Capability profile
reasoning moderate
coding moderate
tool use moderate
instruction following strong
context window strong
multilingual moderate
speed strong
cost efficiency strong
Operator guidance
Route here specifically for workloads that need a very long context window and cannot afford frontier closed-model pricing. The 10M-token context is the primary differentiator — if your task doesn't need it, Llama 4 Maverick or GPT-4o mini are better-rounded alternatives at similar cost.
Use cases
- Ultra-long-context tasks: entire codebases, legal corpora, full books
- Multimodal workflows (image + text) at minimal cost
- Open-weights hosting where context length is the primary constraint
Limitations
- Newer model; fewer production deployments than Llama 3.x series
- Capability ratings are qualitative, not a single cited benchmark run
Citations
Compare models
ComparingLlama 4 Scoutwith
Pick a model above to see the comparison.