Sonar
Language Search-grounded Available Full comparison ↗ perplexity/sonar · by Perplexity
· decoder-only-transformer
Search-grounded model. Unlike a standard base model, this offering combines an underlying LLM with live web retrieval and citations — not a directly comparable like-for-like with a non-search-grounded entry. Billed with the token pricing below plus a per-request search-context fee (see Pricing).
Pricing — 1 offering(s)
Input tokens
- $1.00 / 1M tokens (input) Current 2026-07-26 → present
Output tokens
- $1.00 / 1M tokens (output) Current 2026-07-26 → present
Search context fee (billed on top of token pricing, per query)
- low $5.00 / 1k requests
- medium $8.00 / 1k requests
- high $12.00 / 1k requests
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Sonar fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
The default choice within the Sonar family for cheap, fast, cited web answers to straightforward questions. Step up to Sonar Pro when you need more search results per query, longer context (200K), or better handling of complex/follow-up queries; to Sonar Reasoning Pro when the task needs step-by-step analysis; to Sonar Deep Research for exhaustive multi-source reports. Outside the Sonar family, a general LLM plus your own retrieval layer is more flexible but loses Perplexity's tuned grounding and citation formatting.
Use cases
- Quick factual lookups, topic summaries, product comparisons, and current-events questions grounded in live web sources with citations (Perplexity's own stated best-fit)
- Adding a cited, up-to-date web-answer capability to an app at the lowest price point in the Sonar family
- High-volume, latency-sensitive search answering where per-answer quality needs are modest
Limitations
- Search-grounded only — every call runs a web search and is billed a per-request search-context fee ($5-$12 / 1K requests) on top of token cost; there is no ungrounded mode
- Perplexity explicitly does not recommend it for multi-step analysis, exhaustive research, or detailed-instruction tasks
- Perplexity's current docs do not publish the underlying base model, max output tokens, or citation counts — the Llama 3.3 70B basis is from a Feb 2025 launch post and is not continuously re-confirmed
- The legacy Sonar Chat Completions surface is migrating to the Agent API (support until 2026-09-27 per the Perplexity changelog)
- Capability ratings here are qualitative, derived from Perplexity's own positioning language — no independent published benchmark for the API model specifically