Sonar Deep Research
Language Search-grounded Available Full comparison ↗ perplexity/sonar-deep-research · by Perplexity
· decoder-only-transformer
Search-grounded model. Unlike a standard base model, this offering combines an underlying LLM with live web retrieval and citations — not a directly comparable like-for-like with a non-search-grounded entry. Billed with the token pricing below plus a per-request search-context fee (see Pricing).
Pricing — 1 offering(s)
Input tokens
- $2.00 / 1M tokens (input) Current 2026-07-26 → present
Output tokens
- $8.00 / 1M tokens (output) Current 2026-07-26 → present
Citation tokens
- $2.00 / 1M tokens (citation) Current 2026-07-26 → present
Reasoning tokens
- $3.00 / 1M tokens (reasoning) Current 2026-07-26 → present
Search queries
- $5.00 / 1k requests Current 2026-07-26 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Sonar Deep Research fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Reach for Deep Research only when you actually need an exhaustive, many-source written report and can wait minutes for it. For a cited answer to a hard question at interactive speed, Sonar Reasoning Pro is far cheaper and faster. For simple lookups, Sonar. Because Deep Research bills four usage axes (input, output, citation tokens, reasoning tokens) plus a per-query fee, budget it per-report, not per-token, and cap usage where the answer doesn't warrant the spend.
Use cases
- Comprehensive topic reports and literature reviews synthesised from many sources with citations
- Market, competitive, financial, technology, or health analyses where breadth of sourcing matters more than latency
- One-shot 'research this and write it up' tasks that would otherwise be a human afternoon of searching
Limitations
- Minutes-scale latency — unusable for interactive or real-time applications
- Four billed usage axes plus a per-search-query fee make per-call cost high and hard to predict; the most expensive Sonar variant in practice
- 128K working context despite reading many more source tokens across iterations — very long individual sources can still be truncated
- Underlying base model, exact per-report source/citation counts, and max output length are not published by Perplexity
- Legacy Sonar Chat Completions surface migrating to the Agent API (support until 2026-09-27, Perplexity changelog)
- Capability ratings are qualitative from Perplexity's positioning; no independent published benchmark for the API model