DeepSeek R1
Language Retired Full comparison ↗ deepseek/deepseek-r1 · by DeepSeek
· decoder-only-transformer
Pricing — 1 offering(s)
Input tokens
Deprecated — pricing unavailable
- $0.55 / 1M tokens (input) Historical 2025-01-20 → 2025-11-30
Output tokens
Deprecated — pricing unavailable
- $2.19 / 1M tokens (output) Historical 2025-01-20 → 2025-11-30
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how DeepSeek R1 fits into a cost-aware routing setup
See how →Capability profile
Coding benchmarks
View full leaderboard →| Benchmark | Score | Harness | Source |
|---|---|---|---|
| swe-bench-verified ↗ | 49.2% 2025-01 | agentless | 📄 paper ↗ |
| swe-bench-verified ↗ | 57.6% 2025-05 | agentless | 🏢 vendor ↗ |
| livecodebench ↗ | 65.9% 2025-01 | — | 📄 paper ↗ |
| aider-polyglot ↗ | 56.9% 2025-01 | aider | 📊 leaderboard ↗ |
| aider-polyglot ↗ | 71.4% 2025-05 | aider | 📊 leaderboard ↗ |
| terminal-bench-2-1 | 19.1% 2026-08 | terminus-2 | 🔬 independent ↗ |
| bigcodebench ↗ | 35.1% 2025-01 | — | 📊 leaderboard ↗ |
Operator guidance
Route here when reasoning depth is the primary requirement and o1/o3 cost is prohibitive. Open weights mean it can be self-hosted or accessed via third-party hosts (Together AI, OpenRouter) at significantly lower cost than native DeepSeek API. Avoid for latency-sensitive or simple workloads — the reasoning overhead adds substantial time and token cost.
Use cases
- Complex multi-step reasoning, math, and science problems
- Advanced code generation and algorithm design
- Open-weights reasoning deployment on own infrastructure
Limitations
- The deepseek-reasoner API alias stopped naming this model back on 2025-08-21 (DeepSeek-V3.1's hybrid thinking/non-thinking unification); it now routes to deepseek-v4-flash's thinking mode and retires entirely 2026-07-24 — use open-weights hosts (R1 / R1-0528) for longevity
- Reasoning tokens can 3–5× the effective cost on complex problems
- Chinese lab; data-residency or compliance constraints may apply
- Capability ratings are qualitative, not a single cited benchmark run