Olmo 3 32B Think
Language Available Full comparison ↗ allenai/olmo-3-32b-think · by Allen Institute for AI (Ai2)
· decoder-only-transformer
Pricing — 1 offering(s)
Input tokens
- $0.15 / 1M tokens (input) Current 2025-11-20 → present
Output tokens
- $0.50 / 1M tokens (output) Current 2025-11-20 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Olmo 3 32B Think fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Choose Olmo 3 32B Think when full openness (data + recipe + checkpoints) or self-hosting is the requirement and mid-tier reasoning accuracy is acceptable. For maximum reasoning quality at similar size, Qwen 3 32B and proprietary reasoning models score higher. If you want Ai2's openness with the strongest reasoning, look at the newer Olmo 3.1 32B Think instead of this snapshot. Not the choice for long-context, multilingual, or latency-critical work.
Use cases
- Research and reproducibility work that needs a reasoning model whose full training pipeline and data are inspectable
- Self-hosted reasoning under Apache 2.0 where data cannot leave the environment
- Math / logic / step-by-step analytical tasks at open-model cost where frontier accuracy is not required
- Building on top of a documented model flow — continued pretraining or custom post-training from published checkpoints
Limitations
- Trails top open-weight and proprietary reasoning models — 'best fully open', not 'best'
- Short served context (~66K on this endpoint) and weak multilingual coverage
- Verbose chain-of-thought output directly increases output-token cost
- Superseded for reasoning quality by Ai2's own later Olmo 3.1 32B Think
- No Ai2 first-party per-token API; hosted here via OpenRouter (AllenAI upstream) — see registry note
- Capability ratings are qualitative from Ai2's blog and comparisons; benchmark figures are Ai2-reported