Hermes 4 70B
Language Available Full comparison ↗ nous-research/hermes-4-70b · by Nous Research
· decoder-only-transformer
· Built on meta/llama-3.1-70b
Pricing — 1 offering(s)
Input tokens
- $0.13 / 1M tokens (input) Current 2025-08-26 → present
Output tokens
- $0.40 / 1M tokens (output) Current 2025-08-26 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Hermes 4 70B fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Route here for open-weights reasoning with high steerability at low cost — the Hermes line's traditional niche. For frontier reasoning accuracy, a larger open model (Qwen 3 235B, DeepSeek V4) or a proprietary reasoning model will outperform it. The Meta Llama 3 Community License (below) is more restrictive than the Apache-2.0 of most other open models in this registry — check it if licensing is a constraint.
Use cases
- Open-weights reasoning at 70B scale where steerability and low refusal rates matter (agent backends, structured extraction, roleplay/persona work)
- Self-hosted deployment where a togglable reasoning mode is wanted without running a frontier-size model
- Tasks needing strong math/logic at open-model cost
Limitations
- Meta Llama 3 Community License — a derivative-model license with use restrictions, not a fully-open Apache/MIT license
- Benchmark figures are from Nous's own technical report; no independent third-party leaderboard reproduction found
- No first-party hosted API — served here via OpenRouter (single upstream provider)
- Reasoning-mode benchmarks are substantially higher than non-reasoning mode; a fair comparison depends on which mode the caller uses