Granite 4.0 H Micro
Language Available Full comparison ↗ ibm/granite-4-0-h-micro · by IBM
· decoder-only-transformer
Pricing — 1 offering(s)
Input tokens
- $0.017 / 1M tokens (input) Current 2025-10-02 → present
Output tokens
- $0.11 / 1M tokens (output) Current 2025-10-02 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Granite 4.0 H Micro fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Route here for cost-, latency-, or memory-bound tasks that are mostly instruction-following, tool-calling, or retrieval-grounded rather than open-ended reasoning. For anything needing real analytical depth, step up to a larger Granite 4.0 tier (H Small) or a general mid-size model. Its differentiators versus other cheap small models are the Mamba-2 hybrid efficiency, the strong function-calling story for its size, and IBM's enterprise provenance guarantees.
Use cases
- Edge / on-device / local-first deployments where memory and latency are hard constraints
- High-volume function-calling and RAG pipelines where a small, cheap, instruction-tuned model is sufficient
- Enterprise settings that need Apache-2.0 licensing, signed checkpoints, and ISO 42001 provenance
- The cheap tier in a cascade, handling routing / extraction / tool-dispatch before a larger model
Limitations
- 3B parameters — weak on hard reasoning; not for analysis-heavy work
- No IBM first-party public per-token API (watsonx.ai is enterprise-platform pricing); this entry is hosted via OpenRouter (see registry note)
- Served context (~131K) is lower than the 512K training / 128K eval figures IBM cites
- Capability ratings are qualitative from IBM's announcements; the benchmark claims are IBM-reported