Inkling
Language Available Full comparison ↗ thinking-machines/inkling · by Thinking Machines Lab
· mixture-of-experts
Pricing — 1 offering(s)
Input tokens (64K context)
- $1.87 / 1M tokens (input) Current 2026-07-15 → present
Output tokens (64K context)
- $4.68 / 1M tokens (output) Current 2026-07-15 → present
Input tokens (256K context, thinkingmachines/Inkling:peft:262144)
- $3.74 / 1M tokens (input) Current 2026-07-15 → present
Output tokens (256K context, thinkingmachines/Inkling:peft:262144)
- $9.36 / 1M tokens (output) Current 2026-07-15 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Inkling fits into a cost-aware routing setup
See how →Capability profile
Operator guidance
Now benchmark-backed on five of eight dimensions (see capability_profile) — strong on reasoning, coding, instruction-following, and multilingual, with a mixed tool-use signal. All figures are vendor-published (Thinking Machines' own model card), not yet independently corroborated by a third-party leaderboard, and speed/cost-efficiency are still unrated. The model card positions it for general-purpose multimodal-input (text/image/audio) conversational and coding-assistant use, at a currently promotion-discounted price on Tinker, the only platform serving it today. Re-evaluate once independent corroboration or non-promotional pricing become available.
Use cases
- General-purpose conversational use, instruction-following, and other natural language and multimodal tasks
- Coding assistants
- Chatbots
- RAG (retrieval-augmented generation) systems
Limitations
- May exhibit hallucination and imprecise instruction-following (thinkingmachines.ai/model-card/inkling/)
- Degraded performance in long multi-turn conversations, per the model card
- Not recommended for medical, legal, or safety-critical decision-making without additional fine-tuning, per the model card
- Capability ratings (reasoning/coding/tool-use/instruction-following/multilingual) are sourced from Thinking Machines' own Hugging Face model card, not an independent third-party leaderboard — no independent corroboration found as of the 2026-08-19 SCO-462 sweep. Speed and cost-efficiency remain wholly unrated — no figure found anywhere for either.
- Only served today at 64K/256K context via Tinker, not the full 1,000,000-token architectural context window this entry's context_window field states
- Current pricing reflects a stated 'limited-time 50% discount' on Tinker — treat the registry's current price as temporary, not a stable baseline