Compare AI models on price and capability
Modelglass tracks and verifies pricing and capability ratings for 164 AI models — 34 image, 69 language, 30 video, and 31 audio — across every major provider. This page puts them side by side so you can weigh price against capability before committing to a model.
Each modality section below has a model selector: tick the models you want and the table updates to show their price and a per-dimension capability rating. Ratings use one five-level scale — strong, moderate, weak, variable, unknown — synthesised from published benchmarks, model cards, and independent evaluations so models are comparable across providers. Click any capability name to see how that dimension is judged; the glossary covers the terms in full.
Prices are shown in each provider's real billing unit — per image, per 1M tokens, per second of video, and so on — not forced onto a single number, so where two models bill differently the page says there's no direct comparison rather than inventing one. Every price carries a source and a verified date; the Newly added and Price change views below show what moved in the last 30 days.
Turn this comparison into a cost-aware routing setup
See how →Newly added (last 30 days)
Price changes (last 30 days)
Image models
28 models
28 image models tracked here. Per-image pricing ranges from $0.0037 to $0.13 (21 of 28 priced models bill this way; others use a different unit). 2 added in the last 30 days.
| Capability | Adobe Firefly Image 3
$0.02 / img | Animagine XL 3.1
$0.0037 / img | DALL·E 3
$0.04 / img | FLUX 1.1 [pro]
$0.04 / MP | FLUX.1 [dev]
$0.025 / MP | FLUX.1 [pro]
$0.055 / img | FLUX.1 [schnell]
$0.003 / MP | FLUX.1 Kontext
$0.04 / img | Gemini 2.5 Flash Image
$0.039 / img | Gemini 3 Pro Image
$0.134 / img | Gemini 3.1 Flash Image
$0.045 / img | GPT Image 1
$0.011 / img | GPT Image 2
$0.006 / img | Hunyuan-DiT v1.1 (Distilled)
$0.000975 / s | HunyuanImage 3.0
$0.1 / MP | Ideogram 2.0
$0.05 / img | Ideogram 3.0
$0.06 / img | Imagen 4 | Krea-2-Turbo
$0.008 / MP | Leonardo Phoenix
$0.00257 per_credit | Luma Uni-1.1
$0.0404 / img | Midjourney v7
$10 per_month | Recraft V3
$0.04 / img | Runway Gen-4 Image
$0.01 per_credit | Seedream 4.0
$0.03 / img | Seedream 4.5
$0.04 / img | Seedream 5.0 Pro
$0.045 / img | Stable Diffusion 3.5 Large
$0.065 / img | Stable Diffusion 3.5 Large Turbo
$0.04 / img | Stable Diffusion 3.5 Medium
$0.035 / img | Stable Diffusion XL 1.0
$0.0014 / s | Stable Image Core
$0.03 / img | Stable Image Ultra
$0.08 / img | Z-Image-Turbo
$0.005 / img |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Strong | Moderate est. | Strong | Strong | Strong | Strong | Moderate | Strong est. | Unknown est. | Strong est. | Strong est. | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Unknown est. | Strong | Strong est. | Moderate | Strong | Moderate | Unknown | Strong est. | Strong est. | Strong | Moderate | Moderate | Moderate | — | — | Moderate | |
| How faithfully the image reflects everything the prompt asked for — objects, attributes, relationships, and intent. The single most important axis for most production use, and what alignment benchmarks (GenEval, DPG-Bench, CLIPScore) try to measure. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Weak est. | Moderate | Strong | Strong | Strong | Moderate | Unknown est. | Unknown est. | Unknown est. | Unknown est. | Strong | Strong | Moderate | Strong | Moderate | Strong | Strong | Unknown est. | Strong | Unknown est. | Strong | Moderate | Strong | Unknown | Unknown est. | Strong est. | Strong | Strong | Moderate | Strong | — | — | Strong | |
| How convincingly the model renders real-world scenes, lighting, skin, and materials. Distinct from "looks nice" — a model can be highly aesthetic but stylised rather than photoreal. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Strong est. | Strong | Strong | Strong | Strong | Moderate | Unknown est. | Unknown est. | Unknown est. | Unknown est. | Strong | Strong | — | Strong | Strong | Strong | Strong | Moderate est. | Strong | Unknown est. | Strong | Strong | Strong | Unknown | Unknown est. | Unknown est. | Strong | Strong | — | Strong | — | — | Moderate | |
| The breadth of styles a model can produce (illustration, painting, 3D, anime, graphic design) and how well it follows style instructions. A wide ecosystem of fine-tunes/LoRAs effectively extends this axis. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Weak est. | Strong | Strong | Moderate | Strong | Moderate | Strong est. | Unknown est. | Strong est. | Strong est. | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Unknown est. | Moderate | Unknown est. | Moderate | Strong | Moderate | Strong | Strong est. | Strong est. | Strong | Moderate | Moderate | Weak | — | — | Strong | |
| The ability to render legible, correctly-spelled text inside the image (signs, logos, labels). A long-standing weak spot for diffusion models; newer transformer-backbone models are markedly better. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Moderate est. | Strong | Strong | Strong | Strong | Moderate | Strong est. | Unknown est. | Moderate est. | Unknown est. | Strong | Strong | — | Strong | Moderate | Moderate | Strong | Unknown est. | Moderate | Strong est. | Moderate | Moderate | Moderate | Strong | Strong est. | Strong est. | Strong | Moderate | Moderate | Moderate | — | — | Moderate | |
| Getting multi-object scenes right: correct counts, spatial relationships ("A on top of B"), and binding the right attribute to the right object (the "red cube, blue sphere" problem). Measured by GenEval and T2I-CompBench. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Moderate est. | Moderate | Strong | Moderate | Strong | Moderate | Unknown est. | Weak est. | Strong est. | Strong est. | Moderate | Strong | — | Moderate | Moderate | — | Strong | Unknown est. | Moderate | Moderate est. | Strong | Strong | — | Strong | Unknown est. | Unknown est. | Moderate | Moderate | Moderate | Moderate | — | — | Moderate | |
| The largest / highest-quality native output the model produces before needing upscaling, and the set of aspect ratios it supports well. Matters for hero and print assets. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Strong est. | Weak | Moderate | Moderate | Moderate | Strong | Strong est. | Strong est. | Unknown est. | Strong est. | Weak | Moderate | Strong | Weak | Moderate | — | Moderate | Strong est. | Moderate | Moderate est. | Moderate | Moderate | — | Strong | Unknown est. | Unknown est. | Moderate | Strong | Strong | Moderate | — | — | Strong | |
| How quickly the model produces an image, driven mainly by step count and backbone size. Directly tied to per-image cost on compute-billed hosts and to user-facing latency. What each rating means here
| ||||||||||||||||||||||||||||||||||
Select models to compare
Language models
65 models
65 language models tracked here. Per-1M tokens (input) pricing ranges from $0.017 to $15.00 (64 of 65 priced models bill this way; others use a different unit). 3 repriced in the last 30 days.
| Capability | Claude 3.5 Haiku
$0.8 / 1M in | Claude 3.5 Sonnet
$3 / 1M in | Claude Fable 5
$10 / 1M in | Claude Haiku 4
$1 / 1M in | Claude Opus 4
$5 / 1M in | Claude Opus 4.8
$5 / 1M in | Claude Opus 5
$5 / 1M in | Claude Sonnet 4
$3 / 1M in | Claude Sonnet 4.6
$3 / 1M in | Claude Sonnet 5
$2 / 1M in | Command A
$2.5 / 1M in | Command R+
$3 / 1M in | DeepSeek R1 | DeepSeek V3 | DeepSeek V3.2
$0.28 / 1M in | DeepSeek V4-Flash
$0.14 / 1M in | DeepSeek V4-Pro
$0.435 / 1M in | ERNIE 5.1
$0.59 / 1M in | Gemini 2.0 Flash
$0.1 / 1M in | Gemini 2.5 Flash
$0.3 / 1M in | Gemini 2.5 Pro
$1.25 / 1M in | Gemini 3.1 Pro
$2 / 1M in | Gemini 3.5 Flash
$1.5 / 1M in | GLM-5.1
$1.4 / 1M in | GLM-5.2
$1.4 / 1M in | GPT-4o
$2.5 / 1M in | GPT-4o mini
$0.15 / 1M in | GPT-5.2
$1.75 / 1M in | GPT-5.2-Codex
$1.75 / 1M in | GPT-5.3-Codex
$1.75 / 1M in | GPT-5.4 mini
$0.75 / 1M in | GPT-5.5
$5 / 1M in | GPT-5.5 Pro
$15 / 1M in | GPT-5.6 Luna
$1 / 1M in | GPT-5.6 Sol
$5 / 1M in | GPT-5.6 Terra
$2.5 / 1M in | Granite 4.0 H Micro
$0.017 / 1M in | Grok 3 | Hermes 4 70B
$0.13 / 1M in | Inkling
$1.87 / 1M in | Jamba Large 1.7
$2 / 1M in | Jamba Mini 2
$0.2 / 1M in | K-EXAONE-236B-A23B | Kimi K2.5
$0.6 / 1M in | Kimi K2.6
$0.95 / 1M in | Kimi K2.7 Code
$0.95 / 1M in | Kimi K3
$3 / 1M in | Leanstral 1.5
$0 / 1M tok | Ling-3.0-flash
$0.020999999999999998 / 1M in | Llama 3.3 70B
$0.59 / 1M in | Llama 4 Maverick
$0.27 / 1M in | Llama 4 Scout
$0.18 / 1M in | MiMo-V2.5-Pro
$0.4109589041095891 / 1M in | MiniMax M3
$0.3 / 1M in | Mistral Large 3
$0.5 / 1M in | Mistral Small 4
$0.1 / 1M in | Nova 2 Lite
$0.3 / 1M in | o3
$2 / 1M in | o4-mini
$1.1 / 1M in | Olmo 3 32B Think
$0.15 / 1M in | Qwen 2.5 72B
$1.2 / 1M in | Qwen 3 235B-A22B
$0.455 / 1M in | Reka Flash
$0.8 / 1M in | Solar Pro 3
$0.15 / 1M in | Sonar
$1 / 1M in
Search-grounded
| Sonar Deep Research
$2 / 1M in
Search-grounded
| Sonar Pro
$3 / 1M in
Search-grounded
| Sonar Reasoning Pro
$2 / 1M in
Search-grounded
| Step 3.5 Flash
$0.1 / 1M in |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Context window | 200K tokens | 200K tokens | 1M tokens | 200K tokens | 1M tokens | 1M tokens | 1M tokens | 200K tokens | 1M tokens | 1M tokens | 256K tokens | 128K tokens | 128K tokens | 164K tokens | 128K tokens | 1M tokens | 1M tokens | 128K tokens | 1M tokens | 1M tokens | 1M tokens | 1.048576M tokens | 1.048576M tokens | 200K tokens | 1M tokens | 128K tokens | 128K tokens | 400K tokens | 400K tokens | 400K tokens | 400K tokens | 1.05M tokens | 272K tokens | 1.05M tokens | 1.05M tokens | 1.05M tokens | 131K tokens | 131K tokens | 131K tokens | 1M tokens | 256K tokens | 256K tokens | — tokens | 262K tokens | 262K tokens | 262K tokens | 1M tokens | 256K tokens | 262K tokens | 128K tokens | 1M tokens | 10M tokens | 1M tokens | 1M tokens | 128K tokens | 128K tokens | 1M tokens | 200K tokens | 200K tokens | 66K tokens | 131K tokens | 131K tokens | 128K tokens | 128K tokens | 128K tokens | 128K tokens | 200K tokens | 128K tokens | 262K tokens |
| Moderate | Strong | Strong | Moderate est. | Strong | Strong | Strong est. | Strong | Strong | Strong | Moderate | Moderate est. | Strong | Strong | Strong est. | Moderate est. | Strong | Strong est. | Moderate | Moderate | Strong | Strong | Strong | Strong est. | Moderate est. | Strong | Moderate | Strong | Strong | Strong | Moderate | Strong | — | Weak est. | Strong | Moderate est. | Weak | Strong | Strong | Strong est. | Moderate | Weak | — | Moderate est. | Strong est. | Strong est. | Strong est. | Strong | Moderate est. | Moderate | Moderate | Moderate | Moderate est. | Strong | Strong | Moderate | Moderate | Strong | Strong | Moderate | Moderate | Strong | Unknown est. | Moderate est. | Weak est. | Strong est. | Moderate est. | Strong est. | Strong | |
| Measures the model's ability to solve multi-step logical problems, draw correct inferences, and handle abstract or mathematical reasoning tasks. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Moderate | Moderate | Strong | Moderate est. | Moderate | Strong | Strong est. | Strong | Strong | Strong | Moderate | Unknown est. | Moderate | Moderate | Moderate est. | Moderate est. | Strong | Unknown est. | Moderate | Moderate | Strong | Strong | Strong | Strong est. | Strong est. | Moderate | Moderate | Strong | Strong | Strong | Moderate | Strong | — | Moderate est. | Strong | Moderate est. | Moderate | Moderate | Moderate | Strong est. | Moderate | Weak | — | Strong est. | Strong est. | Strong est. | Strong est. | Moderate | Moderate est. | Moderate | Moderate | Moderate | Strong est. | Strong | Strong | Moderate | Moderate | Strong | Strong | Moderate | Strong | Strong | Unknown est. | Moderate est. | Weak est. | Weak est. | Weak est. | Moderate est. | Strong | |
| Measures ability to write correct, idiomatic code across common programming languages — from simple utility functions to complex algorithmic problems and debugging. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Strong | Strong | Strong | Strong est. | Strong | Strong | Strong est. | Strong | Strong | Strong | Strong | Strong est. | Weak | Moderate | Strong est. | Strong est. | Strong | Unknown est. | Strong | Strong | Strong | Strong | Strong | Strong est. | Strong est. | Strong | Strong | Strong | Strong | Strong | Strong | Strong | — | Strong est. | Strong | Strong est. | Strong | Strong | Moderate | Moderate est. | Moderate | Moderate | — | Unknown est. | Unknown est. | Unknown est. | Strong est. | Strong | Moderate est. | Moderate | Moderate | Moderate | Strong est. | Strong | Strong | Strong | Moderate | Strong | Strong | Moderate | Moderate | Moderate | Moderate est. | Moderate est. | Moderate est. | Strong est. | Moderate est. | Moderate est. | Strong | |
| Measures how reliably the model selects and calls external tools, APIs, and functions — including parameter formatting, multi-turn loops, and chaining dependent calls. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Strong | Strong | Strong | Strong est. | Strong | Strong | Strong est. | Strong | Strong | Strong | Strong | Strong est. | Moderate | Strong | Strong est. | Moderate est. | Moderate | Unknown est. | Strong | Strong | Strong | Strong | Strong | Unknown est. | Unknown est. | Strong | Strong | Strong | Strong | Strong | Strong | Strong | — | — | — | — | Strong | Strong | Strong | Strong est. | Strong | Strong | — | Unknown est. | Unknown est. | Unknown est. | Unknown est. | — | Moderate est. | Strong | Strong | Strong | Moderate est. | — | Strong | Strong | Moderate | Strong | Strong | Moderate | Strong | Strong | Unknown est. | Moderate est. | Moderate est. | Strong est. | Moderate est. | Strong est. | Moderate | |
| Measures how faithfully the model respects explicit constraints — output format, length limits, persona, negative instructions (what NOT to do), and multi-rule system prompts. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Strong | Strong | Strong | Strong est. | Strong | Strong | Strong est. | Strong | Strong | Strong | Strong | Moderate est. | Moderate | Strong | Moderate est. | Strong est. | Strong | Strong est. | Strong | Strong | Strong | Strong | Strong | Strong est. | Strong est. | Moderate | Moderate | Strong | Strong | Strong | Strong | Strong | — | Strong est. | Strong | Strong est. | Moderate | Moderate | Moderate | Unknown est. | Strong | Strong | — | Strong est. | Strong est. | Strong est. | Strong est. | Strong | Moderate est. | Strong | Strong | Strong | Strong est. | Strong | Strong | Strong | Strong | Strong | Strong | Weak | Moderate | Moderate | Weak est. | Moderate est. | Moderate est. | Moderate est. | Strong est. | Moderate est. | Strong | |
| Measures practical usable context length — how accurately the model recalls and synthesises information spread across a long input, not just the advertised token ceiling. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Moderate | Moderate | Strong | Moderate est. | Strong | Strong | Strong est. | Strong | Strong | Strong | Moderate | Moderate est. | Moderate | Moderate | Moderate est. | Moderate est. | Moderate | Unknown est. | Strong | Strong | Strong | Strong | Strong | Unknown est. | Unknown est. | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | — | — | — | — | Moderate | Moderate | Moderate | Strong est. | Moderate | Moderate | — | Unknown est. | Unknown est. | Unknown est. | Unknown est. | — | Moderate est. | Moderate | Moderate | Moderate | Moderate est. | — | Strong | Strong | Moderate | Moderate | Moderate | Weak | Strong | Strong | Moderate est. | Strong est. | Moderate est. | Moderate est. | Moderate est. | Moderate est. | Moderate | |
| Measures output quality across non-English languages — covering generation fluency, translation accuracy, and how well quality holds for lower-resource languages. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Strong | Moderate | Weak | Strong est. | Weak | Weak | Moderate est. | Moderate | Moderate | Moderate | Moderate | Moderate est. | Weak | Moderate | Moderate est. | Strong est. | Moderate | Unknown est. | Strong | Strong | Moderate | Moderate | Strong | Unknown est. | Unknown est. | Moderate | Strong | Moderate | Moderate | Moderate | Strong | Moderate | — | Strong est. | Weak | Moderate est. | Strong | Moderate | Moderate | Unknown est. | Strong | Strong | — | Weak est. | Moderate est. | Weak est. | Weak est. | — | Strong est. | Strong | Strong | Strong | Strong est. | Strong | Moderate | Strong | Strong | Weak | Moderate | Moderate | Moderate | Moderate | Strong est. | Strong est. | Strong est. | Weak est. | Moderate est. | Weak est. | Strong | |
| Measures output generation speed (tokens per second), which determines first-token latency for interactive use and per-token cost efficiency at scale. What each rating means here
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. "—" means no profile data for that dimension.
Select models to compare
Video models
26 models
26 video models tracked here. Per-second pricing ranges from $0.035 to $0.16 (14 of 23 priced models bill this way; others use a different unit).
| Capability | Act Two
$0.05 / s | Aleph 2
$0.01 per_credit | CogVideoX-5B
$0.2 / clip | FLUX 3 Video
$0.06 / s | Gemini Omni Flash
$0.1 / s | Gen-3 Alpha | Gen-4 Turbo
$0.01 per_credit | Gen-4.5
$0.01 per_credit | Grok Imagine Video 1.5
$0.08 / s | Hailuo-02
$0.076 / clip | HappyHorse 1.0
$0.15 / s | HunyuanVideo 1.5
$0.4 / clip | Kling 1.6
$0.04 / s | Kling 2.1
$0.06 / s | LTX Video 0.9.7
$0.048 / clip | LTX-2.3
$0.04 / s | Mochi 1
$0.42 / clip | Pika 2.2
$0.2 / clip | Ray 3.2
$0.3 / clip | Seedance 2
$0.16 / s | Seedance 2.0 | Seedance 2.0 Fast | Seedance 2.0 Mini | Sora 2
$0.1 / s | Veo 2
$0.35 / s | Veo 3
$0.1 / s | Veo 3.1
$0.05 / s | Vidu Q3
$0.035 / s | Wan 2.1
$0.09 / s | Wan 2.5
$0.05 / s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Strong | Moderate | Moderate | Unknown | Moderate est. | Moderate | Moderate est. | Strong est. | Unknown est. | Moderate est. | Strong | Strong est. | Moderate est. | Strong est. | Moderate | Unknown est. | Moderate | Moderate est. | Strong est. | Strong est. | Strong | Unknown est. | Unknown est. | Strong est. | Strong | Strong est. | Strong est. | Unknown est. | Moderate | Moderate est. | |
| Measures how realistic and coherent motion is across frames — including natural movement, fluid transitions, and absence of warping, jitter, or ghosting artifacts. What each rating means here
| ||||||||||||||||||||||||||||||
| Strong | Strong | Moderate | Unknown | Moderate est. | Moderate | Strong est. | Strong est. | Strong est. | Moderate est. | Strong | Strong est. | Moderate est. | Strong est. | Moderate | Unknown est. | Moderate | Moderate est. | Strong est. | Strong est. | Unknown | Unknown est. | Unknown est. | Strong est. | Strong | Strong est. | Strong est. | Unknown est. | Strong | Strong est. | |
| Measures whether subjects, backgrounds, and fine details remain stable frame-to-frame — avoiding identity drift, morphing, or unexpected visual changes mid-clip. What each rating means here
| ||||||||||||||||||||||||||||||
| Weak | Moderate | Moderate | Strong | Strong est. | Moderate | Moderate est. | Strong est. | Moderate est. | Moderate est. | Strong | Strong est. | Moderate est. | Strong est. | Moderate | Moderate est. | Moderate | Moderate est. | Moderate est. | Strong est. | Strong | Unknown est. | Unknown est. | Strong est. | Strong | Strong est. | Strong est. | Moderate est. | Moderate | Moderate est. | |
| Measures how closely the generated clip matches the text description — including subject, action, style, composition, and camera direction instructions. What each rating means here
| ||||||||||||||||||||||||||||||
| Strong | Weak | Weak | Strong | Weak est. | Weak | Weak est. | Weak est. | Strong est. | Weak est. | Strong | Weak est. | Weak est. | Moderate est. | Weak | Strong est. | Weak | Weak est. | Weak est. | Moderate est. | Strong | Unknown est. | Unknown est. | Weak est. | Weak | Strong est. | Strong est. | Strong est. | Weak | Weak est. | |
| Measures whether the model generates synchronised audio — ambient sound, foley, speech, or music — natively alongside the video rather than requiring a separate pipeline. What each rating means here
| ||||||||||||||||||||||||||||||
| Moderate | — | Weak | Moderate | Weak est. | Moderate | — | — | Unknown est. | — | Moderate | — | Weak est. | Moderate est. | Weak | Strong est. | Weak | — | Moderate est. | Strong est. | Unknown | Unknown est. | Unknown est. | Moderate est. | Moderate | Strong est. | Strong est. | Strong est. | — | — | |
| Measures the highest output resolution the model produces natively before quality degrades, independent of any post-processing upscaling. What each rating means here
| ||||||||||||||||||||||||||||||
| Strong | Moderate | Weak | Unknown | Strong est. | Moderate | Strong est. | Weak est. | Unknown est. | Moderate est. | Unknown | Weak est. | Moderate est. | Moderate est. | Strong | Variable est. | Weak | Moderate est. | Moderate est. | Moderate est. | — | — | — | Moderate est. | Moderate | Moderate est. | Moderate est. | Variable est. | Strong | Strong est. | |
| Measures generation latency relative to clip length. Slow inference raises cost per second of output and limits interactive or real-time applications. What each rating means here
| ||||||||||||||||||||||||||||||
| Moderate | Strong | Moderate | Unknown | Weak est. | Moderate | Moderate est. | Moderate est. | Unknown est. | Moderate est. | Strong | Moderate est. | Moderate est. | Moderate est. | Moderate | Moderate est. | Weak | Weak est. | Moderate est. | Strong est. | Strong | Unknown est. | Unknown est. | Moderate est. | Moderate | Moderate est. | Moderate est. | Strong est. | Moderate | Moderate est. | |
| Measures the maximum clip length achievable in a single generation pass while maintaining consistent quality — longer ceilings reduce stitching requirements. What each rating means here
| ||||||||||||||||||||||||||||||
Ratings synthesised from provider benchmarks and independent evaluations. "—" means no profile data for that dimension.
Select models to compare
Audio models
29 models
29 audio models tracked here. Per-1M characters pricing ranges from $4.00 to $100.00 (13 of 29 priced models bill this way; others use a different unit). 1 repriced in the last 30 days.
Text-to-speech
Skip to model selector ↓| Capability | Amazon Polly $4 / 1M chars | Azure Neural TTS $4 / 1M chars | Cartesia Sonic $65 / 1M chars | ElevenLabs Eleven v3 $100 / 1M chars | ElevenLabs Flash v2.5 $50 / 1M chars | ElevenLabs Multilingual v2 $100 / 1M chars | Fish Audio S1 $15 / 1M chars | Fish Audio S2 Pro $15 / 1M chars | Fish Audio S2.1 Pro $15 / 1M chars | Google Cloud TTS $4 / 1M chars | GPT-4o Audio Preview ~$0.096 / minute (est.) | GPT-Audio 1.5 ~$0.0768 / minute (est.) | Grok Voice Think Fast 2.0 $0.08 / minute | Inworld TTS-1.5 Max $35 / 1M chars | OpenAI TTS-1 $15 / 1M chars | OpenAI TTS-1 HD $30 / 1M chars | PlayHT 2.0 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Moderate | Strong | Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | Strong | |
| How natural and human-like the synthesized speech sounds — covering prosody, pacing, intonation, and absence of robotic or mechanical artefacts. What each rating means here
| |||||||||||||||||
| Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Moderate | Unknown | Unknown | Weak | Weak | Strong | |
| The range of distinct voices, accents, ages, and speaking styles available out of the box from the provider's voice library. What each rating means here
| |||||||||||||||||
| Weak | Moderate | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Weak | Weak | Weak | Unknown | Strong | Weak | Weak | Strong | |
| Ability to clone a custom voice from a short audio sample, enabling personalised or brand-consistent TTS output. What each rating means here
| |||||||||||||||||
| Strong | Moderate | Strong | Moderate | Strong | Moderate | Unknown | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Moderate | Moderate | |
| Time-to-first-audio-chunk when using the streaming endpoint — lower latency enables real-time conversational applications. What each rating means here
| |||||||||||||||||
| Strong | Strong | Moderate | Strong | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Strong | Strong | Strong | |
| Number and quality of supported output languages. Strong coverage means high-quality synthesis across many major and minor languages. What each rating means here
| |||||||||||||||||
Select TTS models to compare 16 models
Frequently asked questions
- What can I compare on Modelglass?
- Pricing, architecture, and expert-synthesised capability ratings for 164 AI models — 34 image, 69 language, 30 video and 31 audio — side by side. Pick models with the selector in each section. To compare two specific models with a shareable link, open either model's page and use "Full comparison", or go to /compare/<model>.
- Where do the capability ratings come from?
- They are synthesised from published benchmarks, provider model cards, and independent evaluations, then expressed on one five-level scale — strong, moderate, weak, variable, unknown — so models are comparable across providers. A small "low confidence" mark next to a rating means limited benchmark data was available. Every rating links through to that model's full profile and its citations.
- Why can't I compare some prices directly?
- Providers bill on different units — per image, per megapixel, per 1M tokens, per second of video, per 1,000 characters, per clip. Modelglass shows each model's real billing unit rather than forcing everything onto one number. Where two models bill on different units there is no honest per-unit comparison, and the page says so instead of inventing one.
- How current are the prices?
- Every price carries a source URL and the date it was verified, and a repricing is recorded as a new dated entry — never a silent overwrite. The "Newly added" and "Price change" views at the top of the page surface what moved in the last 30 days. Full price history is available through the paid API.
- Is it free to use?
- Yes. Browsing and comparing on the site needs no account and no API key. The paid product is the read API and MCP feed, for wiring this pricing and capability data into your own routing or tooling.
- What do "Fast", "Standard" and "Premium" mean?
- A rough quality-and-speed tier for image models: Fast is distilled or turbo variants, Standard is the balanced default, Premium is the highest-quality option. It is a filter to narrow the table, not a Modelglass ranking.
- How is this different from comparing two specific models?
- This page is the catalogue view — scan a whole modality at once. For a focused head-to-head with a shareable URL, such as Claude Sonnet 5 vs GPT-5.6, use a model's own /compare/<model> page: it adds a direct-answer summary of the price and capability differences between exactly the models you pick.
- How should I choose between two models that rate similarly?
- Start with price on a matched unit — often the clearest separator. After that the tiebreakers on this page are context window and speed (shown as their own capability rows), how recently the price last moved (a stable price is easier to build against), and lifecycle status: a model marked deprecated or platform-only is filtered out by default for a reason. Each model's own page carries the provenance and limitations behind its ratings.