Compare AI models on price and capability

Modelglass tracks and verifies pricing and capability ratings for 164 AI models — 34 image, 69 language, 30 video, and 31 audio — across every major provider. This page puts them side by side so you can weigh price against capability before committing to a model.

Each modality section below has a model selector: tick the models you want and the table updates to show their price and a per-dimension capability rating. Ratings use one five-level scale — strong, moderate, weak, variable, unknown — synthesised from published benchmarks, model cards, and independent evaluations so models are comparable across providers. Click any capability name to see how that dimension is judged; the glossary covers the terms in full.

Prices are shown in each provider's real billing unit — per image, per 1M tokens, per second of video, and so on — not forced onto a single number, so where two models bill differently the page says there's no direct comparison rather than inventing one. Every price carries a source and a verified date; the Newly added and Price change views below show what moved in the last 30 days.

Turn this comparison into a cost-aware routing setup

See how →

Image models

28 models

28 image models tracked here. Per-image pricing ranges from $0.0037 to $0.13 (21 of 28 priced models bill this way; others use a different unit). 2 added in the last 30 days.

Fast = distilled/turbo · Standard = balanced · Premium = highest quality
Strong Moderate Weak Variable Unknown
Skip to model selector ↓
Capability Adobe Firefly Image 3
$0.02 / img
Animagine XL 3.1
$0.0037 / img
DALL·E 3
$0.04 / img
FLUX 1.1 [pro]
$0.04 / MP
FLUX.1 [dev]
$0.025 / MP
FLUX.1 [pro]
$0.055 / img
FLUX.1 [schnell]
$0.003 / MP
FLUX.1 Kontext
$0.04 / img
Gemini 2.5 Flash Image
$0.039 / img
Gemini 3 Pro Image
$0.134 / img
Gemini 3.1 Flash Image
$0.045 / img
GPT Image 1
$0.011 / img
GPT Image 2
$0.006 / img
Hunyuan-DiT v1.1 (Distilled)
$0.000975 / s
HunyuanImage 3.0
$0.1 / MP
Ideogram 2.0
$0.05 / img
Ideogram 3.0
$0.06 / img
Imagen 4 Krea-2-Turbo
$0.008 / MP
Leonardo Phoenix
$0.00257 per_credit
Luma Uni-1.1
$0.0404 / img
Midjourney v7
$10 per_month
Recraft V3
$0.04 / img
Runway Gen-4 Image
$0.01 per_credit
Seedream 4.0
$0.03 / img
Seedream 4.5
$0.04 / img
Seedream 5.0 Pro
$0.045 / img
Stable Diffusion 3.5 Large
$0.065 / img
Stable Diffusion 3.5 Large Turbo
$0.04 / img
Stable Diffusion 3.5 Medium
$0.035 / img
Stable Diffusion XL 1.0
$0.0014 / s
Stable Image Core
$0.03 / img
Stable Image Ultra
$0.08 / img
Z-Image-Turbo
$0.005 / img
Strong Moderate est. Strong Strong Strong Strong Moderate Strong est. Unknown est. Strong est. Strong est. Strong Strong Moderate Strong Strong Strong Strong Unknown est. Strong Strong est. Moderate Strong Moderate Unknown Strong est. Strong est. Strong Moderate Moderate Moderate Moderate

How faithfully the image reflects everything the prompt asked for — objects, attributes, relationships, and intent. The single most important axis for most production use, and what alignment benchmarks (GenEval, DPG-Bench, CLIPScore) try to measure.

What each rating means here

Strong
Reliably produces outputs that match complex, specific prompts — including spatial relationships, counts, and fine detail.
Moderate
Handles straightforward prompts well, but drops or reinterprets details as requests get longer or more specific.
Weak
Captures the general idea but frequently ignores or invents elements; expect to fight the prompt.
Variable
Adherence swings with prompt phrasing, guidance settings, or seed — sometimes precise, sometimes loose.
Strong Weak est. Moderate Strong Strong Strong Moderate Unknown est. Unknown est. Unknown est. Unknown est. Strong Strong Moderate Strong Moderate Strong Strong Unknown est. Strong Unknown est. Strong Moderate Strong Unknown Unknown est. Strong est. Strong Strong Moderate Strong Strong

How convincingly the model renders real-world scenes, lighting, skin, and materials. Distinct from "looks nice" — a model can be highly aesthetic but stylised rather than photoreal.

What each rating means here

Strong
Renders convincing real-world scenes — believable lighting, skin, and materials that hold up to scrutiny.
Moderate
Looks realistic at a glance but shows tells on close inspection (hands, textures, reflections).
Weak
Outputs read as obviously synthetic or stylised rather than photographic.
Variable
Realism depends heavily on the subject and prompt — strong for some scenes, artificial for others.
Strong Strong est. Strong Strong Strong Strong Moderate Unknown est. Unknown est. Unknown est. Unknown est. Strong Strong Strong Strong Strong Strong Moderate est. Strong Unknown est. Strong Strong Strong Unknown Unknown est. Unknown est. Strong Strong Strong Moderate

The breadth of styles a model can produce (illustration, painting, 3D, anime, graphic design) and how well it follows style instructions. A wide ecosystem of fine-tunes/LoRAs effectively extends this axis.

What each rating means here

Strong
Fluently spans many styles (illustration, painting, 3D, anime, graphic design) and follows style direction closely.
Moderate
Covers common styles competently but has a clear default look it tends to fall back to.
Weak
Locked to a narrow aesthetic; style instructions have limited effect.
Variable
Range expands sharply with fine-tunes/LoRAs — base model is narrower than the wider ecosystem.
Moderate Weak est. Strong Strong Moderate Strong Moderate Strong est. Unknown est. Strong est. Strong est. Strong Strong Strong Strong Strong Strong Strong Unknown est. Moderate Unknown est. Moderate Strong Moderate Strong Strong est. Strong est. Strong Moderate Moderate Weak Strong

The ability to render legible, correctly-spelled text inside the image (signs, logos, labels). A long-standing weak spot for diffusion models; newer transformer-backbone models are markedly better.

What each rating means here

Strong
Renders legible, correctly-spelled text — short signs, logos, and labels come out clean.
Moderate
Manages short words but garbles longer strings, spacing, or unusual fonts.
Weak
Text is typically misspelled or illegible; treat it as decorative only.
Variable
Quality drops fast as the string gets longer — fine for a word, unreliable for a sentence.
Moderate Moderate est. Strong Strong Strong Strong Moderate Strong est. Unknown est. Moderate est. Unknown est. Strong Strong Strong Moderate Moderate Strong Unknown est. Moderate Strong est. Moderate Moderate Moderate Strong Strong est. Strong est. Strong Moderate Moderate Moderate Moderate

Getting multi-object scenes right: correct counts, spatial relationships ("A on top of B"), and binding the right attribute to the right object (the "red cube, blue sphere" problem). Measured by GenEval and T2I-CompBench.

What each rating means here

Strong
Gets multi-object scenes right: correct counts, spatial relationships, and the right attribute bound to the right object.
Moderate
Handles two or three elements but slips on counting, ordering, or attribute binding as scenes grow.
Weak
Struggles with anything beyond a single subject; objects, counts, and positions blur together.
Variable
Accuracy depends on scene complexity — dependable for simple layouts, shaky for crowded ones.
Strong Moderate est. Moderate Strong Moderate Strong Moderate Unknown est. Weak est. Strong est. Strong est. Moderate Strong Moderate Moderate Strong Unknown est. Moderate Moderate est. Strong Strong Strong Unknown est. Unknown est. Moderate Moderate Moderate Moderate Moderate

The largest / highest-quality native output the model produces before needing upscaling, and the set of aspect ratios it supports well. Matters for hero and print assets.

What each rating means here

Strong
Produces large, high-quality native output across many aspect ratios before any upscaling is needed.
Moderate
Solid at standard sizes; needs upscaling for hero or print-scale assets.
Weak
Limited to smaller native outputs; pushing the resolution degrades quality.
Variable
Usable resolution depends on the host/endpoint and chosen settings rather than the model alone.
Moderate Strong est. Weak Moderate Moderate Moderate Strong Strong est. Strong est. Unknown est. Strong est. Weak Moderate Strong Weak Moderate Moderate Strong est. Moderate Moderate est. Moderate Moderate Strong Unknown est. Unknown est. Moderate Strong Strong Moderate Strong

How quickly the model produces an image, driven mainly by step count and backbone size. Directly tied to per-image cost on compute-billed hosts and to user-facing latency.

What each rating means here

Strong
Generates images quickly — low step count or a distilled backbone keeps latency and per-image cost down.
Moderate
Typical generation time; fine for interactive use but not the fastest option.
Weak
Slow to generate, raising both latency and compute-billed cost per image.
Variable
Speed scales with step count and resolution — fast in a few-step mode, slow at maximum quality.
Select models to compare
|

Language models

65 models

65 language models tracked here. Per-1M tokens (input) pricing ranges from $0.017 to $15.00 (64 of 65 priced models bill this way; others use a different unit). 3 repriced in the last 30 days.

Strong Moderate Weak Variable Unknown
Skip to model selector ↓
Capability Claude 3.5 Haiku
$0.8 / 1M in
Claude 3.5 Sonnet
$3 / 1M in
Claude Fable 5
$10 / 1M in
Claude Haiku 4
$1 / 1M in
Claude Opus 4
$5 / 1M in
Claude Opus 4.8
$5 / 1M in
Claude Opus 5
$5 / 1M in
Claude Sonnet 4
$3 / 1M in
Claude Sonnet 4.6
$3 / 1M in
Claude Sonnet 5
$2 / 1M in
Command A
$2.5 / 1M in
Command R+
$3 / 1M in
DeepSeek R1 DeepSeek V3 DeepSeek V3.2
$0.28 / 1M in
DeepSeek V4-Flash
$0.14 / 1M in
DeepSeek V4-Pro
$0.435 / 1M in
ERNIE 5.1
$0.59 / 1M in
Gemini 2.0 Flash
$0.1 / 1M in
Gemini 2.5 Flash
$0.3 / 1M in
Gemini 2.5 Pro
$1.25 / 1M in
Gemini 3.1 Pro
$2 / 1M in
Gemini 3.5 Flash
$1.5 / 1M in
GLM-5.1
$1.4 / 1M in
GLM-5.2
$1.4 / 1M in
GPT-4o
$2.5 / 1M in
GPT-4o mini
$0.15 / 1M in
GPT-5.2
$1.75 / 1M in
GPT-5.2-Codex
$1.75 / 1M in
GPT-5.3-Codex
$1.75 / 1M in
GPT-5.4 mini
$0.75 / 1M in
GPT-5.5
$5 / 1M in
GPT-5.5 Pro
$15 / 1M in
GPT-5.6 Luna
$1 / 1M in
GPT-5.6 Sol
$5 / 1M in
GPT-5.6 Terra
$2.5 / 1M in
Granite 4.0 H Micro
$0.017 / 1M in
Grok 3 Hermes 4 70B
$0.13 / 1M in
Inkling
$1.87 / 1M in
Jamba Large 1.7
$2 / 1M in
Jamba Mini 2
$0.2 / 1M in
K-EXAONE-236B-A23B Kimi K2.5
$0.6 / 1M in
Kimi K2.6
$0.95 / 1M in
Kimi K2.7 Code
$0.95 / 1M in
Kimi K3
$3 / 1M in
Leanstral 1.5
$0 / 1M tok
Ling-3.0-flash
$0.020999999999999998 / 1M in
Llama 3.3 70B
$0.59 / 1M in
Llama 4 Maverick
$0.27 / 1M in
Llama 4 Scout
$0.18 / 1M in
MiMo-V2.5-Pro
$0.4109589041095891 / 1M in
MiniMax M3
$0.3 / 1M in
Mistral Large 3
$0.5 / 1M in
Mistral Small 4
$0.1 / 1M in
Nova 2 Lite
$0.3 / 1M in
o3
$2 / 1M in
o4-mini
$1.1 / 1M in
Olmo 3 32B Think
$0.15 / 1M in
Qwen 2.5 72B
$1.2 / 1M in
Qwen 3 235B-A22B
$0.455 / 1M in
Reka Flash
$0.8 / 1M in
Solar Pro 3
$0.15 / 1M in
Sonar
$1 / 1M in
Search-grounded
Sonar Deep Research
$2 / 1M in
Search-grounded
Sonar Pro
$3 / 1M in
Search-grounded
Sonar Reasoning Pro
$2 / 1M in
Search-grounded
Step 3.5 Flash
$0.1 / 1M in
Context window 200K tokens 200K tokens 1M tokens 200K tokens 1M tokens 1M tokens 1M tokens 200K tokens 1M tokens 1M tokens 256K tokens 128K tokens 128K tokens 164K tokens 128K tokens 1M tokens 1M tokens 128K tokens 1M tokens 1M tokens 1M tokens 1.048576M tokens 1.048576M tokens 200K tokens 1M tokens 128K tokens 128K tokens 400K tokens 400K tokens 400K tokens 400K tokens 1.05M tokens 272K tokens 1.05M tokens 1.05M tokens 1.05M tokens 131K tokens 131K tokens 131K tokens 1M tokens 256K tokens 256K tokens — tokens 262K tokens 262K tokens 262K tokens 1M tokens 256K tokens 262K tokens 128K tokens 1M tokens 10M tokens 1M tokens 1M tokens 128K tokens 128K tokens 1M tokens 200K tokens 200K tokens 66K tokens 131K tokens 131K tokens 128K tokens 128K tokens 128K tokens 128K tokens 200K tokens 128K tokens 262K tokens
Moderate Strong Strong Moderate est. Strong Strong Strong est. Strong Strong Strong Moderate Moderate est. Strong Strong Strong est. Moderate est. Strong Strong est. Moderate Moderate Strong Strong Strong Strong est. Moderate est. Strong Moderate Strong Strong Strong Moderate Strong Weak est. Strong Moderate est. Weak Strong Strong Strong est. Moderate Weak Moderate est. Strong est. Strong est. Strong est. Strong Moderate est. Moderate Moderate Moderate Moderate est. Strong Strong Moderate Moderate Strong Strong Moderate Moderate Strong Unknown est. Moderate est. Weak est. Strong est. Moderate est. Strong est. Strong

Measures the model's ability to solve multi-step logical problems, draw correct inferences, and handle abstract or mathematical reasoning tasks.

What each rating means here

Strong
Reliably solves complex multi-step problems — handles logic puzzles, mathematical proofs, and extended chain-of-thought reasoning with high accuracy.
Moderate
Handles straightforward reasoning tasks well but makes errors on longer inference chains or adversarial logic problems.
Weak
Frequently fails on problems requiring more than one or two reasoning steps; prone to confident-sounding but incorrect answers.
Moderate Moderate Strong Moderate est. Moderate Strong Strong est. Strong Strong Strong Moderate Unknown est. Moderate Moderate Moderate est. Moderate est. Strong Unknown est. Moderate Moderate Strong Strong Strong Strong est. Strong est. Moderate Moderate Strong Strong Strong Moderate Strong Moderate est. Strong Moderate est. Moderate Moderate Moderate Strong est. Moderate Weak Strong est. Strong est. Strong est. Strong est. Moderate Moderate est. Moderate Moderate Moderate Strong est. Strong Strong Moderate Moderate Strong Strong Moderate Strong Strong Unknown est. Moderate est. Weak est. Weak est. Weak est. Moderate est. Strong

Measures ability to write correct, idiomatic code across common programming languages — from simple utility functions to complex algorithmic problems and debugging.

What each rating means here

Strong
Writes correct, idiomatic code for complex tasks and debugs reliably across major languages; handles algorithmic and architectural decisions.
Moderate
Solid at common patterns and languages; struggles with complex algorithms, obscure libraries, or large codebase context.
Weak
Produces working code only for simple, well-trodden patterns; frequent bugs and hallucinated APIs on anything non-trivial.
Strong Strong Strong Strong est. Strong Strong Strong est. Strong Strong Strong Strong Strong est. Weak Moderate Strong est. Strong est. Strong Unknown est. Strong Strong Strong Strong Strong Strong est. Strong est. Strong Strong Strong Strong Strong Strong Strong Strong est. Strong Strong est. Strong Strong Moderate Moderate est. Moderate Moderate Unknown est. Unknown est. Unknown est. Strong est. Strong Moderate est. Moderate Moderate Moderate Strong est. Strong Strong Strong Moderate Strong Strong Moderate Moderate Moderate Moderate est. Moderate est. Moderate est. Strong est. Moderate est. Moderate est. Strong

Measures how reliably the model selects and calls external tools, APIs, and functions — including parameter formatting, multi-turn loops, and chaining dependent calls.

What each rating means here

Strong
Reliably selects and calls the correct tool with well-formed parameters; handles multi-turn tool loops and chained dependent calls gracefully.
Moderate
Gets basic tool calls right but struggles with ambiguous tool selection, malformed prior outputs, or calls requiring context from earlier turns.
Weak
Inconsistent parameter formatting; prone to hallucinating tool names or arguments, or selecting the wrong tool.
Strong Strong Strong Strong est. Strong Strong Strong est. Strong Strong Strong Strong Strong est. Moderate Strong Strong est. Moderate est. Moderate Unknown est. Strong Strong Strong Strong Strong Unknown est. Unknown est. Strong Strong Strong Strong Strong Strong Strong Strong Strong Strong Strong est. Strong Strong Unknown est. Unknown est. Unknown est. Unknown est. Moderate est. Strong Strong Strong Moderate est. Strong Strong Moderate Strong Strong Moderate Strong Strong Unknown est. Moderate est. Moderate est. Strong est. Moderate est. Strong est. Moderate

Measures how faithfully the model respects explicit constraints — output format, length limits, persona, negative instructions (what NOT to do), and multi-rule system prompts.

What each rating means here

Strong
Follows complex multi-constraint instructions precisely — respects format, length limits, persona, and negative instructions consistently.
Moderate
Handles simple instructions well but drifts on multi-constraint prompts or instructions that conflict with the model's trained defaults.
Weak
Often ignores or partially follows instructions; defaults to its own style regardless of stated format or constraint requirements.
Strong Strong Strong Strong est. Strong Strong Strong est. Strong Strong Strong Strong Moderate est. Moderate Strong Moderate est. Strong est. Strong Strong est. Strong Strong Strong Strong Strong Strong est. Strong est. Moderate Moderate Strong Strong Strong Strong Strong Strong est. Strong Strong est. Moderate Moderate Moderate Unknown est. Strong Strong Strong est. Strong est. Strong est. Strong est. Strong Moderate est. Strong Strong Strong Strong est. Strong Strong Strong Strong Strong Strong Weak Moderate Moderate Weak est. Moderate est. Moderate est. Moderate est. Strong est. Moderate est. Strong

Measures practical usable context length — how accurately the model recalls and synthesises information spread across a long input, not just the advertised token ceiling.

What each rating means here

Strong
Reliably attends to information throughout a long context — accurately recalls and synthesises content from early in a large document or conversation.
Moderate
Handles mid-length contexts well; recency bias emerges as context grows — information near the start of a long prompt is sometimes lost.
Weak
Loses track of earlier content in long contexts; practical reliable length is significantly shorter than the advertised token ceiling.
Variable
Recall quality depends on content type and its position in the context — performance varies across providers and benchmark methodologies.
Moderate Moderate Strong Moderate est. Strong Strong Strong est. Strong Strong Strong Moderate Moderate est. Moderate Moderate Moderate est. Moderate est. Moderate Unknown est. Strong Strong Strong Strong Strong Unknown est. Unknown est. Strong Strong Strong Strong Strong Moderate Strong Moderate Moderate Moderate Strong est. Moderate Moderate Unknown est. Unknown est. Unknown est. Unknown est. Moderate est. Moderate Moderate Moderate Moderate est. Strong Strong Moderate Moderate Moderate Weak Strong Strong Moderate est. Strong est. Moderate est. Moderate est. Moderate est. Moderate est. Moderate

Measures output quality across non-English languages — covering generation fluency, translation accuracy, and how well quality holds for lower-resource languages.

What each rating means here

Strong
High-quality outputs across many major languages; reliable at translation and cross-lingual reasoning with consistent grammar and factual accuracy.
Moderate
Good at widely-spoken languages (Spanish, French, German, Chinese) but quality degrades noticeably for lower-resource languages.
Weak
Primarily English-optimised; outputs in other languages may be fluent but are factually unreliable or grammatically inconsistent.
Strong Moderate Weak Strong est. Weak Weak Moderate est. Moderate Moderate Moderate Moderate Moderate est. Weak Moderate Moderate est. Strong est. Moderate Unknown est. Strong Strong Moderate Moderate Strong Unknown est. Unknown est. Moderate Strong Moderate Moderate Moderate Strong Moderate Strong est. Weak Moderate est. Strong Moderate Moderate Unknown est. Strong Strong Weak est. Moderate est. Weak est. Weak est. Strong est. Strong Strong Strong Strong est. Strong Moderate Strong Strong Weak Moderate Moderate Moderate Moderate Strong est. Strong est. Strong est. Weak est. Moderate est. Weak est. Strong

Measures output generation speed (tokens per second), which determines first-token latency for interactive use and per-token cost efficiency at scale.

What each rating means here

Strong
High throughput — fast first-token latency and generation speed well-suited to real-time, interactive, and high-volume applications.
Moderate
Adequate for most use cases; may introduce noticeable latency in latency-sensitive or streaming applications.
Weak
Slow generation; adds significant per-token latency and raises cost at scale.
Variable
Speed varies with generation length, prompt complexity, and server load — check provider latency benchmarks for your specific use case.

Ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. "—" means no profile data for that dimension.

Select models to compare
|

Video models

26 models

26 video models tracked here. Per-second pricing ranges from $0.035 to $0.16 (14 of 23 priced models bill this way; others use a different unit).

Strong Moderate Weak Variable Unknown
Skip to model selector ↓
Capability Act Two
$0.05 / s
Aleph 2
$0.01 per_credit
CogVideoX-5B
$0.2 / clip
FLUX 3 Video
$0.06 / s
Gemini Omni Flash
$0.1 / s
Gen-3 Alpha Gen-4 Turbo
$0.01 per_credit
Gen-4.5
$0.01 per_credit
Grok Imagine Video 1.5
$0.08 / s
Hailuo-02
$0.076 / clip
HappyHorse 1.0
$0.15 / s
HunyuanVideo 1.5
$0.4 / clip
Kling 1.6
$0.04 / s
Kling 2.1
$0.06 / s
LTX Video 0.9.7
$0.048 / clip
LTX-2.3
$0.04 / s
Mochi 1
$0.42 / clip
Pika 2.2
$0.2 / clip
Ray 3.2
$0.3 / clip
Seedance 2
$0.16 / s
Seedance 2.0 Seedance 2.0 Fast Seedance 2.0 Mini Sora 2
$0.1 / s
Veo 2
$0.35 / s
Veo 3
$0.1 / s
Veo 3.1
$0.05 / s
Vidu Q3
$0.035 / s
Wan 2.1
$0.09 / s
Wan 2.5
$0.05 / s
Strong Moderate Moderate Unknown Moderate est. Moderate Moderate est. Strong est. Unknown est. Moderate est. Strong Strong est. Moderate est. Strong est. Moderate Unknown est. Moderate Moderate est. Strong est. Strong est. Strong Unknown est. Unknown est. Strong est. Strong Strong est. Strong est. Unknown est. Moderate Moderate est.

Measures how realistic and coherent motion is across frames — including natural movement, fluid transitions, and absence of warping, jitter, or ghosting artifacts.

What each rating means here

Strong
Smooth, physically plausible motion with natural acceleration, fluid transitions, and minimal warping or temporal artifacts.
Moderate
Motion is mostly convincing but shows subtle jitter, unnatural acceleration, or occasional warping on complex or fast-moving scenes.
Weak
Obvious motion artifacts — jitter, warping, or frame-to-frame inconsistency undermine the realism of the clip.
Strong Strong Moderate Unknown Moderate est. Moderate Strong est. Strong est. Strong est. Moderate est. Strong Strong est. Moderate est. Strong est. Moderate Unknown est. Moderate Moderate est. Strong est. Strong est. Unknown Unknown est. Unknown est. Strong est. Strong Strong est. Strong est. Unknown est. Strong Strong est.

Measures whether subjects, backgrounds, and fine details remain stable frame-to-frame — avoiding identity drift, morphing, or unexpected visual changes mid-clip.

What each rating means here

Strong
Subjects and backgrounds remain stable throughout the clip — no identity drift, face morphing, or background flicker across frames.
Moderate
Generally stable but shows gradual drift or minor flickering on fine details, particularly in longer or complex clips.
Weak
Noticeable identity drift or background inconsistency; subjects may morph or lose visual detail mid-clip.
Weak Moderate Moderate Strong Strong est. Moderate Moderate est. Strong est. Moderate est. Moderate est. Strong Strong est. Moderate est. Strong est. Moderate Moderate est. Moderate Moderate est. Moderate est. Strong est. Strong Unknown est. Unknown est. Strong est. Strong Strong est. Strong est. Moderate est. Moderate Moderate est.

Measures how closely the generated clip matches the text description — including subject, action, style, composition, and camera direction instructions.

What each rating means here

Strong
Closely follows the text description — correct subject, action, style, and camera direction with high fidelity throughout the clip.
Moderate
Captures the main subject and action but may miss stylistic details, camera instructions, or specific secondary elements.
Weak
Captures the general mood or subject but frequently ignores specific actions, style direction, or compositional instructions.
Variable
Adherence varies with prompt complexity — reliable for simple descriptions, inconsistent for compound or highly specific instructions.
Strong Weak Weak Strong Weak est. Weak Weak est. Weak est. Strong est. Weak est. Strong Weak est. Weak est. Moderate est. Weak Strong est. Weak Weak est. Weak est. Moderate est. Strong Unknown est. Unknown est. Weak est. Weak Strong est. Strong est. Strong est. Weak Weak est.

Measures whether the model generates synchronised audio — ambient sound, foley, speech, or music — natively alongside the video rather than requiring a separate pipeline.

What each rating means here

Strong
Generates synchronised, contextually appropriate audio — ambient sound, foley, or music — that matches the on-screen action naturally.
Moderate
Produces audio alongside video but synchronisation or contextual fit is imperfect; works best for simple scenes.
Weak
No native audio support, or generated audio is clearly mismatched with the visual content.
Moderate Weak Moderate Weak est. Moderate Unknown est. Moderate Weak est. Moderate est. Weak Strong est. Weak Moderate est. Strong est. Unknown Unknown est. Unknown est. Moderate est. Moderate Strong est. Strong est. Strong est.

Measures the highest output resolution the model produces natively before quality degrades, independent of any post-processing upscaling.

What each rating means here

Strong
Produces large, high-quality native output at 1080p or above across multiple aspect ratios before any upscaling is needed.
Moderate
Solid at standard sizes; native quality degrades at higher resolutions — upscaling is needed for HD or large-format delivery.
Weak
Limited to lower native resolutions; quality degrades quickly as output dimensions increase.
Variable
Usable resolution depends on the chosen endpoint, plan, or quality setting — check provider specifications for your target output size.
Strong Moderate Weak Unknown Strong est. Moderate Strong est. Weak est. Unknown est. Moderate est. Unknown Weak est. Moderate est. Moderate est. Strong Variable est. Weak Moderate est. Moderate est. Moderate est. Moderate est. Moderate Moderate est. Moderate est. Variable est. Strong Strong est.

Measures generation latency relative to clip length. Slow inference raises cost per second of output and limits interactive or real-time applications.

What each rating means here

Strong
Fast generation — produces clips quickly relative to their length, keeping cost-per-second of output low.
Moderate
Typical generation time; fine for batch or async workflows but noticeable latency in interactive or real-time use.
Weak
Slow to generate, raising both latency and cost-per-second of output significantly relative to clip length.
Variable
Speed scales with clip length and resolution — fast for short clips at standard quality, much slower at maximum length or resolution.
Moderate Strong Moderate Unknown Weak est. Moderate Moderate est. Moderate est. Unknown est. Moderate est. Strong Moderate est. Moderate est. Moderate est. Moderate Moderate est. Weak Weak est. Moderate est. Strong est. Strong Unknown est. Unknown est. Moderate est. Moderate Moderate est. Moderate est. Strong est. Moderate Moderate est.

Measures the maximum clip length achievable in a single generation pass while maintaining consistent quality — longer ceilings reduce stitching requirements.

What each rating means here

Strong
Generates clips of 10 seconds or longer in a single pass with consistent quality throughout.
Moderate
Manages 5–10 seconds reliably; quality or coherence may degrade toward the end of longer clips.
Weak
Limited to short clips (under 5 seconds); stitching multiple clips is required for longer sequences.
Variable
Maximum duration depends on the chosen resolution or quality setting — higher settings reduce the achievable clip length.

Ratings synthesised from provider benchmarks and independent evaluations. "—" means no profile data for that dimension.

Select models to compare
|

Audio models

29 models

29 audio models tracked here. Per-1M characters pricing ranges from $4.00 to $100.00 (13 of 29 priced models bill this way; others use a different unit). 1 repriced in the last 30 days.

Strong Moderate Weak Variable Unknown

Text-to-speech

Skip to model selector ↓
Capability Amazon Polly
$4 / 1M chars
Azure Neural TTS
$4 / 1M chars
Cartesia Sonic
$65 / 1M chars
ElevenLabs Eleven v3
$100 / 1M chars
ElevenLabs Flash v2.5
$50 / 1M chars
ElevenLabs Multilingual v2
$100 / 1M chars
Fish Audio S1
$15 / 1M chars
Fish Audio S2 Pro
$15 / 1M chars
Fish Audio S2.1 Pro
$15 / 1M chars
Google Cloud TTS
$4 / 1M chars
GPT-4o Audio Preview
~$0.096 / minute (est.)
GPT-Audio 1.5
~$0.0768 / minute (est.)
Grok Voice Think Fast 2.0
$0.08 / minute
Inworld TTS-1.5 Max
$35 / 1M chars
OpenAI TTS-1
$15 / 1M chars
OpenAI TTS-1 HD
$30 / 1M chars
PlayHT 2.0
Moderate Strong Moderate Strong Moderate Strong Strong Strong Strong Strong Strong Strong Strong Strong Moderate Strong Strong

How natural and human-like the synthesized speech sounds — covering prosody, pacing, intonation, and absence of robotic or mechanical artefacts.

What each rating means here

Strong
Indistinguishable from human speech in most contexts — natural prosody, expressive intonation, and no robotic artefacts.
Moderate
Clearly intelligible and mostly natural; occasional stiffness in intonation or unnatural pauses on complex sentences.
Weak
Noticeable robotic or mechanical quality; suited for simple utility use cases but not human-facing interactions.
Moderate Strong Moderate Strong Strong Strong Strong Strong Strong Strong Moderate Moderate Unknown Unknown Weak Weak Strong

The range of distinct voices, accents, ages, and speaking styles available out of the box from the provider's voice library.

What each rating means here

Strong
Large library of high-quality voices spanning many accents, ages, and styles — easy to find the right voice for any use case.
Moderate
Good selection for common use cases; narrower range of accents or styles compared to best-in-class providers.
Weak
Few voices available; limited accent or style diversity — expect significant overlap in output character.
Weak Moderate Moderate Strong Strong Strong Strong Strong Strong Weak Weak Weak Unknown Strong Weak Weak Strong

Ability to clone a custom voice from a short audio sample, enabling personalised or brand-consistent TTS output.

What each rating means here

Strong
High-fidelity voice cloning from a short sample (< 30 s); cloned voices retain speaker identity reliably across long outputs.
Moderate
Voice cloning available but requires more audio input or produces less accurate identity preservation across varied content.
Weak
No voice cloning support, or cloning quality is too low for production use.
Strong Moderate Strong Moderate Strong Moderate Unknown Strong Strong Moderate Strong Strong Strong Strong Strong Moderate Moderate

Time-to-first-audio-chunk when using the streaming endpoint — lower latency enables real-time conversational applications.

What each rating means here

Strong
Sub-300 ms time-to-first-audio — suitable for real-time conversational agents and interactive voice applications.
Moderate
300–700 ms first-chunk latency — works for near-real-time use cases but adds noticeable delay in tightly interactive flows.
Weak
High latency (> 700 ms first chunk) — suitable only for offline or batch TTS scenarios, not real-time conversations.
Strong Strong Moderate Strong Strong Strong Moderate Strong Strong Strong Strong Strong Strong Moderate Strong Strong Strong

Number and quality of supported output languages. Strong coverage means high-quality synthesis across many major and minor languages.

What each rating means here

Strong
High-quality output across many major and minor languages with consistent accuracy or naturalness.
Moderate
Good support for widely spoken languages (English, Spanish, French, German, Chinese) but quality degrades for lower-resource languages.
Weak
Primarily English-optimised; other languages may be supported but with noticeably lower quality.
Select TTS models to compare 16 models
|

Frequently asked questions

What can I compare on Modelglass?
Pricing, architecture, and expert-synthesised capability ratings for 164 AI models — 34 image, 69 language, 30 video and 31 audio — side by side. Pick models with the selector in each section. To compare two specific models with a shareable link, open either model's page and use "Full comparison", or go to /compare/<model>.
Where do the capability ratings come from?
They are synthesised from published benchmarks, provider model cards, and independent evaluations, then expressed on one five-level scale — strong, moderate, weak, variable, unknown — so models are comparable across providers. A small "low confidence" mark next to a rating means limited benchmark data was available. Every rating links through to that model's full profile and its citations.
Why can't I compare some prices directly?
Providers bill on different units — per image, per megapixel, per 1M tokens, per second of video, per 1,000 characters, per clip. Modelglass shows each model's real billing unit rather than forcing everything onto one number. Where two models bill on different units there is no honest per-unit comparison, and the page says so instead of inventing one.
How current are the prices?
Every price carries a source URL and the date it was verified, and a repricing is recorded as a new dated entry — never a silent overwrite. The "Newly added" and "Price change" views at the top of the page surface what moved in the last 30 days. Full price history is available through the paid API.
Is it free to use?
Yes. Browsing and comparing on the site needs no account and no API key. The paid product is the read API and MCP feed, for wiring this pricing and capability data into your own routing or tooling.
What do "Fast", "Standard" and "Premium" mean?
A rough quality-and-speed tier for image models: Fast is distilled or turbo variants, Standard is the balanced default, Premium is the highest-quality option. It is a filter to narrow the table, not a Modelglass ranking.
How is this different from comparing two specific models?
This page is the catalogue view — scan a whole modality at once. For a focused head-to-head with a shareable URL, such as Claude Sonnet 5 vs GPT-5.6, use a model's own /compare/<model> page: it adds a direct-answer summary of the price and capability differences between exactly the models you pick.
How should I choose between two models that rate similarly?
Start with price on a matched unit — often the clearest separator. After that the tiebreakers on this page are context window and speed (shown as their own capability rows), how recently the price last moved (a stable price is easier to build against), and lifecycle status: a model marked deprecated or platform-only is filtered out by default for a reason. Each model's own page carries the provenance and limitations behind its ratings.