Compare Z-Image-Turbo
Z-Image-Turbo’s pricing, architecture, and capability ratings side by side with up to three other AI image-generation models. Pick models below — your selection is saved in the URL and is shareable. Z-Image-Turbo full profile ↗
Turn this comparison into a cost-aware routing setup
See how →What can I compare on this page?
Z-Image-Turbo's pricing, architecture, and capability ratings side by side with up to three other AI image-generation models. Add or remove comparison models with the selector below — the selection is saved in the page URL, so a specific comparison is shareable.
Is Z-Image-Turbo or FLUX.1 [schnell] cheaper?
They are billed on different units — Z-Image-Turbo is $0.0050 / image, FLUX.1 [schnell] is from $0.0030 / megapixel — so there is no direct per-unit comparison.
How does Z-Image-Turbo compare to FLUX.1 [schnell]?
Modelglass rates both models across 7 capability dimensions. Z-Image-Turbo rates higher on Photorealism and Text rendering.
Compare with (up to 3)
| Z-Image-Turbo base | Adobe Firefly Image 3 | Animagine XL 3.1 | DALL·E 3 (deprecated) | FLUX 1.1 [pro] | FLUX.1 [dev] | FLUX.1 [pro] | FLUX.1 [schnell] | FLUX.1 Kontext | Gemini 2.5 Flash Image | Gemini 3 Pro Image | Gemini 3.1 Flash Image | GPT Image 1 (deprecated) | GPT Image 2 | Hunyuan-DiT v1.1 (Distilled) | HunyuanImage 3.0 | Ideogram 2.0 | Ideogram 3.0 | Imagen 4 (deprecated) | Krea-2-Turbo | Leonardo Phoenix | Luma Uni-1.1 | Midjourney v7 | Recraft V3 | Runway Gen-4 Image | Seedream 4.0 | Seedream 4.5 | Seedream 5.0 Pro | Stable Diffusion 3.5 Large | Stable Diffusion 3.5 Large Turbo | Stable Diffusion 3.5 Medium | Stable Diffusion XL 1.0 | Stable Image Core | Stable Image Ultra | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Price | $0.0050 / image | $0.020 / image | $0.0037 / image | from $0.040 / image | from $0.040 / megapixel | from $0.025 / megapixel | $0.055 / image | from $0.0030 / megapixel | from $0.040 / image | $0.039 / image | from $0.13 / image | from $0.045 / image | from $0.011 / image | from $0.0060 / image | $0.00097 / second | $0.10 / megapixel | from $0.050 / image | from $0.060 / image | — | $0.0080 / megapixel | $0.0026 / credit | from $0.040 / image | from $10.00 / month | from $0.040 / image | $0.010 / credit | $0.030 / image | $0.040 / image | from $0.045 / image | $0.065 / image | $0.040 / image | $0.035 / image | $0.0014 / second | $0.030 / image | $0.080 / image |
| Creator | Alibaba Tongyi Lab | Adobe | cagliostrolab | OpenAI | Black Forest Labs | Black Forest Labs | Black Forest Labs | Black Forest Labs | Black Forest Labs | Google DeepMind | Google DeepMind | Google DeepMind | OpenAI | OpenAI | Tencent-Hunyuan | Tencent | Ideogram | Ideogram | Google DeepMind | Krea | Leonardo.Ai | Luma AI | Midjourney | Recraft | Runway | ByteDance Seed | ByteDance Seed | ByteDance Seed | Stability AI | Stability AI | Stability AI | Stability AI | Stability AI | Stability AI |
| Architecture | Diffusion transformer (DiT) | Latent diffusion | Latent diffusion | Latent diffusion | Flow matching / rectified flow | Flow matching / rectified flow | Flow matching / rectified flow | Flow matching / rectified flow | Flow matching / rectified flow | Diffusion transformer (DiT) | Diffusion transformer (DiT) | Diffusion transformer (DiT) | Autoregressive | Autoregressive | Diffusion transformer (DiT) | Autoregressive | Latent diffusion | Diffusion transformer (DiT) | Latent diffusion | Diffusion transformer (DiT) | Latent diffusion | Autoregressive | Latent diffusion | Latent diffusion | Diffusion transformer (DiT) | Diffusion transformer (DiT) | Diffusion transformer (DiT) | Diffusion transformer (DiT) | Flow matching / rectified flow | Flow matching / rectified flow | Flow matching / rectified flow | Latent diffusion | — | — |
| Released | 2025 | 2024-04 | 2024-03 | 2023-10 | 2024-10 | 2024-08 | 2024-08 | 2024-08 | 2025-05 | 2025-08 | 2025-11 | 2026-02 | 2025-04 | 2026-04 | 2024-06 | 2025-09 | 2024-08 | 2025-03 | 2025-05 | 2026-06 | 2024-08 | 2026-05 | 2025-04 | 2024-10 | 2025-03 | 2025-09 | 2025-12 | 2026-07 | 2024-10 | 2024-10 | 2024-10 | 2023-07 | 2024-10 | 2024-10 |
| Generation | Current | Current | Previous | Current | Current | Current | Previous | Current | Previous | Previous | Current | Current | Previous | Current | Current | Current | Previous | Current | Current | Current | Current | Current | Current | Current | Current | Previous | Current | Current | Current | Current | Current | Previous | Current | Current |
| Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Moderate | Strong | Unknown | Strong | Strong | Strong | Strong | Moderate | Strong | Strong | Strong | Strong | Unknown | Strong | Strong | Moderate | Strong | Moderate | Unknown | Strong | Strong | Strong | Moderate | Moderate | Moderate | — | — | |
| How faithfully the image reflects everything the prompt asked for — objects, attributes, relationships, and intent. The single most important axis for most production use, and what alignment benchmarks (GenEval, DPG-Bench, CLIPScore) try to measure. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Strong | Weak | Moderate | Strong | Strong | Strong | Moderate | Unknown | Unknown | Unknown | Unknown | Strong | Strong | Moderate | Strong | Moderate | Strong | Strong | Unknown | Strong | Unknown | Strong | Moderate | Strong | Unknown | Unknown | Strong | Strong | Strong | Moderate | Strong | — | — | |
| How convincingly the model renders real-world scenes, lighting, skin, and materials. Distinct from "looks nice" — a model can be highly aesthetic but stylised rather than photoreal. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Strong | Strong | Strong | Strong | Strong | Strong | Moderate | Unknown | Unknown | Unknown | Unknown | Strong | Strong | — | Strong | Strong | Strong | Strong | Moderate | Strong | Unknown | Strong | Strong | Strong | Unknown | Unknown | Unknown | Strong | Strong | — | Strong | — | — | |
| The breadth of styles a model can produce (illustration, painting, 3D, anime, graphic design) and how well it follows style instructions. A wide ecosystem of fine-tunes/LoRAs effectively extends this axis. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Moderate | Weak | Strong | Strong | Moderate | Strong | Moderate | Strong | Unknown | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Strong | Unknown | Moderate | Unknown | Moderate | Strong | Moderate | Strong | Strong | Strong | Strong | Moderate | Moderate | Weak | — | — | |
| The ability to render legible, correctly-spelled text inside the image (signs, logos, labels). A long-standing weak spot for diffusion models; newer transformer-backbone models are markedly better. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Moderate | Moderate | Strong | Strong | Strong | Strong | Moderate | Strong | Unknown | Moderate | Unknown | Strong | Strong | — | Strong | Moderate | Moderate | Strong | Unknown | Moderate | Strong | Moderate | Moderate | Moderate | Strong | Strong | Strong | Strong | Moderate | Moderate | Moderate | — | — | |
| Getting multi-object scenes right: correct counts, spatial relationships ("A on top of B"), and binding the right attribute to the right object (the "red cube, blue sphere" problem). Measured by GenEval and T2I-CompBench. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Moderate | Strong | Moderate | Moderate | Strong | Moderate | Strong | Moderate | Unknown | Weak | Strong | Strong | Moderate | Strong | — | Moderate | Moderate | — | Strong | Unknown | Moderate | Moderate | Strong | Strong | — | Strong | Unknown | Unknown | Moderate | Moderate | Moderate | Moderate | — | — | |
| The largest / highest-quality native output the model produces before needing upscaling, and the set of aspect ratios it supports well. Matters for hero and print assets. What each rating means here
| ||||||||||||||||||||||||||||||||||
| Strong | Moderate | Strong | Weak | Moderate | Moderate | Moderate | Strong | Strong | Strong | Unknown | Strong | Weak | Moderate | Strong | Weak | Moderate | — | Moderate | Strong | Moderate | Moderate | Moderate | Moderate | — | Strong | Unknown | Unknown | Moderate | Strong | Strong | Moderate | — | — | |
| How quickly the model produces an image, driven mainly by step count and backbone size. Directly tied to per-image cost on compute-billed hosts and to user-facing latency. What each rating means here
| ||||||||||||||||||||||||||||||||||
Capability ratings are an expert synthesis across benchmarks, community evaluations, and provider documentation. “—” means no profile data for that dimension.