Gemini 3.1 Flash Image
Image Available Full comparison ↗ google/gemini-3-1-flash-image · by Google DeepMind
· Diffusion transformer (DiT)
Pricing — 1 offering(s)
Standard · 0.5K output
- $0.045 / image Current 2026-07-19 → present
Standard · 1K output
- $0.067 / image Current 2026-07-19 → present
Standard · 2K output
- $0.10 / image Current 2026-07-19 → present
Standard · 4K output
- $0.15 / image Current 2026-07-19 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Gemini 3.1 Flash Image fits into a cost-aware routing setup
See how →About Gemini 3.1 Flash Image
Gemini 3.1 Flash Image — codename "Nano Banana 2" — is Google DeepMind's Gemini-native image generator. Google describes it as natively multimodal: text, image, audio and video share one token space and one transformer, rather than an image pipeline bolted onto the LLM, with diffusion-style iterative denoising happening inside that shared architecture. Secondary technical analysis calls it a "Multimodal Diffusion Transformer"; that has not been confirmed against a Google-published architecture paper, so the `diffusion-transformer` classification here is the closest defensible match, not a verified specific. It is distinct from Google's separate Imagen line, a more conventional dedicated diffusion model.
The "Flash" tier is positioned for speed alongside "Pro-level" quality claims — the speed half being the more verifiable of the two. Per Google DeepMind's own model page it renders legible in-image text with font and style control, does in-image translation and localization, grounds subject accuracy in Gemini's knowledge base and web search rather than prompt parsing alone, upscales output to 2K/4K, and holds up to 5 characters and 14 objects consistent across a series of generations.
There is no independently published GenEval or T2I-CompBench score for it, so the capability ratings in this entry are vendor-stated. A press claim of a #1 Text-to-Image ranking was checked directly against Artificial Analysis's own leaderboard (2026-08-19): Nano Banana 2 sits at #3 of about 147 models tracked (Elo 1320) — strong, but behind GPT Image 2 (high) at Elo 1368. Modelglass records this entry's capability confidence as low.
Capability profile
Operator guidance
Route here for fast, cheaper Gemini-native generation and editing with strong text rendering and real-world knowledge grounding. Step up to Gemini 3 Pro Image when quality/accuracy matters more than speed (see gemini-3-pro-image.yaml), or to Imagen 4 for a more conventional dedicated diffusion pipeline within the same Google Cloud/Vertex AI stack.
Use cases
- Fast iteration where in-image text and localization matter (marketing creative across languages)
- Character-consistent series work — up to 5 characters and 14 objects held consistent across generations, per DeepMind's own model page
- Combined generation + editing in one tool, without switching models
Limitations
- No independently published GenEval / T2I-CompBench score found — capability ratings above are vendor-stated, not third-party verified
- The previously-flagged press claim of a #1 Text-to-Image ranking is now checked directly and corrected: direct fetch of artificialanalysis.ai/image/leaderboard/text-to-image (2026-08-19, SCO-462 sweep) shows Nano Banana 2 at #3 of ~147 models tracked, Elo 1320 — strong, but not #1 (that's GPT Image 2 (high), Elo 1368); the earlier press claim doesn't hold up against a direct fetch of Artificial Analysis's own site.
- Architecture classification (diffusion-transformer) is the closest vocabulary match to secondary technical reporting, not a Google-confirmed specific