← All models

Gemini 3.1 Flash Image

Image Available Full comparison ↗

google/gemini-3-1-flash-image · by Google DeepMind · Diffusion transformer (DiT)

Pricing — 1 offering(s)

Standard · 0.5K output

  • $0.045 / image Current 2026-07-19 → present

Standard · 1K output

  • $0.067 / image Current 2026-07-19 → present

Standard · 2K output

  • $0.10 / image Current 2026-07-19 → present

Standard · 4K output

  • $0.15 / image Current 2026-07-19 → present

Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.

See how Gemini 3.1 Flash Image fits into a cost-aware routing setup

See how →

About Gemini 3.1 Flash Image

Gemini 3.1 Flash Image — codename "Nano Banana 2" — is Google DeepMind's Gemini-native image generator. Google describes it as natively multimodal: text, image, audio and video share one token space and one transformer, rather than an image pipeline bolted onto the LLM, with diffusion-style iterative denoising happening inside that shared architecture. Secondary technical analysis calls it a "Multimodal Diffusion Transformer"; that has not been confirmed against a Google-published architecture paper, so the `diffusion-transformer` classification here is the closest defensible match, not a verified specific. It is distinct from Google's separate Imagen line, a more conventional dedicated diffusion model.

The "Flash" tier is positioned for speed alongside "Pro-level" quality claims — the speed half being the more verifiable of the two. Per Google DeepMind's own model page it renders legible in-image text with font and style control, does in-image translation and localization, grounds subject accuracy in Gemini's knowledge base and web search rather than prompt parsing alone, upscales output to 2K/4K, and holds up to 5 characters and 14 objects consistent across a series of generations.

There is no independently published GenEval or T2I-CompBench score for it, so the capability ratings in this entry are vendor-stated. A press claim of a #1 Text-to-Image ranking was checked directly against Artificial Analysis's own leaderboard (2026-08-19): Nano Banana 2 sits at #3 of about 147 models tracked (Elo 1320) — strong, but behind GPT Image 2 (high) at Elo 1368. Modelglass records this entry's capability confidence as low.

Capability profile

Text rendering strong
Prompt adherence strong
Resolution ceiling strong
Inference speed strong
Photorealism unknown
Artistic style range unknown
Compositional accuracy unknown

Operator guidance

Route here for fast, cheaper Gemini-native generation and editing with strong text rendering and real-world knowledge grounding. Step up to Gemini 3 Pro Image when quality/accuracy matters more than speed (see gemini-3-pro-image.yaml), or to Imagen 4 for a more conventional dedicated diffusion pipeline within the same Google Cloud/Vertex AI stack.

Use cases

Limitations

Citations