Gemini Omni Flash
Video Preview google-deepmind/gemini-omni-flash · by Google DeepMind · Diffusion transformer (DiT)
Pricing — 1 offering(s)
Standard (720p)
- $0.10 / second Current 2026-06-30 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
Capability profile
Operator guidance
Default choice when editing capability and low latency are required and native audio is not needed. At $0.10/s (720p only), this is the cheapest option in the registry with API-level video editing. For silent T2V/I2V at lower cost, use Veo 3.1 Lite ($0.05–0.08/s); for native audio or 4K, use Veo 3.1 Standard/Fast. The Interactions API's stateful session model (previous_interaction_id chains turns) enables multi-turn refinement — note that conversational session management is an API design pattern, not a product-only chat layer. The edit task (V2V) has an active limitation with video input processing as of 2026-07-02 — verify current status before committing to production editing workflows. Geo restriction: externally-uploaded video editing is unavailable in EEA, Switzerland, and UK.
Use cases
- Fast T2V/I2V for latency-sensitive production pipelines
- Multi-turn conversational video editing sessions (Interactions API)
- Reference-conditioned generation anchored to character or subject images
- Cost-efficient video generation with editing capability at $0.10/s
Limitations
- Preview status; pricing and capabilities subject to change before GA
- 720p only — no 1080p or 4K output
- No native audio; audio input references unsupported
- Max 10 seconds per clip; no video extension mode
- edit task (V2V) — video references up to 3s accepted by API schema but not correctly processed (active limitation, 2026-07-02)
- Editing of externally-uploaded video geo-restricted (unavailable EEA, Switzerland, UK)
- Multi-video referencing unsupported
- System instructions, temperature, and top_p parameters not supported
- No mask-guided inpainting
- Input token costs ($1.50/M) apply separately — significant for video editing inputs