Gemini Omni Flash
Google's lasted model for video generation and editing.
Last updated
Google's lasted model for video generation and editing.
Generates from text, image, and video inputs, and lets you refine results through natural language.
Best suited for short clips and iterative editing.
Caps at 10 seconds. Prefer Seedance 2.0 or Kling 3.0 for longer durations or production-grade cinematography.
Multimodal referencing: Combine text, image, and video inputs in one generation to control composition and maintain consistency.
Conversational video editing: Refine and edit videos using natural language, no need to re-generate from a full prompt for small changes.
Real-world knowledge: Draws on Gemini's general knowledge (history, biology, narrative logic) to construct more coherent scenes.
10-second cap: Generations are limited to 10 seconds; longer durations are planned but not yet available.
Limited aspect ratios: 16:9 or 9:16.
Video reference limitation: Short video references (up to 3s) are accepted by the request schema but aren't correctly processed by the model yet.
Consistency across scene changes: Character consistency can degrade during scene changes or panning movements.
Last updated

