For the complete documentation index, see llms.txt. This page is also available as Markdown.

Gemini Omni Flash

Google's lasted model for video generation and editing.

Overview

Generates from text, image, and video inputs, and lets you refine results through natural language.

Best suited for short clips and iterative editing.

Caps at 10 seconds. Prefer Seedance 2.0 or Kling 3.0 for longer durations or production-grade cinematography.

Key updates

  • Multimodal referencing: Combine text, image, and video inputs in one generation to control composition and maintain consistency.

  • Conversational video editing: Refine and edit videos using natural language, no need to re-generate from a full prompt for small changes.

  • Real-world knowledge: Draws on Gemini's general knowledge (history, biology, narrative logic) to construct more coherent scenes.

Weaknesses

  • 10-second cap: Generations are limited to 10 seconds; longer durations are planned but not yet available.

  • Limited aspect ratios: 16:9 or 9:16.

  • Video reference limitation: Short video references (up to 3s) are accepted by the request schema but aren't correctly processed by the model yet.

  • Consistency across scene changes: Character consistency can degrade during scene changes or panning movements.

Last updated