StorySoftware

Google unveils Gemini Omni at I/O, pairing Gemini's reasoning with Veo video generation

Google DeepMindVeo

Google introduced Gemini Omni, a model family creating any output from any mix of inputs, launching first with video generation and editing. Gemini Omni Flash enables conversational video editing where each instruction builds on previous ones, maintaining character consistency and physics. The model rolled out to Gemini app, Google Flow, and YouTube Shorts users.

  • Gemini Omni Flash launched initially with video generation and editing; future support planned for image and audio outputs
  • Conversational video editing maintains character consistency, physics accuracy, and scene context across user instructions
  • Combines images, audio, video and text as inputs for high-quality video generation using Gemini's knowledge and reasoning
  • Knowledge-grounded generation moves beyond pattern matching; demonstrates improved understanding of physical forces like gravity and fluid dynamics
Read the original · Blog

More on Veo

Primer