Google DeepMind announced Gemini Omni, a new natively multimodal model family that generates and edits high-quality video using text, image, audio, and video inputs. The company introduced the first model in the lineup, Gemini Omni Flash, at Google I/O 2026.
According to reporting from Google DeepMind, Gemini Omni Flash allows video editing through natural language prompts across multiple turns. The system maintains character consistency, physics, and scene context as instructions build on previous steps.
Input capabilities and reasoning
The model combines visual rendering with reasoning about real-world physics, including gravity, kinetic energy, and fluid dynamics. It processes reference combinations to generate cohesive outputs and educational explainers. While initial audio inputs are restricted to voice, Google DeepMind plans to roll out additional audio input types alongside future output modalities for images and standalone audio.
Watermarking and distribution
All videos generated by Omni include an imperceptible SynthID digital watermark. Content origin can be verified through the Gemini app, Chrome, and Google Search. Google DeepMind restricted initial speech modification tools to personal digital avatars while testing continues on broader audio editing features.
Availability begins today for Google AI Plus, Pro, and Ultra subscribers globally through the Gemini app and Google Flow. Google DeepMind is also deploying the model at no cost on YouTube Shorts and the YouTube Create App this week, with developer and enterprise API access following in the coming weeks.
