Google DeepMindGA
Gemini Omni Flash
Context
1.05M
Max output
—
Input /1M
—
Output /1M
—
Capabilities
- Input: text, image and video (video up to 10s for editing and extension); output: video
- Output video 3s-10s at 360p, 720p, 1080p or 4K, 24 FPS; 1,048,576-token context window
- Text- and image-to-video generation, conversational editing through the Interactions API, video extension, resolution upscaling and interpolation (vendor-described)
- Paid tier only, USD per million tokens: input $1.50 (text/image/video/audio); output $9.00 text, $17.50 video. Google bills video output at 5,792 tokens per second of 720p video, about $0.10 per second.