Veo
GoogleVeo is Google DeepMind's video generation model line, which generates video with native audio from text and image prompts. It is served through the Gemini API and Google's Flow filmmaking tool in standard, Fast and Lite tiers, with 4K output and portrait formats.
Gemini Omni Flash, which Google presents as a separate Gemini model family and its default for video in the Gemini API, is not part of this entry.
Openness
1 high confidence- weights
- closed(no weights are distributed
- data
- closed
- code
- closed
- license
- Proprietary(Gemini API terms, which bar attempts to extract or replicate the model)
Veo is available only through Google's API and apps, and Google's terms forbid extracting the model. No weights, training data or code are published.
- https://ai.google.dev/gemini-api/docs/video recorded 2026-09-26
"The Gemini API offers two models for generating video, Gemini Omni Flash and Veo"; both are served as API models.
- https://ai.google.dev/gemini-api/terms recorded 2026-09-26
Gemini API terms: users may not "reverse engineer, extract or replicate any component of the Services".
Adoption
not assessedGoogle publishes a count of videos made in its Flow tool but no usage figure for Veo itself, and a hosted model has no download count, so no reading was possible.
- https://blog.google/innovation-and-ai/products/veo-updates-flow/ recorded 2026-09-26
Flow update post: a count of videos generated in Flow, not a Veo usage figure.
Capability
3 medium confidenceVeo 3.1 places in the upper half of the video arena, level with LTX-2.5 on the silent board. Google's newer Gemini Omni Flash and the downloadable MiniMax H3 both place well above it.
- https://artificialanalysis.ai/video/leaderboard/text-to-video recorded 2026-09-26
Without-audio leaderboard entries "Veo 3.1 Lite" (rank 17, Elo 1215.0) and "Veo 3.1" (rank 20).
Verified 2026-09-26