Veo 3.1

Google DeepMind's video model with native audio. Three tiers from $0.05 to $0.60 per second, 4–8 second clips up to 4K, plus video extension.

Updated

Released
Oct 1, 2025
Type
Video Generation Model
License
Proprietary
On this page

Overview

Veo is Google DeepMind's video generation family, and Veo 3.1 is the version developers actually reach through the Gemini API. It generates 4-, 6- or 8-second clips at 24fps with natively generated audio, animates still images, extends existing clips, and accepts reference images to hold a subject consistent across shots.

The interesting thing about Veo 3.1 in August 2026 is the gap between its reputation and its measured standing. It is frequently described as the most capable video generator available. On Artificial Analysis's blind-vote text-to-video arena with audio, checked on 2 August 2026, it sits at rank 11 with an Elo of 1,095 — behind three Kling SKUs, two Alibaba HappyHorse variants, two Wan releases, ByteDance's Seedance, MiniMax H3, and Google's own Gemini Omni Flash, which leads the board at 1,245.

That does not make Veo a bad model. It makes it a model whose strengths are not the ones a preference arena measures. Veo's case rests on tiering, resolution ceiling, and the fact that it is a first-party Google API with a published rate card — a 4K output path and a $0.05-per-second Lite tier are things most of the models above it on the leaderboard do not offer.

The other structural advantage is timing. Sora 2 — the model Veo was positioned against — has its API shutdown scheduled for 24 September 2026 with no announced replacement. For teams on a Western vendor with procurement constraints, the realistic 2026 shortlist is shorter than the leaderboard suggests.

Capabilities

  • Text-to-video with native audio. Audio is generated jointly with the video rather than dubbed afterwards, and every quoted price is the with-audio rate.
  • Image-to-video. Animate a still, which in practice is the most controllable way to get a specific look — generate the frame you want with an image model first, then move it.
  • Video extension. Continue an existing clip past the 8-second ceiling. Available on standard and Fast; not on Lite.
  • Reference images. Up to three, used to guide subject and style consistency across generations. This is the mechanism for keeping a character recognisable between shots.
  • Frame interpolation. Specify first and last frames and let the model generate the motion between them — the closest thing to directing a shot rather than describing it.
  • 4K output. Standard and Fast reach 4K; Lite tops out at 1080p.

Technical Specifications

  • Model IDs: veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-preview
  • Durations: 4s, 6s, 8s
  • Frame rate: 24fps
  • Resolutions: 720p, 1080p, 4K (standard and Fast); 720p, 1080p (Lite)
  • Aspect ratios: 16:9, 9:16
  • Audio: native generation, all tiers
  • Video extension: standard and Fast only
  • Reference images: up to 3
  • Latency: minimum 11 seconds, maximum 6 minutes at peak
  • Retention: generated videos deleted after 2 days
  • Access: paid tier of the Gemini API; also Vertex AI
  • Billing: charged only on successful generation

Note the -preview suffix on all three model IDs. Google's previous generation reached end-of-life on schedule — veo-3.0-generate-001 and veo-3.0-fast-generate-001 shut down on 30 June 2026 — so treat the current IDs as pinned to a version that will eventually be retired the same way.

Use Cases

  • Short-form social video. The 8-second ceiling and 9:16 aspect ratio map directly onto the format, and the Lite tier at $0.05/second makes volume iteration affordable — roughly $0.40 per 8-second draft.
  • Storyboard-to-motion. Reference images plus first/last frame interpolation let a team move from approved stills to motion without re-describing the scene in prose. This is where Veo is genuinely stronger than prompt-only competitors.
  • Product and UI motion. Image-to-video from a rendered asset avoids the "close but wrong" problem that text-to-video has with specific products.
  • Sequences longer than 8 seconds. Video extension is the supported path, and it is a real differentiator against models that only generate isolated clips. Budget for it: extension is billed as new generation.
  • Cost-tiered pipelines. Lite for exploration, standard for the final render. The 8× price gap between Lite 720p and standard 4K is large enough to change how a team works, not just what it pays.

Performance / Benchmarks

Artificial Analysis Video Arena, text-to-video with audio, checked 2 August 2026. Rankings come from blind user votes on pairs of videos generated from the same prompt.

RankModelElo
1Gemini Omni Flash (Google)1,245
2MiniMax H31,238
3Dreamina Seedance 2.0 720p1,223
4Wan2.7-2606121,160
5HappyHorse-1.11,149
7Kling 3.0 1080p (Pro)1,111
11Veo 3.11,095

Three caveats before drawing conclusions from this.

The boards are not one board. Artificial Analysis runs separate with-audio and no-audio leaderboards, and separate text-to-video and image-to-video boards. A model's rank moves between them. Numbers quoted without naming the board — and several circulating figures for Veo, including an Elo of 1,386, do exactly that — are not comparable to the table above.

Elo measures preference, not fitness. A blind voter comparing two 8-second clips is judging aesthetics and prompt adherence. They are not judging 4K support, extension, API reliability, or whether the vendor will still be serving the endpoint next year — all of which decide production use.

Google leads its own board with a different model. Gemini Omni Flash outranks Veo 3.1 by 150 Elo. If preference score is your selection criterion and you want a Google model, Veo is not the one the data points to.

Limitations

  • 8 seconds per generation. Everything longer is stitched through extension, which costs full price per segment and is unavailable on Lite.
  • Two-day retention. Videos are deleted after 48 hours. This catches teams who treat the API as storage.
  • Preview model IDs. All three are -preview, and Google retired the previous generation on a published schedule. Plan for migration.
  • Paid tier only. No free API access; the consumer route is a Google AI subscription.
  • Two aspect ratios. 16:9 and 9:16 only — no square, no cinematic ultrawide.
  • Peak latency up to 6 minutes. Fine for batch, awkward for anything interactive.
  • Mid-table on preference. Ten SKUs rank above it on the with-audio arena. If raw output quality as judged by humans is the deciding factor, the leaderboard does not favour Veo.

Pricing & Access

Per second of generated video, audio included:

Tier720p1080p4K
Veo 3.1$0.40$0.40$0.60
Veo 3.1 Fast$0.10$0.12$0.30
Veo 3.1 Lite$0.05$0.08

What that means per clip, at the maximum 8 seconds: $3.20 on standard 1080p, $4.80 on standard 4K, $0.96 on Fast 1080p, $0.40 on Lite 720p. Billing applies only to successful generations.

The 4–8× spread between Lite and standard is the most consequential number on this page. It makes an exploration-then-render workflow economically obvious, and it is a sharper tiering than most competitors publish.

Deprecated predecessors for reference: Veo 3 standard was $0.40/second and Veo 2 was $0.35/second. Both endpoints are now shut down.

Access is through the paid tier of the Gemini API and through Vertex AI. Consumer access comes with Google AI subscriptions, where generation is metered in credits rather than seconds.

Ecosystem & Tools

  • Gemini API — the primary developer route, with the model IDs above
  • Vertex AI — the enterprise path, with the 4K tier available
  • Google AI subscriptions — consumer access, credit-metered
  • Google Antigravity — Google's agent platform, which handles image generation inline but does not expose Veo as a user-selectable model

Community & Resources

Frequently Asked Questions

Per second of generated video with audio: standard $0.40 at 720p/1080p and $0.60 at 4K; Fast $0.10 at 720p, $0.12 at 1080p, $0.30 at 4K; Lite $0.05 at 720p and $0.08 at 1080p. An 8-second 1080p clip on standard costs $3.20.
4, 6 or 8 seconds per generation at 24fps. Longer sequences are built with video extension, which the standard and Fast tiers support and Lite does not.
No. On Artificial Analysis's text-to-video arena with audio, checked 2 August 2026, Veo 3.1 sits at rank 11 with an Elo of 1,095. Google's own Gemini Omni Flash leads that board at 1,245, and the models above Veo are mostly Chinese.
Yes, natively, on all three tiers. Audio is generated together with the video rather than dubbed on afterwards, and the pricing quoted is the with-audio rate.
No. The legacy veo-3.0-generate-001 and veo-3.0-fast-generate-001 endpoints reached their shutdown date on 30 June 2026. Veo 3.1 is the current developer route.
Two days. Google deletes them after that, so download anything you intend to keep as part of the generation job rather than later.

Explore More Models

Discover other AI models and compare their capabilities.