Overview
Veo 3 is Google’s flagship video generation model. It produces 1080p clips with native audio output, handling dialogue, ambient sound, and lip-sync alignment. The model accepts text prompts and reference images. On Kyma, this model runs on a premium cost tier with slower generation speeds. It supports prompt caching, which bills repeated input prefixes at this model’s cached input rate. All requests route through Kyma’s OpenAI-compatible endpoint using a single API key, include automatic failover, and return exact billing data in usage.cost. The model does not support reasoning, vision analysis, or structured outputs. It is billed per second of generated output and is optimized for final-quality assets rather than high-throughput or low-latency workflows.Specs
Pricing
Use this when
- Hero brand commercials — Generate high-fidelity 1080p promotional clips with synchronized voiceovers and ambient sound.
- Talking head avatars — Produce character videos with accurate lip-sync and native dialogue generation.
- Cinematic scene rendering — Convert text prompts or reference images into broadcast-quality video sequences.
- Synchronized audio output — Output clips where speech, environmental audio, and visual motion are natively aligned.
Not ideal for
Do not use this model for rapid prototyping, low-latency interactive applications, or workflows that require structured JSON outputs.Pick something else when
- You need faster generation speeds: use
veo-3-fast. - You need lower-cost draft iterations: use
seedance-2-fast. - You need high-resolution video without audio: use
hailuo-02-1080p.