Skip to main content

Overview

Veo 3 is Google’s flagship video generation model. It produces 1080p clips with native audio output, handling dialogue, ambient sound, and lip-sync alignment. The model accepts text prompts and reference images. On Kyma, this model runs on a premium cost tier with slower generation speeds. It supports prompt caching, which bills repeated input prefixes at this model’s cached input rate. All requests route through Kyma’s OpenAI-compatible endpoint using a single API key, include automatic failover, and return exact billing data in usage.cost. The model does not support reasoning, vision analysis, or structured outputs. It is billed per second of generated output and is optimized for final-quality assets rather than high-throughput or low-latency workflows.

Specs

Pricing

Use this when

  • Hero brand commercials — Generate high-fidelity 1080p promotional clips with synchronized voiceovers and ambient sound.
  • Talking head avatars — Produce character videos with accurate lip-sync and native dialogue generation.
  • Cinematic scene rendering — Convert text prompts or reference images into broadcast-quality video sequences.
  • Synchronized audio output — Output clips where speech, environmental audio, and visual motion are natively aligned.

Not ideal for

Do not use this model for rapid prototyping, low-latency interactive applications, or workflows that require structured JSON outputs.

Pick something else when

Example

Agent query example

Ask the API which models fit, instead of hardcoding an id:

FAQ

Does this model accept image inputs? Yes, it accepts both text prompts and reference images to guide video generation. How are repeated prompts billed? Kyma caches repeated prompt prefixes and bills them at this model’s cached input rate. Can I test the API without adding a payment method? Yes, signup includes a $0.50 free credit and requires no card.