Skip to main content

Overview

This model converts text descriptions into non-speech audio, producing clips between 0.5 and 22 seconds. It handles prompts for environmental sounds, impacts, and mechanical noises. The context window accepts up to 500 tokens for prompt instructions. On Kyma, the endpoint operates with automatic failover and returns exact usage costs in the response payload. Prompt caching is supported, billing repeated prompt prefixes at this model’s cached input rate. The model runs on a fast, cheap tier and does not support reasoning, vision, or structured JSON outputs. Output is strictly audio. The model does not generate speech or music, and it will not return structured data formats. It is optimized for quick generation of isolated sound assets rather than long-form audio composition.

Specs

Pricing

Use this when

  • Game Foley Generation — Create impact sounds, footsteps, and environmental cues for interactive media.
  • Video Post Production — Add background ambience, weather effects, and mechanical noises to video timelines.
  • UI Interaction Sounds — Generate short clicks, swipes, and notification tones for application interfaces.
  • Podcast Audio Enhancement — Insert transitional whooshes, rain, or crowd noise to match narrative pacing.

Not ideal for

Do not use this model for generating spoken dialogue, singing, or musical compositions.

Pick something else when

Example

Agent query example

Ask the API which models fit, instead of hardcoding an id:

FAQ

What is the maximum length of the generated audio? The model outputs clips ranging from 0.5 to 22 seconds per generation. Does this model support prompt caching? Yes, repeated prompt prefixes are cached and billed at this model’s cached input rate. Can I use this for generating speech or music? No, it is strictly designed for non-speech sound effects and ambient audio.