Overview
This model converts text descriptions into non-speech audio, producing clips between 0.5 and 22 seconds. It handles prompts for environmental sounds, impacts, and mechanical noises. The context window accepts up to 500 tokens for prompt instructions. On Kyma, the endpoint operates with automatic failover and returns exact usage costs in the response payload. Prompt caching is supported, billing repeated prompt prefixes at this model’s cached input rate. The model runs on a fast, cheap tier and does not support reasoning, vision, or structured JSON outputs. Output is strictly audio. The model does not generate speech or music, and it will not return structured data formats. It is optimized for quick generation of isolated sound assets rather than long-form audio composition.Specs
Pricing
Use this when
- Game Foley Generation — Create impact sounds, footsteps, and environmental cues for interactive media.
- Video Post Production — Add background ambience, weather effects, and mechanical noises to video timelines.
- UI Interaction Sounds — Generate short clicks, swipes, and notification tones for application interfaces.
- Podcast Audio Enhancement — Insert transitional whooshes, rain, or crowd noise to match narrative pacing.
Not ideal for
Do not use this model for generating spoken dialogue, singing, or musical compositions.Pick something else when
- You need realistic voice narration: use
eleven-multilingual-v2. - You require background music tracks: use
elevenlabs-music. - You want fast, low-latency voice synthesis: use
eleven-turbo-v2-5.