Overview
This model converts text input into audio output across 32 languages. It operates as a fast, cheap-tier model with strong audio quality, delivering quicker generation than the Multilingual variant while maintaining higher fidelity than the Flash variant. On Kyma, the model runs through an OpenAI-compatible endpoint using a single API key. It supports prompt caching, which bills repeated prompt prefixes at this model’s cached input rate. New accounts include a $0.50 free credit with no card required. Every request receives automatic failover, and responses return exact billing data in usage.cost alongside an X-Kyma-Model header to confirm routing. The model accepts a 5000-token context window and is strictly designed for speech generation. It does not support reasoning, vision, structured outputs, or multimodal inputs, and it will not handle complex logical tasks or data extraction.Specs
Pricing
Use this when
- Podcast narration generation — Produces consistent voiceovers for long-form audio content with balanced latency.
- Multilingual audiobook creation — Converts text to speech across 32 supported languages without sacrificing audio clarity.
- Interactive voice applications — Delivers fast audio responses for conversational agents where moderate latency is acceptable.
- E-learning module voiceovers — Generates clear instructional audio at a lower cost than premium TTS alternatives.
Not ideal for
Do not use this model for real-time conversational voice agents requiring sub-100ms latency or for tasks requiring structured data extraction and logical reasoning.Pick something else when
- You need maximum audio fidelity regardless of cost or speed: use
eleven-v3. - You require the fastest possible generation for high-volume text: use
eleven-flash-v2-5. - You need broader language coverage beyond the supported 32: use
eleven-multilingual-v2.