Overview
Kimi K3 is Moonshot’s 2.8T open-weight multimodal model, designed as a successor to K2.7. It accepts text and image inputs and outputs text, with native support for structured outputs, tool calling, and explicit reasoning steps. The model operates with a 1,048,576-token context window and a maximum output length of 32,768 tokens. On Kyma, the model runs through an OpenAI-compatible endpoint with automatic request failover and exact cost reporting in the usage.cost field. Prompt caching is supported, reducing repeated prefix costs to this model’s cached input rate. The platform serves it at a medium speed tier. The 32,768-token output cap means it is not suited for generating extremely long documents in a single pass. As a premium-tier model, it carries higher per-token costs than lighter alternatives, and its medium throughput requires planning for latency-sensitive applications.Specs
Pricing
Use this when
- Agentic Workflow Execution — Handles multi-step coding and tool execution across extended sessions without losing context.
- Large Repository Analysis — Ingests entire codebases or documentation sets within its million-token window for cross-file reasoning.
- Vision-Enabled Document Review — Processes image inputs alongside text to extract and reason over multimodal data.
- Structured Data Extraction — Outputs strictly formatted JSON or XML for reliable downstream parsing and API integration.
Not ideal for
It is not the right choice for high-throughput, low-latency chat or generating outputs longer than 32,768 tokens.Pick something else when
- You need faster response times and lower costs for simple chat: use
deepseek-v4-flashorqwen3.7-flash. - You require outputs exceeding 32,768 tokens for long-form drafting: use
gpt-5.6-terraorgemini-3.6-flash. - You want a cheaper open-weight model for basic code completion: use
qwen-3-32bordeepseek-v3.