Overview
Gemini 3.5 Flash is the newest Flash-class model from Google, built for long context, multimodal input, and fast reasoning. It sits in Kyma’s frontier-open quality tier and takes text, images, audio, and video as input while returning text. On Kyma it runs through the same OpenAI-compatible endpoint as every other model, with automatic failover if a serving path degrades. Prompt caching is supported, so repeated prompt prefixes — long system prompts, large documents you query more than once — bill at 10% of the input rate. Every response reports its exact cost in usage.cost. The 1M-token context window is the headline capability: whole codebases, long transcripts, or large document sets fit in a single request. Function calling, structured outputs, and reasoning support round it out for agent work, not just one-shot prompts.Specs
Pricing
Use this when
- Video and audio analysis — Feed it video or audio directly — summarize recordings, extract what was said and shown, no separate transcription step.
- Whole-corpus questions — The 1M context takes an entire codebase, contract set, or research archive in one request instead of a chunked retrieval pipeline.
- Vision pipelines — Image understanding with structured outputs turns screenshots, documents, and photos into clean JSON your code consumes.
- Fast reasoning at scale — A fast-tier model that also supports reasoning — multi-step analysis without leaving the fast tier.
- Multimodal agents — Tool calling plus four input modalities lets one agent handle text, screenshots, and recordings without switching models.
Not ideal for
Long single-shot generations — output caps at 8K tokens — or cost-sensitive bulk text work, where its premium pricing buys multimodal range you wouldn’t be using.Pick something else when
- You want the cheapest long-context option: use
gemini-2.5-flash. - You need the strongest tool-heavy agent behavior: use
kimi-k2.6. - You want top reasoning over speed: use
deepseek-v4-pro.