Overview
Gemini 3 Flash is the newest Gemini from Google, built for long-context work and reasoning. It sits in Kyma’s frontier-open quality tier and accepts text, images, audio, and video in a single request, returning text — with extended reasoning available when the problem calls for it. On Kyma, Python apps and the OpenClaw coding agent lead its traffic. Every call gets automatic failover if a serving path degrades, and prompt caching bills repeated prompt prefixes at this model’s cached input rate, which matters at this context size — resending a large cached prefix costs a fraction of the first pass. The 1,048,576-token context window comes with function calling and structured outputs, so the long context is usable inside agent pipelines, not just for one-off summarization.Specs
Pricing
Use this when
- Whole-corpus analysis — The 1M-token window fits entire codebases, document sets, or transcript archives in one request instead of a retrieval pipeline.
- Video and audio understanding — Send recordings, screen captures, or audio directly — no separate transcription step — and ask questions about what’s in them.
- Vision tasks — Screenshots, diagrams, charts, and scanned documents go in as images alongside your text prompt.
- Reasoning over long inputs — Extended reasoning plus the huge context handles analysis that requires holding a lot of material in view at once.
- Long-context agents — Function calling and structured outputs keep multi-step agents reliable even as the working context grows toward the 1M-token window.
Not ideal for
Single responses that need to run very long — output is capped at 8K tokens per request, so generating a book-length draft means chunking; and it sits in the premium cost tier, so high-volume simple tasks are cheaper on a lighter model.Pick something else when
- You want the safer long-context default: use
gemini-2.5-flash. - You want the best overall default: use
qwen-3.6-plus. - You need the strongest coding agent behavior: use
kimi-k2.6.