Gemini Flash
Google’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
Vision · Audio · Tool calling · Thinking · Coding
Last reviewed: 3 September 2026
Overview
Capabilities
- Vision: Yes
- Audio: Yes
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- License
- Proprietary
- Context window
- 1.0M
- Release
- 2026-09
- Modalities
- Text, Image, Audio, Video
- Vision
- Yes
- Audio
- Yes
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- No
- API
- Yes
- Pricing (input)
- ~$0.75 per 1M tokens (3.8 Flash intro through 2026-12-31)
- Pricing (output)
- ~$3.75 per 1M tokens (intro); ~$1.50 / $7.50 from 2027-01-01
Supported modalities
Text · Image · Audio · Video
Context window
1.0M (1,048,576 tokens)
Pricing
Same introductory rates as 3.7 Flash; expire Dec 31, 2026. Verify Google AI / Vertex pricing; Flash-Lite is cheaper for volume. 3.8 Flash Cyber is Fairwind-only — not a public API SKU.
Availability
Pin gemini-3.8-flash. Gemini API / AI Studio, Antigravity, Android Studio, Gemini Enterprise, and Google AI Pro/Ultra (Gemini app, Search AI Mode, Sheets). 3.7 Flash remains supported. 3.8 Flash Cyber is restricted to the Fairwind Program.
Use cases
- High-volume customer support
- Agentic tool loops and coding assistants
- Cost-sensitive RAG generation
- Mobile / edge-adjacent latency budgets
Strengths
- Same Flash speed/cost as 3.7 with stronger coding, agent, and multi-step reasoning per Google’s published evals
- Introductory $0.75/$3.75 per 1M tokens through 2026-12-31
- Large context relative to price (~1.05M)
Limitations
- Weaker than Pro-class on hardest reasoning (3.5 Pro still partner-testing as of 2026-09)
- Closed weights
- Intro token rates rise on 2027-01-01; may use more tokens at higher effort
Related guides
Related benchmarks
Related research
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Original Paper
- Gemini technical reports
Original Paper
Related GitHub
Related tools
Related rankings
Related comparisons
Companies
Explore more models
- Gemini 3.1 ProGoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
- Claude FableAnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
- Claude HaikuAnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
- Claude OpusAnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
- Claude SonnetAnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
- GPT-5.6OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.