Gemini 3.1 Pro
Google’s Pro-class Gemini for native multimodal and long-context work (current pin: gemini-3.1-pro-preview).
Google’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
Why Gemini 3.1 Pro matters
Gemini 3.1 Pro is the Google Cloud / consumer AI Pro escalate when natively multimodal inputs and very long context matter. Flash (3.8) handles volume agents; Pro is for hard reasoning. Gemini 3.5 Pro is still rolling out — verify project access.
Vision · Audio · Tool calling · Thinking · Coding
Last reviewed: 3 September 2026
When to choose Gemini 3.1 Pro
Decision guidance for architects—not a feature list.
Best for
- Native multimodal workloads
- Very long context analysis
- Google Cloud / Vertex stacks
- Research and synthesis over mixed media
Avoid if
- You are standardized only on OpenAI or Anthropic SDKs
- Open weights or fully air-gapped serving are required
Strengths
Qualitative snapshot for architects—not a public ranking.
- Multimodal nativeness★★★★★
- Long context★★★★★
- Google Cloud fit★★★★★
- Open ecosystem lock-in risk★★☆☆☆
- Open weights★☆☆☆☆
Ecosystem
Built by
- Google
Gemini Pro is Google’s flagship multimodal foundation model class.
Competes with
- GPT-5
Competes on frontier multimodal capability and developer APIs.
- Claude Opus
Rival for long-context analysis and enterprise assistant workloads.
Works with
- LangChain
Common orchestration layer in front of Gemini APIs.
Recommended for
- Large language models
Reference multimodal proprietary model in Google’s ecosystem.
- Context Windows
Often cited for very long context multimodal workloads.
How Gemini 3.1 Pro evolved
Key moments in chronological order.
- Model
Current Flash workhorse (gemini-3.8-flash) for long-horizon coding and agents; same intro $0.75/$3.75 per 1M tokens through 2026-12-31. 3.8 Flash Cyber is Fairwind-only.
- Model
Gemini 3.5 Transcribe public preview
Dedicated speech-to-text models (gemini-3.5-transcribe and gemini-3.5-transcribe-live) succeed Chirp for recorded and streaming transcription in Gemini API preview.
- Model
Prior Flash workhorse (gemini-3.7-flash) for coding and agents; intro $0.75/$3.75 per 1M tokens through 2026-12-31. Succeeded as the default workhorse by 3.8 Flash; 3.7 remains supported.
- Deprecation
gemini-2.5-pro shutdown scheduled
Google lists gemini-2.5-pro shutdown for Oct 16, 2026; migrate to gemini-3.1-pro-preview.
- Model
Flash workhorse for agentic/coding/volume; Google confirms 3.5 Pro still testing with partners.
- Model
Current Pro-class pin (gemini-3.1-pro-preview) for hard multimodal reasoning while 3.5 Pro remains in partner testing.
- Model
Gemini 2.5 Pro class (legacy)
Earlier Pro-class Gemini for advanced reasoning and native multimodal workloads.
- Product
Long-context Gemini push
Million-token-class context becomes a Gemini product differentiator.
- Model
Gemini 1.5 era
Native multimodal + long context establishes Gemini’s modern identity.
- Model
Gemini family launches
Google’s unified multimodal foundation-model brand debuts.
- Leadership
Google Brain × DeepMind
Research and model product organizations consolidate under one AI org.
Overview
Capabilities
- Vision: Yes
- Audio: Yes
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- License
- Proprietary
- Context window
- 1.0M
- Parameters
- Undisclosed
- Release
- 2026-02
- Modalities
- Text, Image, Audio, Video
- Vision
- Yes
- Audio
- Yes
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- No
- API
- Yes
- Pricing (input)
- ~$2 / 1M tokens (<200K); ~$4 at ≥200K (3.1 Pro preview)
- Pricing (output)
- ~$12 / 1M tokens (<200K); ~$18 at ≥200K
Supported modalities
Text · Image · Audio · Video
Context window
1.0M (1,048,576 tokens)
Pricing
Verify Google AI / Vertex pricing; Flash tiers are cheaper for volume.
Availability
Pin gemini-3.1-pro-preview (and -customtools when needed). Gemini 3.5 Pro not broadly GA as of 2026-08. Migrate off gemini-2.5-pro before Oct 16 2026. Gemini 3.5 Transcribe (gemini-3.5-transcribe / -live) is a separate speech-to-text preview as of 2026-08-26, not a Pro/Flash generation.
Use cases
- Long-context document and video understanding
- Multimodal research assistants
- Enterprise apps on Vertex AI
- Hard reasoning escalate above Flash
Strengths
- Very large context windows
- Strong native multimodality
- Deep Google Cloud / Vertex integration
Limitations
- Closed weights
- 3.1 Pro remains a preview ID; 3.5 Pro GA delayed
- Feature parity can differ across AI Studio vs Vertex
Related guides
Related benchmarks
Related research
- Gemini technical reports
Original Paper
Related GitHub
Related tools
Related rankings
Related comparisons
Companies
Explore more models
- Gemini FlashGoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
- Claude FableAnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
- Claude OpusAnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
- Claude SonnetAnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
- GPT-5.6OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
- GPT-6 AstraOpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.