DeepSeekOpen SourceCodingReasoning
DeepSeek V3
DeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
Tool calling · Coding
Last reviewed: 24 July 2026
Overview
DeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
Capabilities
- Vision: No
- Audio: No
- Tool calling: Yes
- Thinking: No
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- DeepSeek
- License
- Model License (open weights)
- Context window
- 128K
- Parameters
- MoE (671B total / activated subset)
- Architecture
- Mixture-of-Experts
- Release
- 2024
- Modalities
- Text
- Vision
- No
- Audio
- No
- Tool calling
- Yes
- Thinking
- No
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- DeepSeek API (competitive)
- Pricing (output)
- DeepSeek API (competitive)
Supported modalities
Text
Context window
128K (128,000 tokens)
Pricing
Input: DeepSeek API (competitive)
Output: DeepSeek API (competitive)
Also self-hostable via open weights where license allows.
Availability
API: Yes
Chat UI: Yes
Open weights: Yes
API + Hugging Face / community serving stacks.
Use cases
- Cost-efficient coding assistants
- Self-hosted enterprise chat
- Batch document processing
- Open-weight evaluation baselines
Strengths
- Strong open-weight capability / price ratio
- Good coding performance
- Self-hosting option
Limitations
- Weaker native multimodality than frontier VLMs
- Ecosystem / tooling less mature than OpenAI/Anthropic
Related guides
Related benchmarks
Related research
- DeepSeek-V3 technical report
Original Paper
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- DeepSeek R1DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
- DeepSeek V4DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
- PhiMicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
- Kimi K3Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
- Muse GlimmerMetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.