DeepSeekOpen SourceCodingReasoning
DeepSeek V3
DeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
Tool calling · Coding
Last reviewed: 24 July 2026
Overview
DeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
Capabilities
- Vision: No
- Audio: No
- Tool calling: Yes
- Thinking: No
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- DeepSeek
- License
- Model License (open weights)
- Context window
- 128K
- Parameters
- MoE (671B total / activated subset)
- Architecture
- Mixture-of-Experts
- Release
- 2024
- Modalities
- Text
- Vision
- No
- Audio
- No
- Tool calling
- Yes
- Thinking
- No
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- DeepSeek API (competitive)
- Pricing (output)
- DeepSeek API (competitive)
Supported modalities
Text
Context window
128K (128,000 tokens)
Pricing
Input: DeepSeek API (competitive)
Output: DeepSeek API (competitive)
Also self-hostable via open weights where license allows.
Availability
API: Yes
Chat UI: Yes
Open weights: Yes
API + Hugging Face / community serving stacks.
Use cases
- Cost-efficient coding assistants
- Self-hosted enterprise chat
- Batch document processing
- Open-weight evaluation baselines
Strengths
- Strong open-weight capability / price ratio
- Good coding performance
- Self-hosting option
Limitations
- Weaker native multimodality than frontier VLMs
- Ecosystem / tooling less mature than OpenAI/Anthropic
Related guides
Related benchmarks
Related research
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- DeepSeek R1DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
- PhiMicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
- Qwen3AlibabaAlibaba’s Qwen3 generation — strong multilingual open models spanning chat, reasoning modes, coding, and multimodal variants.
- Mistral LargeMistralMistral’s flagship large model for enterprise reasoning, multilingual chat, and function calling via La Plateforme and cloud partners.
- MixtralMistralMistral’s sparse Mixture-of-Experts open models (e.g. Mixtral 8x7B / 8x22B) — efficient high-quality text generation for self-hosting.