Kimi K3
Moonshot’s 2.8T open MoE multimodal agentic model with 1M context.
Moonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Why Kimi K3 matters
Kimi K3 pushed open weights into the 3T-class MoE tier with competitive long-horizon coding and native vision. It is a primary open alternative when Llama’s ecosystem is not enough and you can accept the custom Kimi K3 License. Also rolling out in GitHub Copilot.
Vision · Tool calling · Thinking · MCP · Coding
Last reviewed: 4 September 2026
When to choose Kimi K3
Decision guidance for architects—not a feature list.
Best for
- Long-horizon open coding agents
- Open multimodal research and products
- 1M-context knowledge work
- Self-host frontier open deployments
Avoid if
- You need MIT/Apache-simple licensing for large-scale MaaS
- You cannot operate very large MoE serving footprints
Strengths
Qualitative snapshot for architects—not a public ranking.
- Open frontier scale★★★★★
- Agentic coding★★★★★
- Multimodal + long context★★★★★
- License simplicity★★☆☆☆
- Serving ease★★☆☆☆
Ecosystem
Built by
- Moonshot AI
Kimi K3 is Moonshot’s open 2.8T MoE multimodal agentic flagship.
Competes with
- Llama 4
Competes as an open generalist when frontier scale and multimodality matter.
- DeepSeek V4
Peer open MoE for agentic coding and 1M-context deployments.
- GPT-5.6
Open weights alternative for long-horizon coding and knowledge work.
Works with
Recommended for
- Large language models
Reference open 3T-class MoE in today’s LLM landscape.
- Llama Models
Compare open-weight deployability and licensing vs Llama stacks.
How Kimi K3 evolved
Key moments in chronological order.
- Platform
Kimi K3 rolls out as a Copilot model option on Pro, Pro+, Max, Business, and Enterprise plans.
- Platform
Kimi K3 on Databricks Unity AI Gateway
Databricks hosts Kimi K3 via Foundation Model API with Unity AI Gateway governance (US hosting; AWS and GCP workspaces).
- Open source
Open weights on Hugging Face
Full 2.8T checkpoint published under the Kimi K3 License (not MIT/Apache).
- API
Kimi K3 API / product launch
Moonshot launches Kimi K3 via platform.kimi.ai and consumer Kimi surfaces.
- Research
KDA + Stable LatentMoE
Kimi Delta Attention and 16-of-896 expert MoE target long-context efficiency.
- Model
Native multimodal + 1M context
Text, image, and video in one open model with a 1,048,576-token window.
Overview
Capabilities
- Vision: Yes
- Audio: No
- Tool calling: Yes
- Thinking: Yes
- MCP: Yes
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- Moonshot AI
- License
- Kimi K3 License (custom; commercial scale conditions)
- Context window
- 1.0M
- Parameters
- MoE 2.8T total / 104B activated
- Architecture
- Stable LatentMoE (KDA + Attention Residuals)
- Release
- 2026-07
- Modalities
- Text, Image, Video
- Vision
- Yes
- Audio
- No
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- Yes
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- Moonshot / Kimi API (verify platform.kimi.ai)
- Pricing (output)
- Moonshot / Kimi API (verify platform.kimi.ai)
Supported modalities
Text · Image · Video
Context window
1.0M (1,048,576 tokens)
Pricing
Also self-host via open weights; large-scale MaaS may need a separate Moonshot agreement.
Availability
API model id kimi-k3; weights at huggingface.co/moonshotai/Kimi-K3. Also rolling out in GitHub Copilot (Pro through Enterprise).
Use cases
- Long-horizon agentic coding
- Open multimodal research and products
- 1M-context knowledge work
- Self-hosted frontier open deployments
Strengths
- 2.8T open MoE with 1M context
- Native multimodal + 1M context
- Competitive agentic coding vs closed frontier peers
Limitations
- Custom license — not MIT/Apache; check commercial clauses
- Very large download / serving footprint
- Ecosystem younger than Llama/Qwen stacks
Related guides
Related benchmarks
Related research
- Kimi K3 model card (Hugging Face)
Original Paper
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- DeepSeek V4DeepSeekDeepSeek’s V4 generation — live Flash is V4.1-Flash (API id deepseek-flash): Causal Encoder–Decoder MoE, 552B backbone (8B active prefill / 16B decode), native image+text, and 1M context. MIT weights on Hugging Face. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp aliases temporarily route here. deepseek-v4-pro still serves V4-Pro-0813 after 2026-09-14 at unchanged Pro rates (DeepSeek withdrew the planned Flash reroute).
- Muse GlimmerMetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
- Qwen3AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
- Claude FableAnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
- Claude OpusAnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.