MiMo V2.6
Xiaomi’s MiMo-V2.6 family — MIT-licensed native-omnimodal MoE with 1M context. Flagship open checkpoint is MiMo-V2.6-Pro-RL (1.02T total / 42B activated); Flash-RL is 309B / 15B activated. Hosted API ids: mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.6-pro-ultraspeed. Xiaomi publishes an Artificial Analysis Intelligence Index score of 46 for Pro; that figure is vendor-published and is not a DataAIHub ranking.
Vision · Audio · Tool calling · Thinking · Coding
Last reviewed: 22 September 2026
Overview
Capabilities
- Vision: Yes
- Audio: Yes
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- Xiaomi
- License
- MIT
- Context window
- 1.0M
- Parameters
- Pro MoE 1.02T/42B active; Flash MoE 309B/15B active
- Architecture
- Sparse MoE (SWA + global attention; MTP drafter)
- Release
- 2026-09
- Modalities
- Text, Image, Video, Audio
- Vision
- Yes
- Audio
- Yes
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- Xiaomi MiMo Open Platform (verify live rates)
- Pricing (output)
- Xiaomi MiMo Open Platform (verify live rates)
Supported modalities
Text · Image · Video · Audio
Context window
1.0M (1,048,576 tokens)
Pricing
Xiaomi says V2.6 keeps V2.5 API list prices and adds Pro UltraSpeed. Self-host via MIT weights. Do not treat third-party rate tables as official.
Availability
Open weights: XiaomiMiMo/MiMo-V2.6-Pro-RL and XiaomiMiMo/MiMo-V2.6-Flash-RL (MIT; collection XiaomiMiMo/mimo-v26). Xiaomi also released MiMo-V2.6-Distill-Qwen-9B plus RL task environments. Hosted: MiMo Open Platform, AI Studio, MiMo Code, MiMo Desktop, and listed on OpenRouter. Model-card serving recipes: SGLang (preferred) and a vLLM image based on the MiMo-V2.5 recipe — the Flash card says stable vLLM may lag. Do not assume Ollama/LM Studio day-one support.
Use cases
- Open-weight omnimodal agents (text, image, video, audio)
- Long-context coding and tool-using agents
- Self-hosted MIT-licensed deployments
- Hosted MiMo API when not self-hosting
Strengths
- MIT license on Pro-RL and Flash-RL weights
- Native text/image/video/audio with 1M context
- Documented SGLang and vLLM serve recipes on the model cards
Limitations
- Pro self-host footprint is multi-node class on the published SGLang recipe
- vLLM support may lag; card points at a MiMo-V2.5 image
- Newer ecosystem than Llama/Qwen; Xiaomi is not yet a DataAIHub company page
- Vendor benchmark/AA claims are not DataAIHub scores
Related guides
Related benchmarks
Related research
- MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement
Original Paper
- MiMo-V2.6-Pro-RL (Hugging Face)
Original Paper
- MiMo-V2.6-Flash-RL (Hugging Face)
Original Paper
Related GitHub
Related tools
Related rankings
Explore more models
- DeepSeek V4DeepSeekDeepSeek’s V4 generation — live Flash is V4.1-Flash (API id deepseek-flash): Causal Encoder–Decoder MoE, 552B backbone (8B active prefill / 16B decode), native image+text, and 1M context. MIT weights on Hugging Face. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp aliases temporarily route here. deepseek-v4-pro still serves V4-Pro-0813 after 2026-09-14 at unchanged Pro rates (DeepSeek withdrew the planned Flash reroute).
- Gemini 3.1 ProGoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
- Kimi K3Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
- Muse GlimmerMetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
- Qwen3AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview—and Qwen3.8-Omni-Flash (hosted 2026-09-17: text/image/audio/video in, text out, 1M context; API id qwen3.8-omni-flash). Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.