Qwen3
Alibaba’s multilingual Qwen3 family: Qwen3.8-Max (2.4T / 95B active; live alias qwen3.8-max → 0902 snapshot as of 2026-09-05) on QwenCloud, open Qwen3.8-2.4T-A95B and Qwen3.8-27B, plus Qwen3.8-Flash-Next (125B / 6B active multimodal MoE)—a Qwen4 architecture preview.
Alibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Why Qwen3 matters
Flash-Next is the efficiency path (open multimodal MoE, native 262K); cloud Flash adds 1M-default context and built-in tools. Use cloud Max for 2.4T-class vision and long-horizon work; treat HF 2.4T as text-first self-host and 27B (Apache-2.0) as the compact open VLM. Still the default non-Llama open stack for many multilingual products.
Vision · Tool calling · Thinking · Coding
Last reviewed: 7 September 2026
When to choose Qwen3
Decision guidance for architects—not a feature list.
Best for
- Multilingual applications
- Self-hosted coding assistants
- Cost-efficient agentic / long-context workloads
- Size-ladder experimentation
Avoid if
- You need a single US frontier proprietary API only
- You cannot navigate Alibaba Cloud vs HF variant differences
Strengths
Qualitative snapshot for architects—not a public ranking.
- Multilingual★★★★★
- Size ladder★★★★★
- Open ecosystem★★★★☆
- Docs consistency★★★☆☆
- Managed US vendor DX★★☆☆☆
Ecosystem
Built by
- Alibaba
Qwen3 is Alibaba’s open multilingual foundation-model family.
Competes with
- Llama 4
Main open-weight peer for self-hosted and fine-tuned deployments.
- DeepSeek R1
Competes when open reasoning and cost efficiency matter.
Works with
- vLLM
Standard high-throughput path for Qwen open weights.
- Hugging Face
Primary distribution and community surface for Qwen variants.
Recommended for
- Large language models
Reference multilingual open-model family alongside Llama.
How Qwen3 evolved
Key moments in chronological order.
- API
Qwen3.8-Max-0902 becomes the live Max alias
Model Studio routes qwen3.8-max to the 0902 snapshot from 2026-09-05 10:00 UTC+8. Alibaba describes gains in coding depth, multi-tool agentic work, and visual understanding; 1M context, thinking mode, and list pricing are unchanged. Pin qwen3.8-max-0902 to test the snapshot explicitly.
- Open source
Qwen3.8-Flash-Next open weights
Multimodal MoE (125B / 6B active + 51B n-gram embeddings) previewing the Qwen4 architecture; native 262K context (extensible to 1M). Cloud Qwen3.8-Flash is the production counterpart with 1M-default context and built-in tools.
- Open source
Dense 27B vision-language model on Hugging Face under Apache-2.0; native 262K context with image and video understanding.
- Open source
Qwen3.8-2.4T-A95B open weights
Text-first 2.4T MoE weights on Hugging Face; cloud Qwen3.8-Max keeps vision, 1M-default context, and built-in tools.
- Model
2.4T / 95B-active Max-class model on QwenCloud (qwen3.8-max) for coding and long-horizon work.
- Model
Qwen3 generation
Multilingual open models spanning chat, reasoning modes, coding, and multimodal variants.
- Product
Thinking / reasoning modes
Explicit reasoning modes make Qwen competitive on hard multi-step tasks.
- Model
Qwen2.5 momentum
Strong coding and size-ladder releases build the Qwen open ecosystem.
- Platform
vLLM serving maturity
Production serving recipes for Qwen solidify across open inference stacks.
- Platform
Hugging Face distribution
Hub becomes the default discovery path for Qwen weights and demos.
Overview
Capabilities
- Vision: Yes
- Audio: No
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- Alibaba
- License
- Qwen3.8-Max License (2.4T-A95B); Apache-2.0 on Qwen3.8-27B; qwen-community-1.0 on Qwen3.8-Flash-Next
- Context window
- 1M
- Parameters
- Max 2.4T/95B active; Flash-Next 125B/6B active + 51B n-gram; 27B dense VLM
- Release
- 2026-08
- Modalities
- Text, Image
- Vision
- Yes
- Audio
- No
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- QwenCloud / DashScope / self-host
- Pricing (output)
- QwenCloud / DashScope / self-host
Supported modalities
Text · Image
Context window
1M (1,000,000 tokens)
Pricing
Qwen3.8-Max ~$2/$6 per 1M on international QwenCloud. Flash production API announced around ¥1/¥3 per 1M (verify live USD region rates). Flash-Next and open Max/27B weights are self-host.
Availability
Cloud: qwen3.8-max aliases to qwen3.8-max-0902 from 2026-09-05 10:00 UTC+8 (Model Studio); pin qwen3.8-max-0902 to test the snapshot. Qwen3.8-Flash is the managed Flash-Next counterpart (1M default, built-in tools). Open weights: Qwen/Qwen3.8-Flash-Next (multimodal MoE; native 262K, extensible to 1M), Qwen/Qwen3.8-2.4T-A95B (text-first), Qwen/Qwen3.8-27B (dense VLM, Apache-2.0). Serve Flash-Next via vLLM / SGLang / TokenSpeed.
Use cases
- Multilingual applications
- Self-hosted coding assistants
- Cost-efficient agentic / long-context workloads
- Open multimodal pipelines
Strengths
- Flash-Next open multimodal MoE (6B active) as a Qwen4 architecture preview
- First Max-class Qwen with published open weights (2.4T-A95B)
- Qwen3.8-27B Apache-2.0 dense VLM for local/self-host coding agents
- Strong multilingual + coding/reasoning size ladder
Limitations
- Flash-Next is an experimental architecture preview; pin versions before production
- Open 2.4T weights are text-only; vision/1M-default/tools on that class are cloud Max features
- Docs/ecosystem split across QwenCloud vs Hugging Face
- 2.4T self-host footprint is datacenter-scale
Related guides
Related benchmarks
Related research
- Qwen3.8-Flash-Next (official blog)
Original Paper
- Qwen3.8-Flash-Next (Hugging Face)
Original Paper
- Qwen3.8-Max: A New Bar for Coding and Cowork
Original Paper
- Qwen3.8-Max-0902 Model Studio update notice
Evaluation
- Qwen3.8-2.4T-A95B (Hugging Face)
Original Paper
- Qwen3.8-27B (Hugging Face)
Original Paper
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- DeepSeek V4DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
- Kimi K3Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
- Muse GlimmerMetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
- Claude FableAnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
- Claude OpusAnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.