AlibabaOpen SourceReasoningCodingMultimodal

Qwen3

Alibaba’s multilingual Qwen3 family: Qwen3.8-Max (2.4T / 95B active; live alias qwen3.8-max → 0902 snapshot as of 2026-09-05) on QwenCloud, open Qwen3.8-2.4T-A95B and Qwen3.8-27B, plus Qwen3.8-Flash-Next (125B / 6B active multimodal MoE)—a Qwen4 architecture preview.

Alibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.

Why Qwen3 matters

Flash-Next is the efficiency path (open multimodal MoE, native 262K); cloud Flash adds 1M-default context and built-in tools. Use cloud Max for 2.4T-class vision and long-horizon work; treat HF 2.4T as text-first self-host and 27B (Apache-2.0) as the compact open VLM. Still the default non-Llama open stack for many multilingual products.

Vision · Tool calling · Thinking · Coding

Last reviewed: 7 September 2026

When to choose Qwen3

Decision guidance for architects—not a feature list.

Best for

  • Multilingual applications
  • Self-hosted coding assistants
  • Cost-efficient agentic / long-context workloads
  • Size-ladder experimentation

Avoid if

  • You need a single US frontier proprietary API only
  • You cannot navigate Alibaba Cloud vs HF variant differences

Strengths

Qualitative snapshot for architects—not a public ranking.

  • Multilingual★★★★★
  • Size ladder★★★★★
  • Open ecosystem★★★★
  • Docs consistency★★★☆☆
  • Managed US vendor DX★★☆☆☆

Ecosystem

Built by

  • Alibaba

    Qwen3 is Alibaba’s open multilingual foundation-model family.

Competes with

  • Llama 4

    Main open-weight peer for self-hosted and fine-tuned deployments.

  • DeepSeek R1

    Competes when open reasoning and cost efficiency matter.

Works with

  • vLLM

    Standard high-throughput path for Qwen open weights.

  • Hugging Face

    Primary distribution and community surface for Qwen variants.

Recommended for

How Qwen3 evolved

Key moments in chronological order.

  1. API

    Qwen3.8-Max-0902 becomes the live Max alias

    Model Studio routes qwen3.8-max to the 0902 snapshot from 2026-09-05 10:00 UTC+8. Alibaba describes gains in coding depth, multi-tool agentic work, and visual understanding; 1M context, thinking mode, and list pricing are unchanged. Pin qwen3.8-max-0902 to test the snapshot explicitly.

  2. Open source

    Qwen3.8-Flash-Next open weights

    Multimodal MoE (125B / 6B active + 51B n-gram embeddings) previewing the Qwen4 architecture; native 262K context (extensible to 1M). Cloud Qwen3.8-Flash is the production counterpart with 1M-default context and built-in tools.

  3. Open source

    Qwen3.8-27B open weights

    Dense 27B vision-language model on Hugging Face under Apache-2.0; native 262K context with image and video understanding.

  4. Open source

    Qwen3.8-2.4T-A95B open weights

    Text-first 2.4T MoE weights on Hugging Face; cloud Qwen3.8-Max keeps vision, 1M-default context, and built-in tools.

  5. Model

    Qwen3.8-Max API

    2.4T / 95B-active Max-class model on QwenCloud (qwen3.8-max) for coding and long-horizon work.

  6. Model

    Qwen3 generation

    Multilingual open models spanning chat, reasoning modes, coding, and multimodal variants.

  7. Product

    Thinking / reasoning modes

    Explicit reasoning modes make Qwen competitive on hard multi-step tasks.

  8. Model

    Qwen2.5 momentum

    Strong coding and size-ladder releases build the Qwen open ecosystem.

  9. Platform

    vLLM serving maturity

    Production serving recipes for Qwen solidify across open inference stacks.

  10. Platform

    Hugging Face distribution

    Hub becomes the default discovery path for Qwen weights and demos.

Overview

Alibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.

Capabilities

  • Vision: Yes
  • Audio: No
  • Tool calling: Yes
  • Thinking: Yes
  • MCP: No
  • Coding: Yes
  • Structured output: Yes

Technical specifications

Provider
Alibaba
License
Qwen3.8-Max License (2.4T-A95B); Apache-2.0 on Qwen3.8-27B; qwen-community-1.0 on Qwen3.8-Flash-Next
Context window
1M
Parameters
Max 2.4T/95B active; Flash-Next 125B/6B active + 51B n-gram; 27B dense VLM
Release
2026-08
Modalities
Text, Image
Vision
Yes
Audio
No
Tool calling
Yes
Thinking
Yes
MCP
No
Open weights
Yes
API
Yes
Pricing (input)
QwenCloud / DashScope / self-host
Pricing (output)
QwenCloud / DashScope / self-host

Supported modalities

Text · Image

Context window

1M (1,000,000 tokens)

Pricing

Input: QwenCloud / DashScope / self-host
Output: QwenCloud / DashScope / self-host

Qwen3.8-Max ~$2/$6 per 1M on international QwenCloud. Flash production API announced around ¥1/¥3 per 1M (verify live USD region rates). Flash-Next and open Max/27B weights are self-host.

Availability

API: Yes
Chat UI: Yes
Open weights: Yes

Cloud: qwen3.8-max aliases to qwen3.8-max-0902 from 2026-09-05 10:00 UTC+8 (Model Studio); pin qwen3.8-max-0902 to test the snapshot. Qwen3.8-Flash is the managed Flash-Next counterpart (1M default, built-in tools). Open weights: Qwen/Qwen3.8-Flash-Next (multimodal MoE; native 262K, extensible to 1M), Qwen/Qwen3.8-2.4T-A95B (text-first), Qwen/Qwen3.8-27B (dense VLM, Apache-2.0). Serve Flash-Next via vLLM / SGLang / TokenSpeed.

Use cases

  • Multilingual applications
  • Self-hosted coding assistants
  • Cost-efficient agentic / long-context workloads
  • Open multimodal pipelines

Strengths

  • Flash-Next open multimodal MoE (6B active) as a Qwen4 architecture preview
  • First Max-class Qwen with published open weights (2.4T-A95B)
  • Qwen3.8-27B Apache-2.0 dense VLM for local/self-host coding agents
  • Strong multilingual + coding/reasoning size ladder

Limitations

  • Flash-Next is an experimental architecture preview; pin versions before production
  • Open 2.4T weights are text-only; vision/1M-default/tools on that class are cloud Max features
  • Docs/ecosystem split across QwenCloud vs Hugging Face
  • 2.4T self-host footprint is datacenter-scale

Related guides

Related benchmarks

Related research

Related GitHub

Related tools

Related rankings

Companies

Explore more models

All models →