FlagshipHosted APIMultimodal in
Qwen3.8-Max
2 Aug 2026
The current hosted flagship. A 2.4T-parameter MoE with ~95B active parameters, aimed at multi-day autonomous coding, research reproduction and parallel agent orchestration. Uses vision during execution to inspect its own intermediate output and self-correct.
- Context
- 1M tokens
- Inputs
- Text, image, video
- Model ID
- qwen3.8-max
- Price
- $2.00 in / $6.00 out per 1M
- Controls
- reasoning_effort: low / medium / xhigh
Open weightsMax-class
Qwen3.8-2.4T-A95B
12 Aug 2026
The first Max-class Qwen ever published with downloadable weights. Hybrid architecture: 92 layers, hidden size 8192, Gated DeltaNet linear attention combined with gated full attention, and a 512-expert MoE layer activating 11 experts per token (10 routed + 1 shared). Text-only, and thinking mode is always on.
- Parameters
- 2.4T total / 95B active
- Context
- 262,144 native → 1,010,000
- Licence
- Model-specific Qwen3.8-Max licence
- Sampling
- temp 1.0, top_p 0.95, top_k 20
Apache 2.0Multimodal
Qwen3.8-27B
14 Aug 2026
The new default local Qwen. A dense 27B model that Alibaba says beats the much larger hosted Qwen3.7-Plus on coding and office tasks, with stronger independent planning for agent work. Thinking mode is on by default and can be switched off per query.
- Parameters
- 27B dense
- Context
- 262K native → 1M with YaRN
- Inputs
- Text, image, video
- Licence
- Apache 2.0
Preview
Qwen3.8-Max-Preview
19 Jul 2026
Early-access build of Qwen3.8-Max, distributed through Alibaba's Token Plan and partner platforms ahead of general availability.
- Status
- Superseded by GA release
Hosted APIMultimodal
Qwen3.7-Plus
1 Jun 2026
Hosted multimodal agent model with a one-million-token context window and notably stronger visual understanding than the 3.5 line. Positioned as the workhorse tier below Max.
- Context
- 1M tokens
- Inputs
- Text, image, video
Hosted API
Qwen3.7-Max
20 May 2026
Proprietary flagship focused on coding, office automation, tool use and autonomous workflows. Never shipped open weights - the 3.7 generation was skipped for open release in favour of 3.8.
- Context
- 1M tokens
- Price
- $2.50 in / $7.50 out per 1M
Hosted API
Qwen3.7-Flash
2026
The cheapest hosted tier, aimed at high-volume classification, extraction and routing workloads that still need a large context budget.
- Context
- 1M tokens
- Price
- $0.03 in / $0.13 out per 1M
Apache 2.0
Qwen3.6-27B
22 Apr 2026
Dense 27B model that, at release, beat the far larger Qwen3.5-397B-A17B on coding suites. Superseded by Qwen3.8-27B.
- Parameters
- 27B dense
- Licence
- Apache 2.0
Apache 2.0Multimodal
Qwen3.6-35B-A3B
15 Apr 2026
Natively multimodal MoE with only ~3B parameters active per token - the cost/quality sweet spot of the 3.6 generation for self-hosting.
- Parameters
- 35B total / 3B active
- Licence
- Apache 2.0
Hosted API
Qwen3.6-Plus
1 Apr 2026
Proprietary tier restricted to Alibaba's own chatbots and Alibaba Cloud customers.
- Availability
- Alibaba Cloud only
Apache 2.0
Qwen3.5-397B-A17B
Mar 2026
The largest open model of the 3.5 generation and, for several months, the strongest downloadable Qwen. Still a common choice for self-hosted frontier-ish workloads.
- Parameters
- 397B total / 17B active
- Context
- 256K
- Benchmarks
- 88.4 GPQA Diamond · 86.7 Tau2-Bench
Apache 2.0Multimodal
Qwen3.5 (0.8B → 122B-A10B)
16 Feb 2026
The first natively multimodal Qwen generation, shipped as a full ladder: 0.8B and 2B for edge and mobile agents, 4B and 9B for local coding, 27B for single-GPU production, then 35B-A3B and 122B-A10B mixtures of experts. Architecture is roughly 75% Gated DeltaNet linear-attention layers and 25% full softmax attention.
- Sizes
- 0.8B, 2B, 4B, 9B, 27B, 35B-A3B, 122B-A10B
- Context
- 256K
- Licence
- Apache 2.0
- Note
- Qwen3.5-9B scores 81.7 GPQA Diamond
Hosted API
Qwen3.5-Plus / Flash
16 Feb 2026
The hosted tiers of the 3.5 generation. Flash remains one of the cheapest hosted models anywhere with a genuine million-token context budget.
- Plus
- $0.40 in / $2.40 out per 1M, 1M ctx
- Flash
- $0.10 in / $0.40 out per 1M, 1M ctx
Hosted APIOmni
Qwen3.5-Omni-Plus
2026
Thinker-Talker MoE architecture of roughly 100B parameters, taking text, image, audio and video and returning streamed text and speech. Supports 113 languages.
- Availability
- Hosted only (DashScope / Model Studio)
Image gen2026
Qwen-Image-3.0
22 Jul 2026
Image generation with an unusual strength: accurate typography inside the image. Handles complex layouts, precise text rendering in 12 languages, full LaTeX pages, and more than 100 art styles.
- Modality
- Text → image, image editing
Speech2026
Qwen-Audio-3.0-TTS
23 Jul 2026
Hosted text-to-speech in Flash and Plus tiers, with expressive inline controls over delivery. Supports 16 languages and continuous output of up to three minutes.
- Tiers
- Flash, Plus
- Languages
- 16
Speech2026
Qwen-Audio-3.0-ASR-Flash
31 Jul 2026
Speech recognition tuned for real deployments: better context consistency, custom hotwords and domain vocabulary, with medical and industrial term recall reported above 93%.
- Features
- Hotwords, domain terms, context consistency
Open weightsImage gen
Qwen-Image
4 Aug 2025
The 20B open-weight text-to-image and image-editing model that established Qwen's focus on commercial-grade Chinese and English text rendering inside generated images.
- Parameters
- 20B
- Licence
- Open weights
Hosted API
Qwen3-Max
24 Sep 2025
The first trillion-parameter Qwen, hosted only. Launched as an Instruct model with a Thinking variant following.
- Parameters
- 1T+
Open weightsVision
Qwen3-VL
22 Sep 2025
Vision-language line with stronger visual perception plus spatial and video reasoning, split into Instruct and Thinking variants. Qwen3-VL-2B-Instruct alone passed 18 million downloads.
- Variants
- Instruct, Thinking
Open weightsOmni
Qwen3-Omni
21 Sep 2025
End-to-end multimodal model accepting text, image, audio and video, streaming both text and speech output for real-time conversation.
- Inputs
- Text, image, audio, video
Open weightsAgentic
Qwen3-Coder
22 Jul 2025
Agent-oriented coding models in 480B-A35B and 30B-A3B sizes, shipped alongside Qwen Code, an open-source CLI for repository-scale, multi-step development tasks. The 30B-A3B build remains a favourite for Apple Silicon via MLX.
- Sizes
- 480B-A35B, 30B-A3B
- Tooling
- Qwen Code CLI
Apache 2.0
Qwen3
28 Apr 2025
The generation that brought hybrid thinking / non-thinking behaviour into the core family. Dense models at 0.6B, 1.7B, 4B, 8B, 14B and 32B plus 30B-A3B and 235B-A22B mixtures of experts, trained on 36 trillion tokens across 119 languages.
- Sizes
- 0.6B–32B dense, 30B-A3B, 235B-A22B
- Training
- 36T tokens, 119 languages
- Licence
- Apache 2.0
LegacyOmni
Qwen2.5-Omni
27 Mar 2025
The first open Omni model - multiple input modalities in, streamed text and speech out. The basis for real-time voice chat in the Qwen apps.
- Licence
- Open weights
LegacyApache 2.0
QwQ-32B
6 Mar 2025
Qwen's dedicated reasoning branch, trained with reinforcement learning for maths, code and problem-solving. Its November 2024 preview was the first open model widely compared to OpenAI's o1.
- Parameters
- 32B
- Licence
- Apache 2.0
LegacyVision
Qwen2.5-VL
26 Jan 2025
Vision-language family at 3B, 7B, 32B and 72B with much better document parsing, chart reading, long-video understanding and localisation. Processes images at arbitrary resolution.
- Sizes
- 3B, 7B, 32B, 72B
LegacyApache 2.0
Qwen2.5 & Qwen2.5-Coder
Sep–Nov 2024
The generation that made Qwen a default open baseline: a coordinated release of general, Coder and Math models with much better structured output. Qwen2.5-Coder shipped in six sizes from 0.5B to 32B and covers 92 programming languages.
- Coder sizes
- 0.5B–32B
- Languages
- 92 programming languages
Legacy
Qwen2 & Qwen1.5
2024
Qwen1.5 (Feb 2024) integrated directly with Hugging Face Transformers and introduced Qwen's first public MoE. Qwen2 (Jun 2024) improved code, maths and multilingual ability and introduced 128K context options.
- Context
- up to 128K
Legacy
Qwen 1st generation
2023
Qwen-7B and Qwen-7B-Chat (3 Aug 2023) were the first public open weights, followed by Qwen-VL (22 Aug), Qwen-14B and the Qwen-Agent framework (25 Sep), and Qwen-72B / Qwen-1.8B (30 Nov).
- Sizes
- 1.8B, 7B, 14B, 72B