Catalog

Every Qwen model, explained

The full Tongyi Qwen lineup as of August 2026 — hosted flagships, open weights, multimodal and media models, and the earlier generations still in wide use. Filter by what you need.

FlagshipHosted APIMultimodal in

Qwen3.8-Max

2 Aug 2026

The current hosted flagship. A 2.4T-parameter MoE with ~95B active parameters, aimed at multi-day autonomous coding, research reproduction and parallel agent orchestration. Uses vision during execution to inspect its own intermediate output and self-correct.

Context
1M tokens
Inputs
Text, image, video
Model ID
qwen3.8-max
Price
$2.00 in / $6.00 out per 1M
Controls
reasoning_effort: low / medium / xhigh
Open weightsMax-class

Qwen3.8-2.4T-A95B

12 Aug 2026

The first Max-class Qwen ever published with downloadable weights. Hybrid architecture: 92 layers, hidden size 8192, Gated DeltaNet linear attention combined with gated full attention, and a 512-expert MoE layer activating 11 experts per token (10 routed + 1 shared). Text-only, and thinking mode is always on.

Parameters
2.4T total / 95B active
Context
262,144 native → 1,010,000
Licence
Model-specific Qwen3.8-Max licence
Sampling
temp 1.0, top_p 0.95, top_k 20
Apache 2.0Multimodal

Qwen3.8-27B

14 Aug 2026

The new default local Qwen. A dense 27B model that Alibaba says beats the much larger hosted Qwen3.7-Plus on coding and office tasks, with stronger independent planning for agent work. Thinking mode is on by default and can be switched off per query.

Parameters
27B dense
Context
262K native → 1M with YaRN
Inputs
Text, image, video
Licence
Apache 2.0
Preview

Qwen3.8-Max-Preview

19 Jul 2026

Early-access build of Qwen3.8-Max, distributed through Alibaba's Token Plan and partner platforms ahead of general availability.

Status
Superseded by GA release
Hosted APIMultimodal

Qwen3.7-Plus

1 Jun 2026

Hosted multimodal agent model with a one-million-token context window and notably stronger visual understanding than the 3.5 line. Positioned as the workhorse tier below Max.

Context
1M tokens
Inputs
Text, image, video
Hosted API

Qwen3.7-Max

20 May 2026

Proprietary flagship focused on coding, office automation, tool use and autonomous workflows. Never shipped open weights - the 3.7 generation was skipped for open release in favour of 3.8.

Context
1M tokens
Price
$2.50 in / $7.50 out per 1M
Hosted API

Qwen3.7-Flash

2026

The cheapest hosted tier, aimed at high-volume classification, extraction and routing workloads that still need a large context budget.

Context
1M tokens
Price
$0.03 in / $0.13 out per 1M
Apache 2.0

Qwen3.6-27B

22 Apr 2026

Dense 27B model that, at release, beat the far larger Qwen3.5-397B-A17B on coding suites. Superseded by Qwen3.8-27B.

Parameters
27B dense
Licence
Apache 2.0
Apache 2.0Multimodal

Qwen3.6-35B-A3B

15 Apr 2026

Natively multimodal MoE with only ~3B parameters active per token - the cost/quality sweet spot of the 3.6 generation for self-hosting.

Parameters
35B total / 3B active
Licence
Apache 2.0
Hosted API

Qwen3.6-Plus

1 Apr 2026

Proprietary tier restricted to Alibaba's own chatbots and Alibaba Cloud customers.

Availability
Alibaba Cloud only
Apache 2.0

Qwen3.5-397B-A17B

Mar 2026

The largest open model of the 3.5 generation and, for several months, the strongest downloadable Qwen. Still a common choice for self-hosted frontier-ish workloads.

Parameters
397B total / 17B active
Context
256K
Benchmarks
88.4 GPQA Diamond · 86.7 Tau2-Bench
Apache 2.0Multimodal

Qwen3.5 (0.8B → 122B-A10B)

16 Feb 2026

The first natively multimodal Qwen generation, shipped as a full ladder: 0.8B and 2B for edge and mobile agents, 4B and 9B for local coding, 27B for single-GPU production, then 35B-A3B and 122B-A10B mixtures of experts. Architecture is roughly 75% Gated DeltaNet linear-attention layers and 25% full softmax attention.

Sizes
0.8B, 2B, 4B, 9B, 27B, 35B-A3B, 122B-A10B
Context
256K
Licence
Apache 2.0
Note
Qwen3.5-9B scores 81.7 GPQA Diamond
Hosted API

Qwen3.5-Plus / Flash

16 Feb 2026

The hosted tiers of the 3.5 generation. Flash remains one of the cheapest hosted models anywhere with a genuine million-token context budget.

Plus
$0.40 in / $2.40 out per 1M, 1M ctx
Flash
$0.10 in / $0.40 out per 1M, 1M ctx
Hosted APIOmni

Qwen3.5-Omni-Plus

2026

Thinker-Talker MoE architecture of roughly 100B parameters, taking text, image, audio and video and returning streamed text and speech. Supports 113 languages.

Availability
Hosted only (DashScope / Model Studio)
Image gen2026

Qwen-Image-3.0

22 Jul 2026

Image generation with an unusual strength: accurate typography inside the image. Handles complex layouts, precise text rendering in 12 languages, full LaTeX pages, and more than 100 art styles.

Modality
Text → image, image editing
Speech2026

Qwen-Audio-3.0-TTS

23 Jul 2026

Hosted text-to-speech in Flash and Plus tiers, with expressive inline controls over delivery. Supports 16 languages and continuous output of up to three minutes.

Tiers
Flash, Plus
Languages
16
Speech2026

Qwen-Audio-3.0-ASR-Flash

31 Jul 2026

Speech recognition tuned for real deployments: better context consistency, custom hotwords and domain vocabulary, with medical and industrial term recall reported above 93%.

Features
Hotwords, domain terms, context consistency
Open weightsImage gen

Qwen-Image

4 Aug 2025

The 20B open-weight text-to-image and image-editing model that established Qwen's focus on commercial-grade Chinese and English text rendering inside generated images.

Parameters
20B
Licence
Open weights
Hosted API

Qwen3-Max

24 Sep 2025

The first trillion-parameter Qwen, hosted only. Launched as an Instruct model with a Thinking variant following.

Parameters
1T+
Open weightsVision

Qwen3-VL

22 Sep 2025

Vision-language line with stronger visual perception plus spatial and video reasoning, split into Instruct and Thinking variants. Qwen3-VL-2B-Instruct alone passed 18 million downloads.

Variants
Instruct, Thinking
Open weightsOmni

Qwen3-Omni

21 Sep 2025

End-to-end multimodal model accepting text, image, audio and video, streaming both text and speech output for real-time conversation.

Inputs
Text, image, audio, video
Open weightsAgentic

Qwen3-Coder

22 Jul 2025

Agent-oriented coding models in 480B-A35B and 30B-A3B sizes, shipped alongside Qwen Code, an open-source CLI for repository-scale, multi-step development tasks. The 30B-A3B build remains a favourite for Apple Silicon via MLX.

Sizes
480B-A35B, 30B-A3B
Tooling
Qwen Code CLI
Apache 2.0

Qwen3

28 Apr 2025

The generation that brought hybrid thinking / non-thinking behaviour into the core family. Dense models at 0.6B, 1.7B, 4B, 8B, 14B and 32B plus 30B-A3B and 235B-A22B mixtures of experts, trained on 36 trillion tokens across 119 languages.

Sizes
0.6B–32B dense, 30B-A3B, 235B-A22B
Training
36T tokens, 119 languages
Licence
Apache 2.0
LegacyOmni

Qwen2.5-Omni

27 Mar 2025

The first open Omni model - multiple input modalities in, streamed text and speech out. The basis for real-time voice chat in the Qwen apps.

Licence
Open weights
LegacyApache 2.0

QwQ-32B

6 Mar 2025

Qwen's dedicated reasoning branch, trained with reinforcement learning for maths, code and problem-solving. Its November 2024 preview was the first open model widely compared to OpenAI's o1.

Parameters
32B
Licence
Apache 2.0
LegacyVision

Qwen2.5-VL

26 Jan 2025

Vision-language family at 3B, 7B, 32B and 72B with much better document parsing, chart reading, long-video understanding and localisation. Processes images at arbitrary resolution.

Sizes
3B, 7B, 32B, 72B
LegacyApache 2.0

Qwen2.5 & Qwen2.5-Coder

Sep–Nov 2024

The generation that made Qwen a default open baseline: a coordinated release of general, Coder and Math models with much better structured output. Qwen2.5-Coder shipped in six sizes from 0.5B to 32B and covers 92 programming languages.

Coder sizes
0.5B–32B
Languages
92 programming languages
Legacy

Qwen2 & Qwen1.5

2024

Qwen1.5 (Feb 2024) integrated directly with Hugging Face Transformers and introduced Qwen's first public MoE. Qwen2 (Jun 2024) improved code, maths and multilingual ability and introduced 128K context options.

Context
up to 128K
Legacy

Qwen 1st generation

2023

Qwen-7B and Qwen-7B-Chat (3 Aug 2023) were the first public open weights, followed by Qwen-VL (22 Aug), Qwen-14B and the Qwen-Agent framework (25 Sep), and Qwen-72B / Qwen-1.8B (30 Nov).

Sizes
1.8B, 7B, 14B, 72B
Naming

How to read a Qwen model name

Qwen3.8

The generation. Higher is newer; generations ship roughly every one to three months in 2026.

27B / 2.4T-A95B

Size. A single figure means a dense model. X-AY means a mixture of experts with X total parameters and Y active per token — cost tracks the active figure, quality tracks somewhere in between.

Max / Plus / Flash

Hosted service tiers, not open weights. Max is the frontier tier, Plus the balanced workhorse, Flash the cheap high-volume option.

Coder / VL / Omni / Audio / Image

Specialisation: code, vision-language, end-to-end multimodal, speech and image generation respectively.

Instruct / Thinking / Base

Post-training. Base is raw pretrained, Instruct follows instructions without visible reasoning, Thinking reasons before answering. Newer generations often fold both into one model with a toggle.

Preview

Early access ahead of general availability. Expect the production release to differ in price, quotas and sometimes quality.

Licensing

What you're allowed to do

Licences vary per model, not per generation. Always read the licence file in the specific repository you download — two models released in the same week can carry different terms.
LicenceTypical modelsCommercial useNotes
Apache 2.0Most open Qwen3, 3.5, 3.6, 3.8-27B releases YesThe permissive default; modification and redistribution allowed with attribution.
Qwen3.8-Max licenceQwen3.8-2.4T-A95B Model-specificA bespoke licence attached to the Max-class open release — read it before deploying.
Tongyi Qianwen licenceSeveral earlier large models Usually, with conditionsSource-available rather than fully open; may include usage thresholds.
Qwen Research licenceSelected research checkpoints NoNon-commercial research use only.
ProprietaryAll Max / Plus / Flash hosted tiers Via API termsNo weights; use is governed by Alibaba Cloud's service agreement.