Updates

Qwen release timeline

Every significant Tongyi Qwen model, product and organisational update from the April 2023 announcement through August 2026 - newest first. Highlighted entries mark the current generation.

Release cadence accelerated sharply in 2026: Alibaba shipped four numbered generations (3.5 → 3.8) between February and August, alongside separate image and audio releases. Anything more than a few months old is likely superseded.
14 August 2026

Qwen3.8-27B released under Apache 2.0

A dense 27B model with 262K native context (extensible to 1M via YaRN), multimodal input across text, images and video, and a thinking mode that is on by default but can be toggled per query. Qwen claims it outperforms the much larger hosted Qwen3.7-Plus on coding and office tasks, and it immediately became the default recommendation for local and single-GPU deployment.

12 August 2026

Day-0 inference support for Qwen3.8-2.4T-A95B

vLLM shipped support on the day of release, making the largest open Qwen servable on standard inference stacks without waiting for a release cycle.

8 August 2026

Qwen3.8-2.4T-A95B open weights

The first Max-class Qwen ever published with downloadable weights: 2.4T total parameters, 95B active, 92 layers, a 512-expert MoE activating 11 experts per token, and a hybrid of Gated DeltaNet linear attention with gated full attention. Text-only, always in thinking mode, 262,144 native context extensible to 1,010,000 tokens.

2 August 2026

Qwen3.8-Max reaches general availability

The hosted flagship launched with a 1M-token context window, native text/image/video input, a reasoning_effort control and built-in tools. Alibaba positioned it for multi-day autonomous coding, research reproduction and parallel agent orchestration, claiming leadership on agentic computer-use benchmarks. Priced at $2 in / $6 out per million tokens.

31 July 2026

Qwen-Audio-3.0-ASR-Flash

Speech recognition with improved context consistency, custom hotwords and domain-term handling; reported recall above 93% on medical and industrial terminology.

23 July 2026

Qwen-Audio-3.0-TTS in Flash and Plus tiers

Hosted text-to-speech across 16 languages with expressive inline delivery controls and continuous output of up to three minutes.

22 July 2026

Qwen-Image-3.0

Image generation handling complex layouts, precise text rendering in 12 languages, full LaTeX pages and more than 100 art styles.

19 July 2026

Qwen3.8-Max-Preview

Early access through Alibaba's Token Plan and partner platforms, two weeks ahead of general availability.

1 June 2026

Qwen3.7-Plus

Hosted multimodal agent model with a one-million-token context window and substantially stronger visual understanding than the 3.5 line.

20 May 2026

Qwen3.7-Max

Proprietary flagship targeting coding, office automation, tool use and autonomous workflows. The 3.7 generation never received an open-weight release - it was skipped in favour of 3.8.

22 April 2026

Qwen3.6-27B

A dense 27B Apache 2.0 model that beat the previous generation's 397B-A17B flagship on coding suites - an early sign of how fast the small-dense line was improving.

15 April 2026

Qwen3.6-35B-A3B

Natively multimodal mixture of experts under Apache 2.0, with roughly 3B parameters active per token.

1 April 2026

Qwen3.6-Plus

Proprietary model made available only through Alibaba's own chatbots and Alibaba Cloud.

March 2026

Qwen3.5-397B-A17B and a leadership change

The largest open model of the 3.5 generation arrived (88.4 GPQA Diamond, 86.7 Tau2-Bench). In the same month Qwen division head Lin Junyang departed, prompting public debate about the future of Alibaba's open-weight commitment.

16 February 2026

Qwen3.5 - the first natively multimodal generation

A full ladder of open-weight models from 0.8B to 122B-A10B, all with 256K context under Apache 2.0, handling text, images, video, reasoning and agent tasks natively rather than through bolted-on adapters. The hosted Qwen3.5-Plus and Qwen3.5-Flash tiers launched alongside.

January 2026

Qwen App wired into Alibaba's ecosystem

A mobile update connected Qwen to Taobao and Taobao Instant Commerce, Alipay, Fliggy and Amap, turning the assistant into an executor: ordering food, booking travel, placing phone calls with transcripts, and batch-processing up to 100 documents. Alibaba also reorganised its AI operations into a group called Token Hub under CEO Eddie Wu.

17 November 2025

Qwen App public beta

Alibaba's consumer assistant entered public beta in China and passed 100 million monthly active users within two months.

24 September 2025

Qwen3-Max

Qwen's first trillion-parameter model, hosted only, launched as an Instruct release with a Thinking variant to follow.

22 September 2025

Qwen3-VL

Vision-language models with stronger visual perception plus spatial and video reasoning, in separate Instruct and Thinking variants.

21 September 2025

Qwen3-Omni

End-to-end multimodal model accepting text, images, audio and video and streaming both text and speech back.

4 August 2025

Qwen-Image

A 20B open-weight text-to-image and editing model with unusual attention to rendering readable text - in both Chinese and English - inside generated images.

22 July 2025

Qwen3-Coder and the Qwen Code CLI

Agentic coding models at 480B-A35B and 30B-A3B, paired with an open-source command-line tool for repository-scale, multi-step development work.

28 April 2025

Qwen3

Dense models at 0.6B, 1.7B, 4B, 8B, 14B and 32B plus 30B-A3B and 235B-A22B MoE variants, all Apache 2.0, trained on 36 trillion tokens across 119 languages. This is the generation that brought hybrid thinking / non-thinking behaviour into the core family.

27 March 2025

Qwen2.5-Omni

The first open Omni model: multiple input modalities in, streamed text and speech out, enabling real-time voice conversation.

6 March 2025

QwQ-32B

An open reasoning model trained with reinforcement learning for maths, coding and problem-solving.

29 January 2025

Qwen2.5-Max

A large-scale mixture-of-experts model, released days after DeepSeek's R1 shook the market.

26 January 2025

Qwen2.5-VL

Vision-language family at 3B, 7B, 32B and 72B with much better document parsing, chart reading, long video understanding and localisation.

28 November 2024

QwQ-32B-Preview

The experimental model that opened Qwen's dedicated reasoning branch and was the first open model widely benchmarked against OpenAI's o1.

12 November 2024

Qwen2.5-Coder

The coding family expanded to six sizes from 0.5B to 32B, covering 92 programming languages.

19 September 2024

Qwen2.5

A coordinated release of general, Coder and Math models with much better structured output - the generation that made Qwen a default open baseline worldwide.

29 August 2024

Qwen2-VL

Added dynamic image resolution, video understanding and visual localisation at 2B, 7B and 72B.

9 August 2024

Qwen2-Audio

Apache 2.0 audio-language model interpreting speech, sound and music directly, without a separate transcription step.

7 June 2024

Qwen2

Second generation, with better code, maths and multilingual ability and 128K context options.

16 April 2024

CodeQwen1.5

The first dedicated Qwen coding family, separating code-focused training from general chat.

4 February 2024

Qwen1.5

Wider size range and direct integration with Hugging Face Transformers; later added Qwen's first public mixture-of-experts model.

30 November 2023

Qwen-72B and Qwen-1.8B

The first generation extended in both directions at once - a large flagship and a compact model for constrained hardware.

25 September 2023

Qwen-14B and Qwen-Agent

A mid-sized model plus an open-source agent framework for tool-using applications. Tongyi Qianwen also opened to the Chinese public this month after regulatory approval.

22 August 2023

Qwen-VL and Qwen-VL-Chat

The first vision-language models, supporting image understanding, text reading and visual grounding.

3 August 2023

Qwen-7B and Qwen-7B-Chat

The first publicly released Qwen open weights - base and chat variants. The beginning of the open-weight strategy that now accounts for over 40 million downloads.

11 April 2023

Tongyi Qianwen announced

Alibaba Cloud unveiled its large language model with enterprise beta access in China.

Reading the trend

Three things the timeline shows

Small models are eating large ones

Qwen3.6-27B beat the previous generation's 397B flagship on coding; Qwen3.8-27B beats the hosted 3.7-Plus. If you deployed a huge open model six months ago, a dense 27B is probably now cheaper and better.

The open/closed line keeps moving

2026 opened with proprietary Plus and Max tiers and worried commentary about Alibaba retreating from open source. It ended the first half by publishing Max-class weights - under a bespoke licence rather than Apache 2.0.

Agents, not chat

Every 2026 release is framed around autonomy: computer use, terminal tasks, long-horizon execution, tool calling, and in the consumer app, actually placing orders and phone calls.