Benchmarks

How Qwen measures up

Reported results for Qwen3.8-Max against the frontier models it was launched against, August 2026. Qwen leads on agentic computer use and instruction following, sits mid-pack on terminal work and science reasoning, and trails on hard enterprise coding.

Read these carefully. Figures marked are third-party comparisons rather than numbers published by the model's own vendor, and every score here is self-reported or press-reported rather than independently replicated. Benchmark harnesses, prompting and scaffolding differ between labs, so treat two-point gaps as noise and use these to shortlist candidates for your own evaluation, not to settle an argument.

SWE-bench Pro

Complex enterprise software engineering. Higher is better; scale 0–100.
Claude Fable 5
≈80.0
Qwen3.8-Max
67.7
GPT-5.6 Sol
≈64.6
GLM-5.2
62.1
Claude Fable 5Qwen3.8-MaxGPT-5.6 SolGLM-5.2

Qwen's weakest relative showing - roughly twelve points behind the leader.

Terminal-Bench 2.1

Command-line and shell task completion. Higher is better; scale 0–100.
GPT-5.6 Sol
≈88.8
Qwen3.8-Max
86.6
Claude Fable 5
≈84.6
GLM-5.2
81.0
GPT-5.6 SolQwen3.8-MaxClaude Fable 5GLM-5.2

The highest score of any model with downloadable weights.

OSWorld-Verified

Autonomous desktop GUI interaction (“computer use”). Higher is better; scale 0–100.
Qwen3.8-Max
86.1
Claude Fable 5
≈85.0
GPT-5.6 Sol
≈83.2
Qwen3.8-MaxClaude Fable 5GPT-5.6 Sol

The headline claim of the 3.8 launch - Qwen leads on agentic computer use.

GPQA Diamond

Graduate-level science reasoning. Higher is better; scale 0–100.
Gemini 3.1 Pro
94.3
GPT-5.6 Sol
≈94.1
Qwen3.8-Max
92.6
Claude Fable 5
92.6
GLM-5.2
91.2
Gemini 3.1 ProGPT-5.6 SolQwen3.8-MaxClaude Fable 5GLM-5.2

Effectively saturated - all five models sit within three points.

IFBench

Precise instruction following under constraints. Higher is better; scale 0–100.
Qwen3.8-Max
82.8
GPT-5.6 Sol
≈72.7
Claude Fable 5
≈63.5
Qwen3.8-MaxGPT-5.6 SolClaude Fable 5

Qwen's clearest lead: it does what it is told, literally, more reliably than its rivals.

Table view

All comparison figures

Same data as the charts above. ≈ marks third-party approximations; - means no comparable published figure.

BenchmarkQwen3.8-MaxGPT-5.6 SolClaude Fable 5GLM-5.2Gemini 3.1 Pro
SWE-bench Pro67.7≈64.6≈80.062.1-
Terminal-Bench 2.186.6≈88.8≈84.681.0-
OSWorld-Verified86.1≈83.2≈85.0--
GPQA Diamond92.6≈94.192.691.294.3
IFBench82.8≈72.7≈63.5--
Scorecard

Qwen3.8-Max in full

The vendor-published result set for the current hosted flagship. PaperBench - reproducing published machine-learning research end to end - is the most striking number, and the one that best explains the model's positioning around long-horizon autonomous work rather than single-turn answers.

BenchmarkWhat it measuresScore
OSWorld-VerifiedComputer use / GUI agents86.1
PaperBenchReproducing published ML research93.0
Terminal-Bench 2.1Terminal and shell tasks86.6
SWE-bench ProEnterprise-scale bug fixing67.7
GPQA DiamondGraduate-level science QA92.6
IFBenchInstruction following82.8

Open-weight leaderboard context

Among models you can actually download, the picture is simpler:

  • Qwen3.8-2.4T-A95B - the strongest open Qwen, and the first Max-class open release.
  • Qwen3.8-27B - Artificial Analysis intelligence score 52, versus 38 for Qwen3.6-27B four months earlier.
  • Qwen3.5-397B-A17B - 88.4 GPQA Diamond, 86.7 Tau2-Bench; the previous open flagship.
  • Qwen3.5-9B - 81.7 GPQA Diamond from a model that fits on one consumer GPU.

Choosing without benchmarks

Pick on constraints first: does it need to run on your hardware, in your jurisdiction, under a licence your legal team accepts? Then test the two or three survivors on fifty real examples from your own workload. That comparison beats every table on this page.