Reported results for Qwen3.8-Max against the frontier models it was launched against, August 2026. Qwen leads on agentic computer use and instruction following, sits mid-pack on terminal work and science reasoning, and trails on hard enterprise coding.
Qwen's weakest relative showing - roughly twelve points behind the leader.
The highest score of any model with downloadable weights.
The headline claim of the 3.8 launch - Qwen leads on agentic computer use.
Effectively saturated - all five models sit within three points.
Qwen's clearest lead: it does what it is told, literally, more reliably than its rivals.
Same data as the charts above. ≈ marks third-party approximations; - means no comparable published figure.
| Benchmark | Qwen3.8-Max | GPT-5.6 Sol | Claude Fable 5 | GLM-5.2 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| SWE-bench Pro | 67.7 | ≈64.6 | ≈80.0 | 62.1 | - |
| Terminal-Bench 2.1 | 86.6 | ≈88.8 | ≈84.6 | 81.0 | - |
| OSWorld-Verified | 86.1 | ≈83.2 | ≈85.0 | - | - |
| GPQA Diamond | 92.6 | ≈94.1 | 92.6 | 91.2 | 94.3 |
| IFBench | 82.8 | ≈72.7 | ≈63.5 | - | - |
The vendor-published result set for the current hosted flagship. PaperBench - reproducing published machine-learning research end to end - is the most striking number, and the one that best explains the model's positioning around long-horizon autonomous work rather than single-turn answers.
| Benchmark | What it measures | Score |
|---|---|---|
| OSWorld-Verified | Computer use / GUI agents | 86.1 |
| PaperBench | Reproducing published ML research | 93.0 |
| Terminal-Bench 2.1 | Terminal and shell tasks | 86.6 |
| SWE-bench Pro | Enterprise-scale bug fixing | 67.7 |
| GPQA Diamond | Graduate-level science QA | 92.6 |
| IFBench | Instruction following | 82.8 |
Among models you can actually download, the picture is simpler:
Pick on constraints first: does it need to run on your hardware, in your jurisdiction, under a licence your legal team accepts? Then test the two or three survivors on fifty real examples from your own workload. That comparison beats every table on this page.