Qwen flagship successor to 3.7 Max. AA Intelligence Index 58.1 with a large agentic jump (58.4), at GPQA 92.7.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 68 models released within six months of it, across 4 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
aaAgentic | 58.4% | #4 / 16 |
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 58.1% | #7 / 22 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 87.9% | #10 / 62 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 88% | #12 / 51 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 71.8% | #16 / 22 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 88.6% | #19 / 61 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 92.8% | #20 / 95 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 43.1% | #24 / 97 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisQwen3.8 Max (0902) | 43.1% |
MMMU-Pro Artificial AnalysisQwen3.8 Max (0902) | 82.8% |
Terminal-Bench 2.1 Artificial AnalysisQwen3.8 Max (0902) | 88.8% |
Terminal-Bench 4.0 Artificial AnalysisQwen3.8 Max (0902) | 38.9% |
SciCode Artificial AnalysisQwen3.8 Max (0902) | 52.1% |
AA-LCR Artificial AnalysisQwen3.8 Max (0902) | 80.3% |
CritPt Artificial AnalysisQwen3.8 Max (0902) | 17.7% |
ITBench SRE Artificial AnalysisQwen3.8 Max (0902) | 40.3% |