Open-weight successor to GLM 5.2 with a large agentic gain: AA Agentic Index 59.1, up from 5.2, at GPQA 91.7.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 63 models released within six months of it, across 4 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
aaAgentic | 59.1% | #2 / 16 |
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 59.5% | #5 / 22 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 74.8% | #10 / 22 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 86% | #21 / 51 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 91.7% | #29 / 95 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 42.3% | #32 / 97 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 86.8% | #35 / 61 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 80.5% | #43 / 62 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGLM-5.3 (Max) | 42.3% |
Terminal-Bench 2.1 Artificial AnalysisGLM-5.3 (Max) | 83.9% |
Terminal-Bench 4.0 Artificial AnalysisGLM-5.3 (Max) | 41.9% |
SciCode Artificial AnalysisGLM-5.3 (Max) | 59% |
AA-LCR Artificial AnalysisGLM-5.3 (Max) | 79.7% |
CritPt Artificial AnalysisGLM-5.3 (Max) | 19.1% |
ITBench SRE Artificial AnalysisGLM-5.3 (Max) | 46.1% |