Very large Qwen MoE (2.4T total, 95B active) posting the family's highest GPQA at 93.5.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with models released within six months of it.
| Benchmark | Score | Rank |
|---|---|---|
aaAgentic | 57.1% | #6 / 16 |
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 57.7% | #8 / 22 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 93.5% | #12 / 95 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 71.9% | #15 / 22 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 42.4% | #30 / 97 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisQwen3.8 2.4T A95B | 42.4% |
Terminal-Bench 2.1 Artificial AnalysisQwen3.8 2.4T A95B | 82% |
Terminal-Bench 4.0 Artificial AnalysisQwen3.8 2.4T A95B | 11.1% |
SciCode Artificial AnalysisQwen3.8 2.4T A95B | 54.1% |
AA-LCR Artificial AnalysisQwen3.8 2.4T A95B | 80.3% |
CritPt Artificial AnalysisQwen3.8 2.4T A95B | 20% |