Small open-weight Qwen 3.8 tier, holding GPQA 90.5 at 27B parameters.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with models released within six months of it.
| Benchmark | Score | Rank |
|---|---|---|
aaAgentic | 50.9% | #10 / 16 |
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 52% | #17 / 22 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 68.1% | #20 / 22 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 83.9% | #27 / 51 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 84% | #33 / 62 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 90.5% | #37 / 95 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 84.3% | #46 / 61 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 33.9% | #56 / 97 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisQwen3.8 27B (Xhigh) | 33.9% |
MMMU-Pro Artificial AnalysisQwen3.8 27B (Xhigh) | 76.3% |
Terminal-Bench 2.1 Artificial AnalysisQwen3.8 27B (Xhigh) | 79.8% |
Terminal-Bench 4.0 Artificial AnalysisQwen3.8 27B (Xhigh) | 5.6% |
SciCode Artificial AnalysisQwen3.8 27B (Xhigh) | 46.6% |
AA-LCR Artificial AnalysisQwen3.8 27B (Xhigh) | 82% |
CritPt Artificial AnalysisQwen3.8 27B (Xhigh) | 5.4% |