Highest AA Intelligence Index of the August frontier wave at 60.9, and the catalog's top GPQA at 94.9.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 66 models released within six months of it, across 5 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 60.9% | #2 / 22 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 94.9% | #3 / 95 |
aaAgentic | 58.7% | #3 / 16 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 76.8% | #5 / 22 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 88.2% | #7 / 62 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 89.4% | #11 / 61 |
ARC-AGIARC Prize Novel reasoning tasks requiring fluid intelligence | 68.3% | #18 / 39 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 42.9% | #26 / 97 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGrok 4.6 (High) | 42.9% |
Terminal-Bench 2.1 Artificial AnalysisGrok 4.6 (High) | 88.4% |
Terminal-Bench 4.0 Artificial AnalysisGrok 4.6 (High) | 21.2% |
SciCode Artificial AnalysisGrok 4.6 (High) | 56.5% |
AA-LCR Artificial AnalysisGrok 4.6 (High) | 80.3% |
CritPt Artificial AnalysisGrok 4.6 (High) | 17.1% |