Meta's next Superintelligence Labs iteration, lifting AA Intelligence Index from 50.6 to 56.8.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 66 models released within six months of it, across 3 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 56.8% | #11 / 22 |
aaAgentic | 49.3% | #12 / 16 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 72.2% | #14 / 22 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 88.3% | #20 / 61 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 86.1% | #20 / 51 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 45.5% | #21 / 97 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 90.4% | #39 / 95 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisMuse Spark 1.2 (Xhigh) | 45.5% |
Terminal-Bench 2.1 Artificial AnalysisMuse Spark 1.2 (Xhigh) | 80.1% |
Terminal-Bench 4.0 Artificial AnalysisMuse Spark 1.2 (Xhigh) | 7.1% |
SciCode Artificial AnalysisMuse Spark 1.2 (Xhigh) | 57.4% |
AA-LCR Artificial AnalysisMuse Spark 1.2 (Xhigh) | 79% |
CritPt Artificial AnalysisMuse Spark 1.2 (Xhigh) | 17.7% |