SpaceXAI frontier model focused on coding, agentic tasks, and professional knowledge work.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 71 models released within six months of it, across 4 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 89.2% | #14 / 61 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 87.4% | #14 / 62 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 93.1% | #15 / 95 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 42.7% | #28 / 97 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 61.8% | #48 / 51 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGrok 4.5 (High) | 42.7% |
MMMU-Pro Artificial AnalysisGrok 4.5 (High) | 80.4% |
Terminal-Bench 2.1 Artificial AnalysisGrok 4.5 (High) | 81.6% |
Terminal-Bench 4.0 Artificial AnalysisGrok 4.5 (High) | 10.6% |
SciCode Artificial AnalysisGrok 4.5 (High) | 55% |
AA-LCR Artificial AnalysisGrok 4.5 (High) | 79.3% |
CritPt Artificial AnalysisGrok 4.5 (High) | 15.4% |