Upgrade to GPT-6 Sol announced at DevDay one week after it launched. OpenAI reports near-Astra performance on agentic coding, computer use, and professional work at one-fifth of the token price of GPT-6 Astra.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 58 models released within six months of it, across 4 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
ARC-AGIARC Prize Novel reasoning tasks requiring fluid intelligence | 97.1% | #4 / 39 |
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 52.9% | #8 / 95 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGPT-6.1 Sol (Max) | 52.9% |
MMMU-Pro Artificial AnalysisGPT-6.1 Sol (Max) | 86% |
Terminal-Bench 4.0 Artificial AnalysisGPT-6.1 Sol (Max) | 56.1% |
SciCode Artificial AnalysisGPT-6.1 Sol (Max) | 54.2% |
AA-LCR Artificial AnalysisGPT-6.1 Sol (Max) | 83% |
CritPt Artificial AnalysisGPT-6.1 Sol (Max) | 31.7% |