Flagship GPT-5.6 model for difficult reasoning, coding, cybersecurity, science, and agentic work. Listed base pricing applies to prompts up to 272K tokens; promotional rates run at least through November 21, 2026.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 71 models released within six months of it, across 7 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
Terminal Agentic terminal coding tasks requiring multi-step execution | 88.8% | #1 / 58 |
AA Coding (Historical) Archived release-era coding composite, excluded from current rankings because comparable current-series coverage is unavailable | 80% | #1 / 22 |
MMMU College-level multimodal reasoning across 30+ disciplines | 88.8% | #6 / 51 |
AA Intelligence (Historical) Archived release-era composite, not comparable to the current versioned Intelligence Index used for rankings | 58.9% | #6 / 22 |
ARC-AGI Novel reasoning tasks requiring fluid intelligence | 93.8% | #6 / 39 |
GPQA PhD-level science questions even experts struggle with | 94.1% | #7 / 95 |
OSWorld Computer use in real desktop environments | 62.6% | #7 / 13 |
HLE Challenging multidisciplinary questions evaluated by Artificial Analysis | 49.5% | #9 / 97 |
MMLU-Pro Harder 10-option successor to MMLU; more reasoning-focused | 89.1% | #15 / 61 |
LiveCodeBench Contamination-free competitive programming (filtered by cutoff date) | 82.6% | #35 / 62 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGPT-5.6 Sol (Max) | 49.5% |
MMMU-Pro Artificial AnalysisGPT-5.6 Sol (Max) | 83.4% |
Terminal-Bench 2.1 Artificial AnalysisGPT-5.6 Sol (Max) | 88% |
Terminal-Bench 4.0 Artificial AnalysisGPT-5.6 Sol (Max) | 39.9% |
SciCode Artificial AnalysisGPT-5.6 Sol (Max) | 57.1% |
AA-LCR Artificial AnalysisGPT-5.6 Sol (Max) | 84% |
CritPt Artificial AnalysisGPT-5.6 Sol (Max) | 32.3% |
ITBench SRE Artificial AnalysisGPT-5.6 Sol (Max) | 56.2% |
IFBench Artificial AnalysisGPT-5.6 Sol (Max) | 72.7% |