Most cost-efficient GPT-5.6 model for responsive, high-volume production workloads.
| Benchmark | Score | Rank |
|---|---|---|
Terminal Agentic terminal coding tasks requiring multi-step execution | 84.7% | #4 / 56 |
AA Coding Independent composite of implementation, terminal use, and real-codebase agent performance | 74.6% | #7 / 11 |
AA Intelligence Independent composite of agentic work, coding, scientific reasoning, and general capability | 51.2% | #8 / 11 |
ARC-AGI Novel reasoning tasks requiring fluid intelligence | 60.2% | #12 / 28 |
OSWorld Computer use in real desktop environments | 45.6% | #13 / 13 |
MMMU College-level multimodal reasoning across 30+ disciplines | 85% | #16 / 53 |
GPQA PhD-level science questions even experts struggle with | 91.1% | #19 / 81 |
hle | 37.2% | #20 / 69 |
MMLU-Pro Harder 10-option successor to MMLU; more reasoning-focused | 86% | #33 / 52 |