Most cost-efficient GPT-5.6 model for responsive, high-volume production workloads.
| Benchmark | Score | Rank |
|---|---|---|
Terminal Agentic terminal coding tasks requiring multi-step execution | 84.7% | #4 / 56 |
AA Coding Independent composite of implementation, terminal use, and real-codebase agent performance | 74.6% | #10 / 18 |
OSWorld Computer use in real desktop environments | 45.6% | #13 / 13 |
ARC-AGI Novel reasoning tasks requiring fluid intelligence | 60.2% | #14 / 30 |
AA Intelligence Independent composite of agentic work, coding, scientific reasoning, and general capability | 51.2% | #15 / 18 |
MMMU College-level multimodal reasoning across 30+ disciplines | 85% | #19 / 57 |
GPQA PhD-level science questions even experts struggle with | 91.1% | #24 / 88 |
hle | 37.2% | #26 / 76 |
MMLU-Pro Harder 10-option successor to MMLU; more reasoning-focused | 86% | #38 / 58 |