Frontier Anthropic model with strong coding, reasoning, and professional-work performance.
| Benchmark | Score | Rank |
|---|---|---|
hleArtificial Analysis | 45.7% | #5 / 79 |
Terminal Agentic terminal coding tasks requiring multi-step execution | 78.9% | #7 / 56 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 89.6% | #7 / 58 |
OSWorld Computer use in real desktop environments | 54.8% | #9 / 13 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 87.8% | #9 / 59 |
AA Coding Independent composite of implementation, terminal use, and real-codebase agent performance | 72.5% | #12 / 21 |
AA Intelligence Independent composite of agentic work, coding, scientific reasoning, and general capability | 55.7% | #13 / 21 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 86.6% | #14 / 59 |
GPQA PhD-level science questions even experts struggle with | 92% | #21 / 91 |