Open-weight native-vision flagship with 2.8T parameters and a one-million-token context window.
| Benchmark | Score | Rank |
|---|---|---|
AA Coding Independent composite of implementation, terminal use, and real-codebase agent performance | 76.2% | #8 / 22 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 93.5% | #10 / 95 |
AA Intelligence Independent composite of agentic work, coding, scientific reasoning, and general capability | 57.1% | #10 / 22 |
MMMUvals.ai College-level multimodal reasoning across 30+ disciplines | 88.2% | #11 / 51 |
aaAgentic | 50.1% | #11 / 16 |
hleArtificial Analysis | 46.9% | #11 / 85 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 87.2% | #16 / 62 |
MMLU-Provals.ai Harder 10-option successor to MMLU; more reasoning-focused | 88% | #22 / 61 |