Frontier reasoning model for computer use, coding, science, and professional work. Rolled out in phases starting September 3; independent coding and standard MMMU coverage remain pending in the tracked feeds.
| Benchmark | Score | Rank |
|---|---|---|
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 96.1% | #1 / 95 |
ARC-AGIARC Prize Novel reasoning tasks requiring fluid intelligence | 97.9% | #2 / 33 |
hleArtificial Analysis | 54.7% | #4 / 85 |