Open-weight coding model optimized for long-horizon software engineering and agentic tool use.
| Benchmark | Score | Rank |
|---|---|---|
TerminalArtificial Analysis Agentic terminal coding tasks requiring multi-step execution | 44.7% | #23 / 56 |
GPQAArtificial Analysis PhD-level science questions even experts struggle with | 89.6% | #27 / 81 |
hleArtificial Analysis | 32.8% | #29 / 69 |
LiveCodeBenchvals.ai Contamination-free competitive programming (filtered by cutoff date) | 82.1% | #33 / 54 |