Claude 5.5 model for agentic coding and knowledge work. Independent measurements track max effort with default safeguards and fallback behavior.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
| Benchmark | Score | Rank |
|---|---|---|
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 61.4% | #1 / 91 |
ARC-AGIARC Prize Novel reasoning tasks requiring fluid intelligence | 97.9% | #2 / 36 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 61.4% |
MMMU-Pro Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 87.7% |
Terminal-Bench 4.0 Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 59.6% |
SciCode Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 66.9% |
AA-LCR Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 84.7% |
CritPt Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 31.7% |
ITBench SRE Artificial AnalysisClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Verified 2026-09-25 | 38.2% |
Terminal-Bench 4.0 Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks; Xhigh effort; standard error 2.6 points Verified 2026-09-25 | 66.4% |
FrontierCode 1.1 Main Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks Verified 2026-09-25 | 54.4% |
CursorBench 4.0 Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks Verified 2026-09-25 | 57.8% |
GDPval-AA 2.1 Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks Verified 2026-09-25 | 1846 Elo |
AutomationBench Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks; Zapier early-access evaluation; no fallback models Verified 2026-09-25 | 40% |
Humanity's Last Exam Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks; With tools Verified 2026-09-25 | 67.7% |
Terminal-Bench-Science 0.1 Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks Verified 2026-09-25 | 58.7% |
OSWorld 2.0 Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks; Partial reward Verified 2026-09-25 | 81.8% |
Chartography Provider reportAdaptive thinking, max effort unless noted; production safeguards with Opus 4.8/5 fallbacks; With tools Verified 2026-09-25 | 89% |