First Gemini 4 model, aimed at sustained reasoning, software engineering, and cyber defense. Rolling out first to trusted cyber defenders, with a one-million-token output limit and introductory pricing of $2 input and $10 output per million tokens.
SI v3 uses weighted percentages from one Artificial Analysis configuration. Ranking requires all five benchmarks. Partial scores average only available weights. Scores are not comparable to SI v2 or a probability of reaching singularity.
Compared with 56 models released within six months of it, across 3 benchmarks where scores spread out.
| Benchmark | Score | Rank |
|---|---|---|
HLEArtificial Analysis Challenging multidisciplinary questions evaluated by Artificial Analysis | 57.1% | #3 / 95 |
| Benchmark / source | Score |
|---|---|
HLE Artificial AnalysisGemini 4 Argon (High) | 57.1% |
Terminal-Bench 4.0 Artificial AnalysisGemini 4 Argon (High) | 57.1% |
SciCode Artificial AnalysisGemini 4 Argon (High) | 61.8% |
AA-LCR Artificial AnalysisGemini 4 Argon (High) | 79.7% |
CritPt Artificial AnalysisGemini 4 Argon (High) | 27.1% |