Google has 18 models tracked on this site, from Chinchilla 70B in 2022 to Gemini 4 Argon, released Sep 30, 2026. Its highest-ranked model, Gemini 3.8 Flash, is #10 of 39 on the Singularity Index with a score of 50.1.
Newest first. Bars show each model's Singularity Index score out of 100. Provisional scores have partial benchmark coverage and are not ranked yet.
First Gemini 4 model, aimed at sustained reasoning, software engineering, and cyber defense. Rolling out first to trusted cyber defenders, with a one-million-token output limit and introductory pricing of $2 input and $10 output per million tokens.
Reasoning and coding workhorse for long-running agent tasks. Introductory pricing of $0.75 input and $3.75 output per million tokens lasts through December 31, 2026.
Flash-tier successor to 3.6 posting GPQA 94.5 and ARC-AGI-2 85.3 — frontier-level reasoning at the cost-efficient tier.
High-throughput Flash model for scaled agentic search and document processing, with 350 output tokens per second reported by Artificial Analysis.
Successor to 3.5 Flash, pushing Flash-tier reasoning to GPQA 92.8 while holding the cost-efficient position.
Flash-tier model with near-frontier reasoning: GPQA 92.2, ARC-AGI-1 96%. AA Intelligence Index 55.3 — top of the cost-efficient tier.
Apache 2.0 dense 31B from the Gemma 4 family. Native vision + audio, 140+ languages. Best open-weight from Google Q2 2026.
Highest GPQA Diamond score ever at 94.3%, doubled ARC-AGI-2 to 77.1%
Flash-tier model outperforming previous-gen Pro on most benchmarks
First model to break 1500 Arena Elo, 100% on AIME with code execution
First Flash model with thinking, near-Pro performance at one-eighth the price
Built-in thinking mode, native audio and video understanding
Near-frontier performance at flash pricing, native tool use and code execution
First 1M token context window, processed entire codebases in one pass
Google's first natively multimodal model, launched the Gemini brand
NOT YET SCOREDPowered Bard and Google Workspace AI, strong multilingual performance
NOT YET SCOREDLargest dense model at launch, first to show chain-of-thought reasoning at scale
NOT YET SCOREDProved most LLMs were undertrained, reshaped scaling strategy industry-wide
NOT YET SCORED