OpenAI has 20 models tracked on this site, from GPT-3 175B in 2020 to GPT-6.1 Sol, released Sep 29, 2026. Its highest-ranked model, GPT-6 Astra, is #2 of 39 on the Singularity Index with a score of 62.1.
Newest first. Bars show each model's Singularity Index score out of 100. Provisional scores have partial benchmark coverage and are not ranked yet.
Upgrade to GPT-6 Sol announced at DevDay one week after it launched. OpenAI reports near-Astra performance on agentic coding, computer use, and professional work at one-fifth of the token price of GPT-6 Astra.
Lower-cost GPT-6 model for coding and everyday work. Independent measurements track max effort; missing evaluations are not estimated.
Reasoning model for coding and professional work. Independent measurements track max effort.
Frontier reasoning model for computer use, coding, science, and professional work. Rolled out in phases starting September 3. Independent evaluations include SciCode and MMMU-Pro.
Most cost-efficient GPT-5.6 model for responsive, high-volume production workloads. Listed base pricing applies to prompts up to 272K tokens.
Flagship GPT-5.6 model for difficult reasoning, coding, cybersecurity, science, and agentic work. Listed base pricing applies to prompts up to 272K tokens; promotional rates run at least through November 21, 2026.
Balanced GPT-5.6 model for everyday professional work with strong capability and lower cost than Sol. Listed base pricing applies to prompts up to 272K tokens.
Frontier reasoning model that preceded GPT-5.6, with strong coding, agentic, and ARC-AGI performance.
First mainline model incorporating GPT-5.3-Codex coding capabilities. Native computer use, 1M context, surpasses human baseline on OSWorld. Codex branch ends here.
Final specialized Codex release before coding was folded into mainline GPT-5.4. Peak on LiveCodeBench and Terminal-Bench.
First model to score 100% on AIME 2025 and 80% on SWE-bench
Incremental upgrade with improved reliability and instruction following
Unified reasoning and chat, 400K context, 94.6% on AIME 2025
Full reasoning model with tool use, first to break 80% on GPQA Diamond
Reasoning at 75% lower cost than o1, made chain-of-thought economically viable
First reasoning model, uses chain-of-thought at inference time to solve hard problems
Natively multimodal with voice, 2x faster and 50% cheaper than GPT-4 Turbo
First multimodal GPT, passed the bar exam, defined the frontier for a year
NOT YET SCOREDLaunched ChatGPT, fastest consumer product to 100M users in history
NOT YET SCOREDFirst large language model to demonstrate emergent few-shot learning
NOT YET SCORED