Gemini Series Benchmark Evolution: From Gemini 1.0 to Gemini 3.5 Flash β A Complete Trend Analysis
A comprehensive longitudinal analysis of Gemini benchmark performance across the entire series (Gemini 1.0 through Gemini 3.5 Flash), tracking 20+ metrics from December 2023 to May 2026. Reveals Google's strategic evolution from native multimodality to agentic coding dominance, with Gemini 3.1 Pro achieving a 248% leap on ARC-AGI-2 and Gemini 3.5 Flash leading in multi-step tool workflows.
Executive Summary
Since the launch of Gemini 1.0 in December 2023, Google DeepMind has released a succession of models that have progressively redefined the frontier of AI capability β not just in raw reasoning, but in native multimodality, agentic tool use, and long-context understanding. This article compiles and analyzes benchmark data from the entire Gemini series (Gemini 1.0, 1.5, 2.0, 2.5, 3.0, 3.1, 3.5) across 20+ metrics, drawing exclusively from official Google DeepMind model cards, technical reports, and announcement posts.
The data reveals four distinct phases in the Gemini evolution:
- Multimodal Foundation (Gemini 1.0 β 1.5): Establishing native multimodality β Gemini 1.0 Ultra achieved 90.0% on MMLU (first model to outperform human experts), while Gemini 1.5 introduced the 1M-token context window and dominated long-context benchmarks.
- Reasoning Breakthrough (Gemini 2.0 β 2.5): The deep reasoning inflection β Gemini 2.5 Pro achieved 86.4% on GPQA Diamond, 88.0% on AIME 2025, and 21.6% on Humanity's Last Exam, establishing Google as the leader in scientific reasoning.
- Abstract Reasoning Leap (Gemini 3.0 β 3.1): The ARC-AGI-2 explosion β Gemini 3.1 Pro scored 77.1% on ARC-AGI-2 (verified by ARC Prize), a 248% improvement over Gemini 3 Pro's 31.1%, while reaching 94.3% on GPQA Diamond and 80.6% on SWE-bench Verified.
- Agentic Dominance (Gemini 3.5 Flash): The workflow king emerges β Gemini 3.5 Flash reached 83.6% on MCP Atlas (multi-step tool workflows), 76.2% on Terminal-bench 2.1, and 1656 Elo on GDPval-AA, establishing Google as the leader in agentic automation.
Key finding: The Gemini line has pursued a clear trajectory from "best multimodal model" to "best agentic reasoning engine." Gemini 3.1 Pro is the most abstract-reasoning-capable model ever released (77.1% ARC-AGI-2), while Gemini 3.5 Flash optimizes for real-world multi-step tool use. The overall story is one of Google shifting from "make the model understand everything" to "make the model do everything."
1. Release Timeline and Cadence
The cadence accelerated from 15 months (1.0 β 1.5) to just 3 months (3.1 β 3.5). This compression reflects a mature development pipeline focused on closing specific capability gaps in agentic tool use and abstract reasoning.
| Version | Release Date | Days Since Previous | Key Focus |
|---|---|---|---|
| Gemini 1.0 | Dec 2023 | β | Native multimodality, MMLU record |
| Gemini 1.5 | Mar 2024 | ~120 | 1M context window, long-context dominance |
| Gemini 2.0 | Feb 2025 | ~340 | Efficiency, tool use, function calling |
| Gemini 2.5 Pro | May 2025 | ~90 | Deep reasoning, science, math |
| Gemini 2.5 Flash | Sep 2025 | ~120 | Reasoning at scale, hybrid thinking |
| Gemini 3 Pro | Nov 2025 | ~60 | Generational leap, multimodal reasoning |
| Gemini 3.1 Pro | Feb 2026 | ~90 | ARC-AGI-2 explosion, abstract reasoning |
| Gemini 3.5 Flash | May 2026 | ~90 | Agentic workflows, multi-step tool use |
2. General Reasoning: MMLU and Beyond
2.1 MMLU β From Human-Level to Saturation
| Version | MMLU (5-shot) | Global MMLU Lite | MMLU-Pro |
|---|---|---|---|
| Gemini 1.0 Ultra | 90.0% | β | β |
| Gemini 1.5 Pro | 85.9% | β | β |
| Gemini 2.5 Pro | β | 89.2% | β |
| Gemini 3 Pro | β | β | 90.1% |
| Gemini 3.1 Pro | β | β | 91.0% |
Analysis: Gemini 1.0 Ultra's 90.0% MMLU score was historic β the first model to outperform human experts. The apparent dip in Gemini 1.5 Pro (85.9%) reflects a methodology change (different prompt format), not a regression. By Gemini 3.1 Pro, MMLU-Pro reached 91.0%, indicating near-saturation on general knowledge benchmarks. The trajectory shows Google prioritizing harder benchmarks (ARC-AGI-2, GPQA Diamond) over continuing to optimize MMLU.
2.2 GPQA Diamond β The Science Benchmark
| Version | GPQA Diamond | Ξ vs Previous |
|---|---|---|
| Gemini 1.5 Pro | 46.7% | β |
| Gemini 2.5 Pro | 86.4% | +39.7pp |
| Gemini 3.1 Pro | 94.3% | +7.9pp |
Analysis: The jump from 1.5 Pro to 2.5 Pro (+39.7pp) is the largest single-cycle gain in the entire Gemini series on any benchmark. This reflects Google's focused investment in scientific reasoning between 2024 and 2025. At 94.3%, Gemini 3.1 Pro leads all models on GPQA Diamond, surpassing GPT-5.2 (92.4%) and Claude Opus 4.6 (91.3%).
3. Mathematics: From GSM8k to AIME
3.1 AIME 2025 β Competitive Math
| Version | AIME 2025 | AIME 2024 |
|---|---|---|
| Gemini 2.5 Pro | 88.0% | 92.0% |
| Gemini 3.1 Pro | 91.0% | β |
Analysis: Gemini 2.5 Pro achieved 88.0% on AIME 2025, competitive with OpenAI o3 (88.9%) and far ahead of Claude 4 Sonnet (70.5%). Gemini 3.1 Pro pushed to 91.0%, establishing Google as the leader in competitive mathematics.
3.2 GSM8k and MATH β Foundational Math
| Version | GSM8k | MATH |
|---|---|---|
| Gemini 1.5 Pro | 88.9% | 76.9% |
| Gemini 2.5 Pro | β | β |
Analysis: Gemini 1.5 Pro's math scores were solid but not leading β Claude 3 Opus dominated both benchmarks at the time. Google's strategy shifted from optimizing foundational math (GSM8k/MATH) to tackling harder competitions (AIME), where Gemini 2.5 Pro and 3.1 Pro achieved dominance.
4. Abstract Reasoning: The ARC-AGI-2 Story
4.1 The 248% Leap
| Version | ARC-AGI-2 | Ξ vs Previous | Verification |
|---|---|---|---|
| Gemini 3 Pro | 31.1% | β | ARC Prize Verified |
| Gemini 3.1 Pro | 77.1% | +46.0pp (+148%) | ARC Prize Verified |
Analysis: This is the most dramatic single-benchmark improvement in the history of LLM development. Gemini 3.1 Pro more than doubled Gemini 3 Pro's ARC-AGI-2 score in just 3 months. At 77.1%, it leads all commercially available models, surpassing Claude Opus 4.6 (68.8%), GPT-5.2 (52.9%), and GPT-5.5 (84.6% on a different variant). The ARC Prize verification adds credibility β this is not a self-reported score.
Why this matters: ARC-AGI-2 is specifically designed to resist pattern memorization and test genuine fluid reasoning. A 46pp improvement suggests a fundamental architectural breakthrough, not just more training data.
5. Software Engineering: The Agentic Coding Narrative
5.1 SWE-bench Verified β From 64% to 81%
| Version | SWE-bench Verified | Ξ vs Previous |
|---|---|---|
| Gemini 2.5 Pro | 63.8% | β |
| Gemini 3.1 Pro | 80.6% | +16.8pp |
| Gemini 3.5 Flash | ~80.6% | ~0pp |
Analysis: The +16.8pp jump from 2.5 Pro to 3.1 Pro is the largest single-cycle gain in the Gemini series on coding benchmarks. At 80.6%, Gemini 3.1 Pro ties Claude Opus 4.6 and trails GPT-5.5 (88.7%) on this specific benchmark. However, Gemini 3.5 Flash's strength lies in harder, more agentic benchmarks (see below).
5.2 SWE-bench Pro β The Harder Frontier
| Version | SWE-bench Pro | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 49.6% | β |
| Gemini 3.1 Pro | 54.2% | +4.6pp |
| Gemini 3.5 Flash | 55.1% | +0.9pp |
Analysis: SWE-bench Pro shows more incremental improvement, suggesting diminishing returns on pure coding benchmarks. Google's strategy shifted toward agentic tool use (MCP Atlas, Terminal-bench) rather than continuing to optimize SWE-bench Pro.
5.3 Terminal-bench β Agentic Terminal Coding
| Version | Terminal-bench | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 58.0% | β |
| Gemini 3.1 Pro | 70.3% | +12.3pp |
| Gemini 3.5 Flash | 76.2% | +5.9pp |
Analysis: Terminal-bench shows consistent improvement across all three versions, with Gemini 3.5 Flash reaching 76.2% β just behind GPT-5.5 (78.2%). This benchmark better reflects real-world software engineering workflows than SWE-bench alone.
6. Agentic Tool Use: The Gemini 3.5 Flash Advantage
6.1 MCP Atlas β Multi-Step Tool Workflows
| Version | MCP Atlas | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 62.0% | β |
| Gemini 3.1 Pro | 78.2% | +16.2pp |
| Gemini 3.5 Flash | 83.6% | +5.4pp |
Analysis: MCP Atlas is where Gemini 3.5 Flash truly shines. At 83.6%, it leads all models on multi-step tool workflows, surpassing Claude Opus 4.7 (79.1%) and GPT-5.5 (75.3%). This is the benchmark that best captures real-world agentic automation β the ability to chain multiple tool calls into coherent workflows.
6.2 OSWorld β Agentic Computer Use
| Version | OSWorld-Verified | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 65.1% | β |
| Gemini 3.1 Pro | 76.2% | +11.1pp |
| Gemini 3.5 Flash | 78.4% | +2.2pp |
Analysis: Gemini 3.5 Flash reaches 78.4% on OSWorld, competitive with GPT-5.5 (78.7%) and Claude Opus 4.7 (78.0%). The improvements from 3 Pro to 3.1 Pro (+11.1pp) were dramatic, but 3.5 Flash shows diminishing returns on pure computer use.
7. Multimodal Understanding
7.1 MMMU-Pro β Multimodal Reasoning
| Version | MMMU-Pro | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 81.2% | β |
| Gemini 3.1 Pro | 80.5% | -0.7pp |
| Gemini 3.5 Flash | 83.6% | +3.1pp |
Analysis: A slight regression from 3 Pro to 3.1 Pro (-0.7pp), followed by recovery in 3.5 Flash (+3.1pp). At 83.6%, Gemini 3.5 Flash leads on MMMU-Pro, surpassing GPT-5.5 (81.2%). This reflects Google's continued investment in native multimodality as a differentiator.
7.2 CharXiv β Chart Reasoning
| Version | CharXiv | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 80.3% | β |
| Gemini 3.1 Pro | 83.3% | +3.0pp |
| Gemini 3.5 Flash | 84.2% | +0.9pp |
Analysis: Consistent improvement across all versions, with Gemini 3.5 Flash reaching 84.2% β competitive with GPT-5.5 (84.1%). Chart reasoning remains a strength of the Gemini line.
8. Long-Context Performance
8.1 MRCR v2 β Needle-in-a-Haystack
| Version | MRCR 128k (avg) | MRCR 1M (pointwise) |
|---|---|---|
| Gemini 3 Pro | 67.2% | 22.1% |
| Gemini 3.1 Pro | 84.9% | 26.3% |
| Gemini 3.5 Flash | 77.3% | 26.6% |
Analysis: Gemini 3.1 Pro dominated at 128k context (84.9%), but Gemini 3.5 Flash shows a regression (-7.6pp). At 1M context, both 3.1 Pro and 3.5 Flash score ~26%, significantly ahead of Claude Opus 4.7 (not supported). This reflects a trade-off: 3.5 Flash optimized for agentic workflows at the cost of pure long-context retrieval.
9. Humanity's Last Exam β The Ultimate Reasoning Test
| Version | HLE (full set) | Ξ vs Previous |
|---|---|---|
| Gemini 2.5 Pro | 21.6% | β |
| Gemini 3 Pro | 33.7% | +12.1pp |
| Gemini 3.1 Pro | 44.4% | +10.7pp |
| Gemini 3.5 Flash | 40.2% | -4.2pp |
Analysis: Gemini 3.1 Pro peaked at 44.4%, trailing Claude Opus 4.7 (46.9%) but leading GPT-5.5 (41.4%). The regression in 3.5 Flash (-4.2pp) suggests a strategic trade-off: optimizing for agentic tool use (MCP Atlas +5.4pp) at the cost of pure reasoning (HLE -4.2pp).
10. Expert Tasks and Real-World Workflows
10.1 GDPval-AA β Economically Valuable Knowledge Work
| Version | GDPval-AA (Elo) | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 1204 | β |
| Gemini 3.1 Pro | 1314 | +110 |
| Gemini 3.5 Flash | 1656 | +342 |
Analysis: The +342 Elo jump from 3.1 Pro to 3.5 Flash is the largest single-cycle gain in the entire Gemini series on any benchmark. At 1656 Elo, Gemini 3.5 Flash trails only Claude Opus 4.7 (1753) and GPT-5.5 (1769) on economically valuable knowledge work. This is the benchmark that best captures real-world professional tasks.
10.2 Finance Agent v2
| Version | Finance Agent v2 | Ξ vs Previous |
|---|---|---|
| Gemini 3 Pro | 42.6% | β |
| Gemini 3.1 Pro | 43.0% | +0.4pp |
| Gemini 3.5 Flash | 57.9% | +14.9pp |
Analysis: Gemini 3.5 Flash's +14.9pp jump on Finance Agent v2 is dramatic. At 57.9%, it leads all models, surpassing Claude Opus 4.7 (51.5%) and GPT-5.5 (51.8%). This suggests 3.5 Flash is particularly well-suited for financial analysis and decision-making workflows.
11. The Trade-off Matrix: What Each Version Gave Up
| Version | Gained | Traded Off |
|---|---|---|
| 1.0 β 1.5 | 1M context window, long-context dominance | Slight MMLU regression (methodology change) |
| 1.5 β 2.5 Pro | Deep reasoning (GPQA +39.7pp), math (AIME 88%) | Multimodal edge (Claude led MMMU) |
| 2.5 Pro β 3.1 Pro | ARC-AGI-2 (+148%), SWE-bench (+16.8pp) | Some multimodal (MMMU-Pro -0.7pp) |
| 3.1 Pro β 3.5 Flash | MCP Atlas (+5.4pp), GDPval-AA (+342 Elo) | HLE (-4.2pp), MRCR 128k (-7.6pp) |
Key insight: Each generation made explicit trade-offs. Gemini 3.5 Flash sacrificed some pure reasoning (HLE) and long-context retrieval (MRCR) to gain agentic tool use (MCP Atlas) and real-world workflow performance (GDPval-AA, Finance Agent). This is a strategic choice, not a regression.
12. References and Resources
Official Google DeepMind Sources
- Gemini 1.0 Announcement: blog.google/innovation-and-ai/technology/ai/google-gemini-ai/ (Dec 2023)
- Gemini 1.5 Technical Report: storage.googleapis.com/deepmind-media/gemini/gemini_v1_5_report.pdf (Feb 2024)
- Gemini 2.0 Flash Model Card: storage.googleapis.com/deepmind-media/Model-Cards/Gemini-2-0-Flash-Model-Card.pdf (Apr 2025)
- Gemini 2.5 Pro Model Card: storage.googleapis.com/deepmind-media/Model-Cards/Gemini-2-5-Pro-Model-Card.pdf (Jun 2025)
- Gemini 3 Pro Model Card: storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf (Nov 2025)
- Gemini 3.1 Pro Model Card: deepmind.google/models/model-cards/gemini-3-1-pro/ (Feb 2026)
- Gemini 3.5 Flash Benchmarks: deepmind.google/models/gemini/ (May 2026)
- ARC Prize Verification: ARC-AGI-2 scores for Gemini 3 Pro (31.1%) and Gemini 3.1 Pro (77.1%) are ARC Prize Verified
Cross-References
- Gpt Series Benchmark Evolution 4 To 55 Complete Trend Analysis 2026 05 30 β GPT series benchmark evolution
- Claude Opus Benchmark Evolution 41 To 48 Complete Trend Analysis 2026 05 29 β Claude Opus benchmark evolution
- Frontier Convergence Five Models Mimo Qwen V4 Gpt55 Opus47 2026 04 28 β Frontier model convergence analysis
13. Future Directions
The Gemini trajectory suggests three likely directions for future releases:
- Agentic Specialization: Gemini 3.5 Flash's dominance on MCP Atlas and Finance Agent suggests future versions will double down on multi-step tool workflows, potentially reaching 90%+ on MCP Atlas.
- Reasoning Recovery: The HLE regression in 3.5 Flash (-4.2pp) may be corrected in a future "Pro" variant that combines agentic tool use with pure reasoning.
- Multimodal Integration: Gemini's native multimodality remains a differentiator. Future versions may integrate video understanding and real-time audio more deeply into agentic workflows.
The Gemini line has evolved from "best multimodal model" to "best agentic reasoning engine." The next question is whether Google will continue this trajectory or split into separate specialized lines (Pro for reasoning, Flash for agentic workflows).