Compare · ModelsLive · 2 picked · head to head

Gemini 2.0 Flash vs Gemini 2.0 Flash (Dec 2024)

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Gemini 2.0 Flash wins 19 of 19 shared benchmarks. Leads in coding · reasoning · knowledge.

Category leads
coding·Gemini 2.0 Flashreasoning·Gemini 2.0 Flashknowledge·Gemini 2.0 Flashmath·Gemini 2.0 Flashlanguage·Gemini 2.0 Flashagentic·Gemini 2.0 Flash
Hype vs Reality
Gemini 2.0 Flash
#129 by perf·no signal
QUIET
Gemini 2.0 Flash (Dec 2024)
#130 by perf·no signal
QUIET
Best value
Gemini 2.0 Flash
192.0 pts/$
$0.25/M
Gemini 2.0 Flash (Dec 2024)
no price
Vendor risk
Google DeepMind logo
Google DeepMind
$4.00T·Tier 1
Low risk
Google DeepMind logo
Google DeepMind
$4.00T·Tier 1
Low risk
Head to head
Gemini 2.0 FlashGemini 2.0 Flash (Dec 2024)
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
Gemini 2.0 Flash
38.2
Gemini 2.0 Flash (Dec 2024)
38.2
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Gemini 2.0 Flash
1.3
Gemini 2.0 Flash (Dec 2024)
1.3
CadEval
CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge.
Gemini 2.0 Flash
30.0
Gemini 2.0 Flash (Dec 2024)
30.0
Fiction.LiveBench
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
Gemini 2.0 Flash
61.1
Gemini 2.0 Flash (Dec 2024)
61.1
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Gemini 2.0 Flash
1.7
Gemini 2.0 Flash (Dec 2024)
1.7
GeoBench
GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding.
Gemini 2.0 Flash
77.0
Gemini 2.0 Flash (Dec 2024)
77.0
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 2.0 Flash
52.2
Gemini 2.0 Flash (Dec 2024)
52.2
HELM · GPQA
Gemini 2.0 Flash
55.6
Gemini 2.0 Flash (Dec 2024)
55.6
HELM · IFEval
Gemini 2.0 Flash
84.1
Gemini 2.0 Flash (Dec 2024)
84.1
HELM · MMLU-Pro
Gemini 2.0 Flash
73.7
Gemini 2.0 Flash (Dec 2024)
73.7
HELM · Omni-MATH
Gemini 2.0 Flash
45.9
Gemini 2.0 Flash (Dec 2024)
45.9
HELM · WildBench
Gemini 2.0 Flash
80.0
Gemini 2.0 Flash (Dec 2024)
80.0
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Gemini 2.0 Flash
71.5
Gemini 2.0 Flash (Dec 2024)
71.5
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Gemini 2.0 Flash
82.2
Gemini 2.0 Flash (Dec 2024)
82.2
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemini 2.0 Flash
72.9
Gemini 2.0 Flash (Dec 2024)
72.9
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemini 2.0 Flash
31.0
Gemini 2.0 Flash (Dec 2024)
31.0
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Gemini 2.0 Flash
17.3
Gemini 2.0 Flash (Dec 2024)
17.3
The Agent Company
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
Gemini 2.0 Flash
11.4
Gemini 2.0 Flash (Dec 2024)
11.4
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 2.0 Flash
25.8
Gemini 2.0 Flash (Dec 2024)
25.8
Full benchmark table
BenchmarkGemini 2.0 FlashGemini 2.0 Flash (Dec 2024)
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
38.238.2
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
1.31.3
CadEval
CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge.
30.030.0
Fiction.LiveBench
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
61.161.1
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
1.71.7
GeoBench
GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding.
77.077.0
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
52.252.2
HELM · GPQA
55.655.6
HELM · IFEval
84.184.1
HELM · MMLU-Pro
73.773.7
HELM · Omni-MATH
45.945.9
HELM · WildBench
80.080.0
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
71.571.5
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
82.282.2
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
72.972.9
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
31.031.0
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
17.317.3
The Agent Company
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
11.411.4
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
25.825.8
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemini 2.0 Flash$0.10$0.401.0M tokens (~500 books)$1.75
Google DeepMind logoGemini 2.0 Flash (Dec 2024)