Compare · ModelsLive · 2 picked · head to head
Gemini 2.0 Flash (Dec 2024) vs Gemini 2.0 Flash
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Gemini 2.0 Flash (Dec 2024) wins on 19/19 benchmarks
Gemini 2.0 Flash (Dec 2024) wins 19 of 19 shared benchmarks. Leads in coding · reasoning · knowledge.
Category leads
coding·Gemini 2.0 Flash (Dec 2024)reasoning·Gemini 2.0 Flash (Dec 2024)knowledge·Gemini 2.0 Flash (Dec 2024)math·Gemini 2.0 Flash (Dec 2024)language·Gemini 2.0 Flash (Dec 2024)agentic·Gemini 2.0 Flash (Dec 2024)
Hype vs Reality
Attention vs performance
Gemini 2.0 Flash (Dec 2024)
#130 by perf·no signal
Gemini 2.0 Flash
#129 by perf·no signal
Best value
Gemini 2.0 Flash
Gemini 2.0 Flash (Dec 2024)
—
no price
Gemini 2.0 Flash
192.0 pts/$
$0.25/M
Vendor risk
Who is behind the model
Google DeepMind
$4.00T·Tier 1
Google DeepMind
$4.00T·Tier 1
Head to head
19 benchmarks · 2 models
Gemini 2.0 Flash (Dec 2024)Gemini 2.0 Flash
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
Gemini 2.0 Flash (Dec 2024)
38.2
Gemini 2.0 Flash
38.2
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Gemini 2.0 Flash (Dec 2024)
1.3
Gemini 2.0 Flash
1.3
CadEval
CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge.
Gemini 2.0 Flash (Dec 2024)
30.0
Gemini 2.0 Flash
30.0
Fiction.LiveBench
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
Gemini 2.0 Flash (Dec 2024)
61.1
Gemini 2.0 Flash
61.1
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Gemini 2.0 Flash (Dec 2024)
1.7
Gemini 2.0 Flash
1.7
GeoBench
GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding.
Gemini 2.0 Flash (Dec 2024)
77.0
Gemini 2.0 Flash
77.0
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 2.0 Flash (Dec 2024)
52.2
Gemini 2.0 Flash
52.2
HELM · GPQA
Gemini 2.0 Flash (Dec 2024)
55.6
Gemini 2.0 Flash
55.6
HELM · IFEval
Gemini 2.0 Flash (Dec 2024)
84.1
Gemini 2.0 Flash
84.1
HELM · MMLU-Pro
Gemini 2.0 Flash (Dec 2024)
73.7
Gemini 2.0 Flash
73.7
HELM · Omni-MATH
Gemini 2.0 Flash (Dec 2024)
45.9
Gemini 2.0 Flash
45.9
HELM · WildBench
Gemini 2.0 Flash (Dec 2024)
80.0
Gemini 2.0 Flash
80.0
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Gemini 2.0 Flash (Dec 2024)
71.5
Gemini 2.0 Flash
71.5
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Gemini 2.0 Flash (Dec 2024)
82.2
Gemini 2.0 Flash
82.2
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemini 2.0 Flash (Dec 2024)
72.9
Gemini 2.0 Flash
72.9
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemini 2.0 Flash (Dec 2024)
31.0
Gemini 2.0 Flash
31.0
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Gemini 2.0 Flash (Dec 2024)
17.3
Gemini 2.0 Flash
17.3
The Agent Company
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
Gemini 2.0 Flash (Dec 2024)
11.4
Gemini 2.0 Flash
11.4
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 2.0 Flash (Dec 2024)
25.8
Gemini 2.0 Flash
25.8
Full benchmark table
| Benchmark | Gemini 2.0 Flash (Dec 2024) | Gemini 2.0 Flash |
|---|---|---|
Aider polyglot Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework. | 38.2 | 38.2 |
ARC-AGI-2 ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. | 1.3 | 1.3 |
CadEval CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge. | 30.0 | 30.0 |
Fiction.LiveBench Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination. | 61.1 | 61.1 |
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 1.7 | 1.7 |
GeoBench GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding. | 77.0 | 77.0 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 52.2 | 52.2 |
HELM · GPQA | 55.6 | 55.6 |
HELM · IFEval | 84.1 | 84.1 |
HELM · MMLU-Pro | 73.7 | 73.7 |
HELM · Omni-MATH | 45.9 | 45.9 |
HELM · WildBench | 80.0 | 80.0 |
Lech Mazur Writing Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication. | 71.5 | 71.5 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 82.2 | 82.2 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 72.9 | 72.9 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 31.0 | 31.0 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 17.3 | 17.3 |
The Agent Company The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows. | 11.4 | 11.4 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 25.8 | 25.8 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $0.10 | $0.40 | 1.0M tokens (~500 books) | $1.75 |
People also compared