Compare · ModelsLive · 3 picked · head to head
Gemini 1.5 Pro (May 2024) vs DeepSeek V3 vs Gemini 1.5 Pro (Feb 2024)
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
DeepSeek V3 wins on 9/17 benchmarks
DeepSeek V3 wins 9 of 17 shared benchmarks. Leads in knowledge · math · coding.
Category leads
reasoning·Gemini 1.5 Pro (May 2024)knowledge·DeepSeek V3language·Gemini 1.5 Pro (May 2024)math·DeepSeek V3coding·DeepSeek V3arena·DeepSeek V3agentic·Gemini 1.5 Pro (May 2024)
Hype vs Reality
Attention vs performance
Gemini 1.5 Pro (May 2024)
#165 by perf·no signal
DeepSeek V3
#59 by perf·no signal
Gemini 1.5 Pro (Feb 2024)
#167 by perf·no signal
Best value
DeepSeek V3
Gemini 1.5 Pro (May 2024)
—
no price
DeepSeek V3
118.0 pts/$
$0.50/M
Gemini 1.5 Pro (Feb 2024)
—
no price
Vendor risk
Mixed exposure
One or more vendors flagged
Google DeepMind
$4.00T·Tier 1
DeepSeek
$3.4B·Tier 1
Google DeepMind
$4.00T·Tier 1
Head to head
17 benchmarks · 3 models
Gemini 1.5 Pro (May 2024)DeepSeek V3Gemini 1.5 Pro (Feb 2024)
BBH
Gemini 1.5 Pro (May 2024) leads by +2.3
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
Gemini 1.5 Pro (May 2024)
85.6
DeepSeek V3
83.3
Gemini 1.5 Pro (Feb 2024)
78.7
GPQA diamond
DeepSeek V3 leads by +14.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 1.5 Pro (May 2024)
27.8
DeepSeek V3
42.0
Gemini 1.5 Pro (Feb 2024)
27.8
HELM · GPQA
DeepSeek V3 leads by +0.4
Gemini 1.5 Pro (May 2024)
53.4
DeepSeek V3
53.8
Gemini 1.5 Pro (Feb 2024)
53.4
HELM · IFEval
Gemini 1.5 Pro (May 2024)
83.7
DeepSeek V3
83.2
Gemini 1.5 Pro (Feb 2024)
83.7
HELM · MMLU-Pro
Gemini 1.5 Pro (May 2024)
73.7
DeepSeek V3
72.3
Gemini 1.5 Pro (Feb 2024)
73.7
HELM · Omni-MATH
DeepSeek V3 leads by +3.9
Gemini 1.5 Pro (May 2024)
36.4
DeepSeek V3
40.3
Gemini 1.5 Pro (Feb 2024)
36.4
HELM · WildBench
DeepSeek V3 leads by +1.8
Gemini 1.5 Pro (May 2024)
81.3
DeepSeek V3
83.1
Gemini 1.5 Pro (Feb 2024)
81.3
MATH level 5
DeepSeek V3 leads by +24.1
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Gemini 1.5 Pro (May 2024)
40.8
DeepSeek V3
64.8
Gemini 1.5 Pro (Feb 2024)
40.8
MMLU
DeepSeek V3 leads by +1.7
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemini 1.5 Pro (May 2024)
81.2
DeepSeek V3
82.9
Gemini 1.5 Pro (Feb 2024)
76.9
OTIS Mock AIME 2024-2025
DeepSeek V3 leads by +9.0
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemini 1.5 Pro (May 2024)
6.7
DeepSeek V3
15.8
Gemini 1.5 Pro (Feb 2024)
6.7
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Gemini 1.5 Pro (May 2024)
12.5
DeepSeek V3
2.7
Gemini 1.5 Pro (Feb 2024)
12.5
WeirdML
DeepSeek V3 leads by +13.9
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 1.5 Pro (May 2024)
22.2
DeepSeek V3
36.1
Gemini 1.5 Pro (Feb 2024)
22.2
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Gemini 1.5 Pro (May 2024)
0.8
Gemini 1.5 Pro (Feb 2024)
0.8
Chatbot Arena Elo · Overall
DeepSeek V3 leads by +36.0
DeepSeek V3
1358.5
Gemini 1.5 Pro (Feb 2024)
1322.5
Balrog
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Gemini 1.5 Pro (May 2024)
21.0
Gemini 1.5 Pro (Feb 2024)
21.0
CadEval
CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge.
Gemini 1.5 Pro (May 2024)
34.0
Gemini 1.5 Pro (Feb 2024)
34.0
The Agent Company
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
Gemini 1.5 Pro (May 2024)
3.4
Gemini 1.5 Pro (Feb 2024)
3.4
Full benchmark table
| Benchmark | Gemini 1.5 Pro (May 2024) | DeepSeek V3 | Gemini 1.5 Pro (Feb 2024) |
|---|---|---|---|
BBH BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans. | 85.6 | 83.3 | 78.7 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 27.8 | 42.0 | 27.8 |
HELM · GPQA | 53.4 | 53.8 | 53.4 |
HELM · IFEval | 83.7 | 83.2 | 83.7 |
HELM · MMLU-Pro | 73.7 | 72.3 | 73.7 |
HELM · Omni-MATH | 36.4 | 40.3 | 36.4 |
HELM · WildBench | 81.3 | 83.1 | 81.3 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 40.8 | 64.8 | 40.8 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 81.2 | 82.9 | 76.9 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 6.7 | 15.8 | 6.7 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 12.5 | 2.7 | 12.5 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 22.2 | 36.1 | 22.2 |
ARC-AGI-2 ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. | 0.8 | — | 0.8 |
Chatbot Arena Elo · Overall | — | 1358.5 | 1322.5 |
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 21.0 | — | 21.0 |
CadEval CadEval · evaluates the ability to generate and reason about Computer-Aided Design code, testing spatial reasoning and engineering knowledge. | 34.0 | — | 34.0 |
The Agent Company The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows. | 3.4 | — | 3.4 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $0.20 | $0.80 | 131K tokens (~66 books) | $3.50 | |
| — | — | — | — |