Compare · ModelsLive · 2 picked · head to head

Gemini 1.5 Pro (May 2024) vs DeepSeek V3

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

DeepSeek V3 wins 8 of 12 shared benchmarks. Leads in knowledge · math · coding.

Category leads
reasoning·Gemini 1.5 Pro (May 2024)knowledge·DeepSeek V3language·Gemini 1.5 Pro (May 2024)math·DeepSeek V3coding·DeepSeek V3
Hype vs Reality
Gemini 1.5 Pro (May 2024)
#165 by perf·no signal
QUIET
DeepSeek V3
#59 by perf·no signal
QUIET
Best value
Gemini 1.5 Pro (May 2024)
no price
DeepSeek V3
118.0 pts/$
$0.50/M
Vendor risk
One or more vendors flagged
Google DeepMind logo
Google DeepMind
$4.00T·Tier 1
Low risk
DeepSeek logo
DeepSeek
$3.4B·Tier 1
Higher risk
Head to head
Gemini 1.5 Pro (May 2024)DeepSeek V3
BBH
Gemini 1.5 Pro (May 2024) leads by +2.3
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
Gemini 1.5 Pro (May 2024)
85.6
DeepSeek V3
83.3
GPQA diamond
DeepSeek V3 leads by +14.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 1.5 Pro (May 2024)
27.8
DeepSeek V3
42.0
HELM · GPQA
DeepSeek V3 leads by +0.4
Gemini 1.5 Pro (May 2024)
53.4
DeepSeek V3
53.8
HELM · IFEval
Gemini 1.5 Pro (May 2024) leads by +0.5
Gemini 1.5 Pro (May 2024)
83.7
DeepSeek V3
83.2
HELM · MMLU-Pro
Gemini 1.5 Pro (May 2024) leads by +1.4
Gemini 1.5 Pro (May 2024)
73.7
DeepSeek V3
72.3
HELM · Omni-MATH
DeepSeek V3 leads by +3.9
Gemini 1.5 Pro (May 2024)
36.4
DeepSeek V3
40.3
HELM · WildBench
DeepSeek V3 leads by +1.8
Gemini 1.5 Pro (May 2024)
81.3
DeepSeek V3
83.1
MATH level 5
DeepSeek V3 leads by +24.1
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Gemini 1.5 Pro (May 2024)
40.8
DeepSeek V3
64.8
MMLU
DeepSeek V3 leads by +1.7
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemini 1.5 Pro (May 2024)
81.2
DeepSeek V3
82.9
OTIS Mock AIME 2024-2025
DeepSeek V3 leads by +9.0
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemini 1.5 Pro (May 2024)
6.7
DeepSeek V3
15.8
SimpleBench
Gemini 1.5 Pro (May 2024) leads by +9.8
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Gemini 1.5 Pro (May 2024)
12.5
DeepSeek V3
2.7
WeirdML
DeepSeek V3 leads by +13.9
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 1.5 Pro (May 2024)
22.2
DeepSeek V3
36.1
Full benchmark table
BenchmarkGemini 1.5 Pro (May 2024)DeepSeek V3
BBH
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
85.683.3
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
27.842.0
HELM · GPQA
53.453.8
HELM · IFEval
83.783.2
HELM · MMLU-Pro
73.772.3
HELM · Omni-MATH
36.440.3
HELM · WildBench
81.383.1
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
40.864.8
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
81.282.9
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
6.715.8
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
12.52.7
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
22.236.1
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemini 1.5 Pro (May 2024)
DeepSeek logoDeepSeek V3$0.20$0.80131K tokens (~66 books)$3.50