Compare · ModelsLive · 2 picked · head to head
DeepSeek R1 Distill Qwen 32B vs Gemini 2.5 Pro
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Gemini 2.5 Pro wins on 4/4 benchmarks
Gemini 2.5 Pro wins 4 of 4 shared benchmarks. Leads in knowledge · math.
Category leads
knowledge·Gemini 2.5 Promath·Gemini 2.5 Pro
Hype vs Reality
Attention vs performance
DeepSeek R1 Distill Qwen 32B
#163 by perf·no signal
Gemini 2.5 Pro
#116 by perf·no signal
Best value
Gemini 2.5 Pro
DeepSeek R1 Distill Qwen 32B
n/a
no price
Gemini 2.5 Pro
9.0 pts/$
$5.63/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
Google DeepMind
$4.20T·Tier 1
Head to head
4 benchmarks · 2 models
DeepSeek R1 Distill Qwen 32BGemini 2.5 Pro
Balrog
Gemini 2.5 Pro leads by +23.8
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
DeepSeek R1 Distill Qwen 32B
19.5
Gemini 2.5 Pro
43.3
Chess Puzzles
Gemini 2.5 Pro leads by +15.8
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek R1 Distill Qwen 32B
0.0
Gemini 2.5 Pro
15.8
GPQA diamond
Gemini 2.5 Pro leads by +28.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek R1 Distill Qwen 32B
52.2
Gemini 2.5 Pro
80.4
OTIS Mock AIME 2024-2025
Gemini 2.5 Pro leads by +29.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek R1 Distill Qwen 32B
55.5
Gemini 2.5 Pro
84.7
Full benchmark table
| Benchmark | DeepSeek R1 Distill Qwen 32B | Gemini 2.5 Pro |
|---|---|---|
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 19.5 | 43.3 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 0.0 | 15.8 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 52.2 | 80.4 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 55.5 | 84.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $1.25 | $10.00 | 1.0M tokens (~524 books) | $34.38 |