Compare · ModelsLive · 2 picked · head to head
DeepSeek R1 Distill Qwen 32B vs Grok 4
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Grok 4 wins on 4/4 benchmarks
Grok 4 wins 4 of 4 shared benchmarks. Leads in knowledge · math.
Category leads
knowledge·Grok 4math·Grok 4
Hype vs Reality
Attention vs performance
DeepSeek R1 Distill Qwen 32B
#163 by perf·no signal
Grok 4
#87 by perf·#15 by attention
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
xAI
$250.0B·Tier 1
Head to head
4 benchmarks · 2 models
DeepSeek R1 Distill Qwen 32BGrok 4
Balrog
Grok 4 leads by +24.1
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
DeepSeek R1 Distill Qwen 32B
19.5
Grok 4
43.6
Chess Puzzles
Grok 4 leads by +24.2
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek R1 Distill Qwen 32B
0.0
Grok 4
24.2
GPQA diamond
Grok 4 leads by +30.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek R1 Distill Qwen 32B
52.2
Grok 4
82.7
OTIS Mock AIME 2024-2025
Grok 4 leads by +28.5
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek R1 Distill Qwen 32B
55.5
Grok 4
84.0
Full benchmark table
| Benchmark | DeepSeek R1 Distill Qwen 32B | Grok 4 |
|---|---|---|
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 19.5 | 43.6 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 0.0 | 24.2 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 52.2 | 82.7 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 55.5 | 84.0 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| — | — | — | — |
People also compared