Compare · ModelsLive · 3 picked · head to head
Claude 3.5 Sonnet vs DeepSeek V3 vs GPT-4 (older v0314)
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
DeepSeek V3 wins on 9/19 benchmarks
DeepSeek V3 wins 9 of 19 shared benchmarks. Leads in math · reasoning.
Category leads
arena·Claude 3.5 Sonnetknowledge·Claude 3.5 Sonnetmath·DeepSeek V3coding·Claude 3.5 Sonnetgeneral·Claude 3.5 Sonnetlanguage·Claude 3.5 Sonnetreasoning·DeepSeek V3
Hype vs Reality
Attention vs performance
Claude 3.5 Sonnet
#190 by perf·no signal
DeepSeek V3
#70 by perf·no signal
GPT-4 (older v0314)
#81 by perf·no signal
Best value
DeepSeek V3
71.4x better value than GPT-4 (older v0314)
Claude 3.5 Sonnet
n/a
no price
DeepSeek V3
87.2 pts/$
$0.64/M
GPT-4 (older v0314)
1.2 pts/$
$45.00/M
Vendor risk
Mixed exposure
One or more vendors flagged
Anthropic
$965.0B·Tier 1
DeepSeek
$3.4B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
19 benchmarks · 3 models
Claude 3.5 SonnetDeepSeek V3GPT-4 (older v0314)
Chatbot Arena Elo · Overall
Claude 3.5 Sonnet leads by +15.5
Claude 3.5 Sonnet
1374.0
DeepSeek V3
1358.4
GPT-4 (older v0314)
1285.8
GPQA diamond
DeepSeek V3 leads by +3.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude 3.5 Sonnet
38.7
DeepSeek V3
42.0
GPT-4 (older v0314)
14.3
MMLU
DeepSeek V3 leads by +0.9
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Claude 3.5 Sonnet
82.0
DeepSeek V3
82.9
GPT-4 (older v0314)
81.9
OTIS Mock AIME 2024-2025
DeepSeek V3 leads by +9.3
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude 3.5 Sonnet
6.4
DeepSeek V3
15.8
GPT-4 (older v0314)
0.5
Aider · Code Editing
Claude 3.5 Sonnet leads by +18.0
Claude 3.5 Sonnet
84.2
GPT-4 (older v0314)
66.2
Aider polyglot
Claude 3.5 Sonnet leads by +3.2
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
Claude 3.5 Sonnet
51.6
DeepSeek V3
48.4
Dtbench
Claude 3.5 Sonnet leads by +5.0
Claude 3.5 Sonnet
46.3
DeepSeek V3
41.3
FrontierMath-2025-02-28-Private
DeepSeek V3 leads by +1.2
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Claude 3.5 Sonnet
1.8
DeepSeek V3
3.0
HELM · GPQA
Claude 3.5 Sonnet leads by +2.7
Claude 3.5 Sonnet
56.5
DeepSeek V3
53.8
HELM · IFEval
Claude 3.5 Sonnet leads by +2.4
Claude 3.5 Sonnet
85.6
DeepSeek V3
83.2
HELM · MMLU-Pro
Claude 3.5 Sonnet leads by +5.4
Claude 3.5 Sonnet
77.7
DeepSeek V3
72.3
HELM · Omni-MATH
DeepSeek V3 leads by +12.7
Claude 3.5 Sonnet
27.6
DeepSeek V3
40.3
HELM · WildBench
DeepSeek V3 leads by +3.9
Claude 3.5 Sonnet
79.2
DeepSeek V3
83.1
Lech Mazur Writing
Claude 3.5 Sonnet leads by +3.3
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Claude 3.5 Sonnet
80.3
DeepSeek V3
77.0
MATH level 5
DeepSeek V3 leads by +13.2
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Claude 3.5 Sonnet
51.7
DeepSeek V3
64.8
Metr Time Horizons
DeepSeek V3 leads by +7.2
Claude 3.5 Sonnet
40.1
DeepSeek V3
47.4
SimpleBench
Claude 3.5 Sonnet leads by +10.3
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Claude 3.5 Sonnet
13.0
DeepSeek V3
2.7
WeirdML
DeepSeek V3 leads by +5.1
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude 3.5 Sonnet
31.0
DeepSeek V3
36.1
Winogrande
GPT-4 (older v0314) leads by +4.6
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
DeepSeek V3
70.4
GPT-4 (older v0314)
75.0
Full benchmark table
| Benchmark | Claude 3.5 Sonnet | DeepSeek V3 | GPT-4 (older v0314) |
|---|---|---|---|
Chatbot Arena Elo · Overall | 1374.0 | 1358.4 | 1285.8 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 38.7 | 42.0 | 14.3 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 82.0 | 82.9 | 81.9 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 6.4 | 15.8 | 0.5 |
Aider · Code Editing | 84.2 | — | 66.2 |
Aider polyglot Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework. | 51.6 | 48.4 | — |
Dtbench | 46.3 | 41.3 | — |
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 1.8 | 3.0 | — |
HELM · GPQA | 56.5 | 53.8 | — |
HELM · IFEval | 85.6 | 83.2 | — |
HELM · MMLU-Pro | 77.7 | 72.3 | — |
HELM · Omni-MATH | 27.6 | 40.3 | — |
HELM · WildBench | 79.2 | 83.1 | — |
Lech Mazur Writing Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication. | 80.3 | 77.0 | — |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 51.7 | 64.8 | — |
Metr Time Horizons | 40.1 | 47.4 | — |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 13.0 | 2.7 | — |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 31.0 | 36.1 | — |
Winogrande WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs. | — | 70.4 | 75.0 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $0.26 | $1.03 | 164K tokens (~82 books) | $4.50 | |
| $30.00 | $60.00 | 8K tokens (~4 books) | $375.00 |
People also compared