Compare · ModelsLive · 2 picked · head to head

Claude 3.5 Sonnet vs Claude Opus 4.5

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Claude Opus 4.5 wins 16 of 16 shared benchmarks. Leads in arena · knowledge · coding.

Category leads
arena·Claude Opus 4.5knowledge·Claude Opus 4.5coding·Claude Opus 4.5general·Claude Opus 4.5math·Claude Opus 4.5safety·Claude Opus 4.5reasoning·Claude Opus 4.5
Hype vs Reality
Claude 3.5 Sonnet
#190 by perf·no signal
QUIET
Claude Opus 4.5
#179 by perf·#7 by attention
OVERHYPED
Best value
Claude 3.5 Sonnet
n/a
no price
Claude Opus 4.5
2.8 pts/$
$15.00/M
Vendor risk
Anthropic logo
Anthropic
$965.0B·Tier 1
Medium risk
Anthropic logo
Anthropic
$965.0B·Tier 1
Medium risk
Head to head
Claude 3.5 SonnetClaude Opus 4.5
Chatbot Arena Elo · Overall
Claude Opus 4.5 leads by +95.8
Claude 3.5 Sonnet
1374.0
Claude Opus 4.5
1469.8
Balrog
Claude Opus 4.5 leads by +10.9
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Claude 3.5 Sonnet
32.6
Claude Opus 4.5
43.5
Cybench
Claude Opus 4.5 leads by +64.5
Cybench · evaluates AI on real Capture-The-Flag cybersecurity challenges, testing vulnerability analysis, exploitation, and security reasoning.
Claude 3.5 Sonnet
17.5
Claude Opus 4.5
82.0
Dtbench
Claude Opus 4.5 leads by +36.8
Claude 3.5 Sonnet
46.3
Claude Opus 4.5
83.1
FrontierMath-2025-02-28-Private
Claude Opus 4.5 leads by +18.9
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Claude 3.5 Sonnet
1.8
Claude Opus 4.5
20.7
FrontierMath-Tier-4-2025-07-01-Private
Claude Opus 4.5 leads by +4.2
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
Claude 3.5 Sonnet
0.0
Claude Opus 4.5
4.2
GeoBench
Claude Opus 4.5 leads by +13.0
GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding.
Claude 3.5 Sonnet
62.0
Claude Opus 4.5
75.0
GPQA diamond
Claude Opus 4.5 leads by +42.7
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude 3.5 Sonnet
38.7
Claude Opus 4.5
81.4
GSO-Bench
Claude Opus 4.5 leads by +21.9
GSO-Bench · evaluates AI models on real-world open-source software engineering tasks, testing the ability to understand and resolve actual GitHub issues.
Claude 3.5 Sonnet
4.6
Claude Opus 4.5
26.5
HLE
Claude Opus 4.5 leads by +21.4
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
Claude 3.5 Sonnet
0.0
Claude Opus 4.5
21.4
Metr Time Horizons
Claude Opus 4.5 leads by +34.8
Claude 3.5 Sonnet
40.1
Claude Opus 4.5
75.0
OTIS Mock AIME 2024-2025
Claude Opus 4.5 leads by +79.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude 3.5 Sonnet
6.4
Claude Opus 4.5
86.1
Fortress
Claude Opus 4.5 leads by +0.6
Claude 3.5 Sonnet
13.0
Claude Opus 4.5
13.6
SimpleBench
Claude Opus 4.5 leads by +41.4
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Claude 3.5 Sonnet
13.0
Claude Opus 4.5
54.4
VPCT
Claude Opus 4.5 leads by +10.0
VPCT (Visual Pattern Completion Test) · tests visual reasoning and pattern recognition by having models complete visual sequences and transformations.
Claude 3.5 Sonnet
0.0
Claude Opus 4.5
10.0
WeirdML
Claude Opus 4.5 leads by +32.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude 3.5 Sonnet
31.0
Claude Opus 4.5
63.7
Full benchmark table
BenchmarkClaude 3.5 SonnetClaude Opus 4.5
Chatbot Arena Elo · Overall
1374.01469.8
Balrog
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
32.643.5
Cybench
Cybench · evaluates AI on real Capture-The-Flag cybersecurity challenges, testing vulnerability analysis, exploitation, and security reasoning.
17.582.0
Dtbench
46.383.1
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
1.820.7
FrontierMath-Tier-4-2025-07-01-Private
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
0.04.2
GeoBench
GeoBench · tests geographic knowledge and spatial reasoning across countries, landmarks, coordinates, and geopolitical understanding.
62.075.0
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
38.781.4
GSO-Bench
GSO-Bench · evaluates AI models on real-world open-source software engineering tasks, testing the ability to understand and resolve actual GitHub issues.
4.626.5
HLE
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
0.021.4
Metr Time Horizons
40.175.0
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
6.486.1
Fortress
13.013.6
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
13.054.4
VPCT
VPCT (Visual Pattern Completion Test) · tests visual reasoning and pattern recognition by having models complete visual sequences and transformations.
0.010.0
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
31.063.7
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Anthropic logoClaude 3.5 Sonnet————
Anthropic logoClaude Opus 4.5$5.00$25.00200K tokens (~100 books)$100.00
People also compared